mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 09:24:43 +08:00
* feat(tts): voxcpm voice-design types + deterministic voice id helpers (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): emit + persist per-agent voiceDesign descriptor (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): voxcpm voice registration backend client (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): voxcpm-voice ensure/register endpoint (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): client auto-voice registration + reference-clip cache (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): reference registered voice id in vLLM-Omni speech (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): thread voiceDesign + backend through tts call sites (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): match vLLM-Omni voice registration contract (#670) Live e2e against the real backend revealed the assumed multipart contract was wrong: POST /v1/audio/voices needs name + consent + audio_sample (not voice_id/file), and there is no per-name GET (405) — existence must list /v1/audio/voices and check membership. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(tts): make auto-voice register-once pattern provider-neutral (#670) The descriptor + register-once/reference-by-id pattern is not VoxCPM-specific. De-couple it into a provider-neutral seam so other registration-capable providers (ElevenLabs/MiniMax/Doubao voice cloning, …) can plug in: - lib/audio/voice-design.ts: VoiceDesign + buildVoiceDesignPrompt / normalizeVoiceDesign / getDeterministicVoiceId (neutral 'auto-<hash>' id, namespaced by providerId). - lib/audio/voice-registration.ts: VoiceRegistrationAdapter interface + providerId->adapter registry + supportsVoiceRegistration. - lib/audio/voice-registration-client.ts: neutral ensureRegisteredVoice() + IndexedDB clip cache. - app/api/generate/voice (replaces .../voxcpm-voice): dispatches by providerId. - voxcpm-registration.ts becomes the VoxCPM adapter (sole registered provider). - AgentConfig.voiceDesign + DB cache table renamed neutral. Behavior-preserving; VoxCPM-specific bits (inline (prompt)text, backend kinds, capability gate) stay in the voxcpm modules. Full suite green (674). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): build the voice clip Blob from an inline Uint8Array (#670) A factored-out helper returning Uint8Array widened to Uint8Array<ArrayBufferLike>, which next build (stricter than bare tsc) rejects as a BlobPart. Inline the buffer like the rest of the repo does. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): managed TTS providers resolve model server-side (#670) A server-managed TTS provider's model was still client-driven, but the managed-provider settings UI hides the model field — so a backend whose model id isn't the client default (e.g. VoxCPM/vLLM-Omni serving the full model path) always 500'd. resolveTTSModel() makes the model authoritative from server config (${PREFIX}_MODELS, first entry) for managed providers, like key/baseUrl; unmanaged/unconfigured providers keep the client model unchanged. Used by the tts and voice routes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): single agent-voice resolver + stable teacher narration voice (#670) Route all TTS voice resolution (narration, discussion, preview) through one resolveAgentVoiceOptions(agent, ...) that reads the agent profile and registers + references the voice by id. Eagerly warm up generated agents' voices on save. Critically fixes the teacher narration drifting (male/female jumps): the registry is always seeded with DEFAULT_AGENTS, so the old narration lookup find(role==='teacher') returned the default teacher (no voiceDesign) instead of the generated one — so narration never registered a voice and fell back to the inline prompt. pickNarratorAgent() now prefers the teacher carrying a voiceDesign. Also: drop language from the deterministic voice id (descriptor already encodes it) so narration (directive) and discussion (locale) resolve the same id; log the effective registeredVoiceId. Regression tests for pickNarratorAgent. Verified e2e: agent-profiles -> 1 voice registration -> 16 narration TTS all referencing the same registeredVoiceId; voice present on the backend. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): fall back to persona as the voice seed for agents without a voiceDesign (#670) Preset/default agents carry no LLM voiceDesign. Rather than add a static field, derive the bootstrap descriptor from the agent's persona when voiceDesign is absent, so they still register one stable reference voice (stable-but-generic; persona is not a vocal spec). Generated agents keep their LLM voiceDesign. Also hardens replay: stage snapshots drop voiceDesign but keep persona. Unit-tested: resolveAgentVoiceOptions uses real voiceDesign when present, persona otherwise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): match Auto Voice by its localized label in the agent voice picker (#670) The picker filtered on voice.name ('Auto Voice'), but Auto Voice is shown via its localized label (自动音色). Searching '自动' returned '没有匹配音色'. Match the localized label for the Auto Voice option so it's findable in any language. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): address code-review round 1 (#670) - strip parentheses in the voice-design prompt so a paren in the descriptor/ persona can't corrupt the (prompt)text bootstrap delimiter - gate POST /api/generate/voice on isServerTTSProviderDisabled (#665): a force-disabled provider was off for the tts route but not this sibling - check voiceExists before re-registering a client-cached clip, so a cached voice that's still live isn't needlessly re-uploaded every session - dedup concurrent ensureRegisteredVoice calls via an in-flight promise map (eager warm-up + first utterance no longer double bootstrap/register) - warm up only the narrator (teacher), not every generated agent, to avoid synthesizing voices at save time for agents that may never speak - collapse the duplicated vLLM-Omni speech payload into one object Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(tts): prettier formatting (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): key auto-voice memo by (voiceId, backend) so a base-URL change re-registers (#670) Addresses review (cosarah): registeredThisSession/inFlight were keyed by voiceId alone, so switching the VoxCPM base URL mid-session made ensureRegisteredVoice short-circuit and return an id registered only on the old backend — TTS then sent that stale registeredVoiceId and skipped the inline fallback, failing with voice-not-found on the new backend. Memo key now includes the base URL; the IndexedDB clip cache stays keyed by voiceId (the reference clip is backend-independent and reused to re-register elsewhere). Regression test added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): include API key in the auto-voice memo key (#670) Addresses review (cosarah, round 2): memoKeyFor was keyed by (voiceId, baseUrl) but not the API key. Registration/existence checks and speech calls are auth-scoped, so switching account/key on the same base URL could reuse a registeredVoiceId from the old credentials and skip re-validation. Memo key now includes the API key (in-memory only, never persisted/logged). Regression test covers the key-switch case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
杨慎
parent
893f080f51
commit
d7c0d80afe
@@ -12,6 +12,7 @@ import { createLogger } from '@/lib/logger';
|
||||
import { apiError, apiSuccess } from '@/lib/server/api-response';
|
||||
import { resolveModelFromRequest } from '@/lib/server/resolve-model';
|
||||
import { AGENT_COLOR_PALETTE } from '@/lib/constants/agent-defaults';
|
||||
import { normalizeVoiceDesign } from '@/lib/audio/voice-design';
|
||||
|
||||
const log = createLogger('Agent Profiles API');
|
||||
|
||||
@@ -128,6 +129,10 @@ Requirements:
|
||||
- Use the "path" value as the avatar field in the output
|
||||
- Each agent must be assigned one color from this list: ${JSON.stringify(AGENT_COLOR_PALETTE)}
|
||||
- Each agent must have a different color
|
||||
- Each agent needs a "voiceDesign" object describing their VOCAL identity (not personality), written following the language directive and consistent with the persona, as three short comma-free phrases:
|
||||
- "identity": gender + age + role (e.g. "middle-aged male teacher")
|
||||
- "texture": pitch + vocal quality (e.g. "warm low-pitched slightly husky")
|
||||
- "delivery": emotion + pace (e.g. "calm measured encouraging")
|
||||
${voicePrompt}
|
||||
|
||||
Return a JSON object with this exact structure:
|
||||
@@ -137,6 +142,7 @@ Return a JSON object with this exact structure:
|
||||
"name": "string",
|
||||
"role": "teacher" | "assistant" | "student",
|
||||
"persona": "string (2-3 sentences)",
|
||||
"voiceDesign": { "identity": "string", "texture": "string", "delivery": "string" },
|
||||
"avatar": "string (from available list)",
|
||||
"color": "string (hex color from palette)",
|
||||
"priority": number (10 for teacher, 7 for assistant, 4-6 for student)${voiceJsonField}
|
||||
@@ -170,6 +176,7 @@ Return a JSON object with this exact structure:
|
||||
color: string;
|
||||
priority: number;
|
||||
voice?: string;
|
||||
voiceDesign?: unknown;
|
||||
}>;
|
||||
};
|
||||
|
||||
@@ -211,6 +218,8 @@ Return a JSON object with this exact structure:
|
||||
}
|
||||
}
|
||||
|
||||
const voiceDesign = normalizeVoiceDesign(agent.voiceDesign);
|
||||
|
||||
return {
|
||||
id: `gen-${nanoid(8)}`,
|
||||
name: agent.name,
|
||||
@@ -221,6 +230,7 @@ Return a JSON object with this exact structure:
|
||||
priority:
|
||||
agent.priority ?? (agent.role === 'teacher' ? 10 : agent.role === 'assistant' ? 7 : 5),
|
||||
...(voiceConfig ? { voiceConfig } : {}),
|
||||
...(voiceDesign ? { voiceDesign } : {}),
|
||||
};
|
||||
});
|
||||
|
||||
|
||||
@@ -14,6 +14,7 @@ import {
|
||||
isServerTTSProviderDisabled,
|
||||
resolveTTSApiKey,
|
||||
resolveTTSBaseUrl,
|
||||
resolveTTSModel,
|
||||
} from '@/lib/server/provider-config';
|
||||
import type { TTSProviderId } from '@/lib/audio/types';
|
||||
import { createLogger } from '@/lib/logger';
|
||||
@@ -68,10 +69,15 @@ export async function POST(req: NextRequest) {
|
||||
|
||||
const voxcpmVoicePrompt =
|
||||
typeof ttsProviderOptions?.voicePrompt === 'string' ? ttsProviderOptions.voicePrompt : '';
|
||||
const voxcpmRegisteredVoiceId =
|
||||
typeof ttsProviderOptions?.registeredVoiceId === 'string'
|
||||
? ttsProviderOptions.registeredVoiceId
|
||||
: '';
|
||||
if (
|
||||
ttsProviderId === VOXCPM_TTS_PROVIDER_ID &&
|
||||
ttsVoice === VOXCPM_AUTO_VOICE_ID &&
|
||||
!voxcpmVoicePrompt.trim()
|
||||
!voxcpmVoicePrompt.trim() &&
|
||||
!voxcpmRegisteredVoiceId.trim()
|
||||
) {
|
||||
return apiError(
|
||||
'VOXCPM_AUTO_VOICE_REQUIRES_CONTEXT',
|
||||
@@ -93,10 +99,10 @@ export async function POST(req: NextRequest) {
|
||||
const apiKey = resolveTTSApiKey(ttsProviderId, managed ? undefined : ttsApiKey || undefined);
|
||||
const baseUrl = resolveTTSBaseUrl(ttsProviderId, clientBaseUrl);
|
||||
|
||||
// Build TTS config
|
||||
// Build TTS config (managed providers may pin the model server-side)
|
||||
const config = {
|
||||
providerId: ttsProviderId as TTSProviderId,
|
||||
modelId: ttsModelId,
|
||||
modelId: resolveTTSModel(ttsProviderId, ttsModelId),
|
||||
voice: ttsVoice,
|
||||
speed: ttsSpeed ?? 1.0,
|
||||
apiKey,
|
||||
@@ -105,7 +111,8 @@ export async function POST(req: NextRequest) {
|
||||
};
|
||||
|
||||
log.info(
|
||||
`Generating TTS: provider=${ttsProviderId}, model=${ttsModelId || 'default'}, voice=${ttsVoice}, audioId=${audioId}, textLen=${text.length}`,
|
||||
`Generating TTS: provider=${ttsProviderId}, model=${config.modelId || 'default'}, voice=${ttsVoice}, ` +
|
||||
`registeredVoiceId=${voxcpmRegisteredVoiceId || 'none'}, audioId=${audioId}, textLen=${text.length}`,
|
||||
);
|
||||
|
||||
// Generate audio
|
||||
|
||||
@@ -0,0 +1,153 @@
|
||||
/**
|
||||
* Auto-voice registration API (provider-neutral).
|
||||
*
|
||||
* Idempotently ensures an agent's deterministic voice id is registered on the
|
||||
* selected TTS provider's backend so later TTS can reference it by id (stable
|
||||
* timbre, lean payload). Dispatches to the provider's VoiceRegistrationAdapter;
|
||||
* no provider is named here. Folds bootstrap + register + existence-check +
|
||||
* register-on-invalid into one call:
|
||||
* - client supplies a cached reference clip → (re)register it under voiceId;
|
||||
* - else if the voice already exists → no-op;
|
||||
* - else synthesize the descriptor once, register, and return the clip so the
|
||||
* client can cache it.
|
||||
*
|
||||
* POST /api/generate/voice
|
||||
*/
|
||||
|
||||
import { NextRequest } from 'next/server';
|
||||
import {
|
||||
isServerConfiguredProvider,
|
||||
isServerTTSProviderDisabled,
|
||||
resolveTTSApiKey,
|
||||
resolveTTSBaseUrl,
|
||||
resolveTTSModel,
|
||||
} from '@/lib/server/provider-config';
|
||||
import { createLogger } from '@/lib/logger';
|
||||
import { apiError, apiSuccess } from '@/lib/server/api-response';
|
||||
import { validateUrlForSSRF } from '@/lib/server/ssrf-guard';
|
||||
import { normalizeVoiceDesign } from '@/lib/audio/voice-design';
|
||||
import {
|
||||
getVoiceRegistrationAdapter,
|
||||
type VoiceRegistrationConfig,
|
||||
} from '@/lib/audio/voice-registration';
|
||||
|
||||
const log = createLogger('Voice Registration API');
|
||||
|
||||
export const maxDuration = 30;
|
||||
|
||||
export async function POST(req: NextRequest) {
|
||||
let providerId: string | undefined;
|
||||
let voiceId: string | undefined;
|
||||
try {
|
||||
const body = (await req.json()) as {
|
||||
providerId?: string;
|
||||
voiceId?: string;
|
||||
descriptor?: unknown;
|
||||
language?: string;
|
||||
referenceAudioBase64?: string;
|
||||
mimeType?: string;
|
||||
ttsApiKey?: string;
|
||||
ttsBaseUrl?: string;
|
||||
ttsModelId?: string;
|
||||
};
|
||||
providerId = typeof body.providerId === 'string' ? body.providerId : undefined;
|
||||
voiceId = typeof body.voiceId === 'string' ? body.voiceId.trim() : undefined;
|
||||
const design = normalizeVoiceDesign(body.descriptor);
|
||||
|
||||
if (!providerId) {
|
||||
return apiError('MISSING_REQUIRED_FIELD', 400, 'providerId is required');
|
||||
}
|
||||
if (!voiceId) {
|
||||
return apiError('MISSING_REQUIRED_FIELD', 400, 'voiceId is required');
|
||||
}
|
||||
if (!design && !body.referenceAudioBase64) {
|
||||
return apiError(
|
||||
'MISSING_REQUIRED_FIELD',
|
||||
400,
|
||||
'descriptor or referenceAudioBase64 is required',
|
||||
);
|
||||
}
|
||||
|
||||
// A server-force-disabled provider is off for everyone (#665), same as the TTS route.
|
||||
if (isServerTTSProviderDisabled(providerId)) {
|
||||
return apiError('PROVIDER_DISABLED', 403, 'This TTS provider is disabled by the server');
|
||||
}
|
||||
|
||||
const adapter = getVoiceRegistrationAdapter(providerId);
|
||||
if (!adapter) {
|
||||
return apiError(
|
||||
'INVALID_REQUEST',
|
||||
400,
|
||||
`Provider "${providerId}" does not support voice registration`,
|
||||
);
|
||||
}
|
||||
|
||||
// Managed providers are admin-owned: ignore any client-sent key/baseUrl.
|
||||
const managed = isServerConfiguredProvider('tts', providerId);
|
||||
const clientBaseUrl = managed ? undefined : body.ttsBaseUrl || undefined;
|
||||
if (clientBaseUrl) {
|
||||
const ssrfError = await validateUrlForSSRF(clientBaseUrl);
|
||||
if (ssrfError) {
|
||||
return apiError('INVALID_URL', 403, ssrfError);
|
||||
}
|
||||
}
|
||||
|
||||
const apiKey = resolveTTSApiKey(providerId, managed ? undefined : body.ttsApiKey || undefined);
|
||||
const baseUrl = resolveTTSBaseUrl(providerId, clientBaseUrl);
|
||||
if (!baseUrl) {
|
||||
return apiError('MISSING_REQUIRED_FIELD', 400, 'TTS base URL is required');
|
||||
}
|
||||
|
||||
const cfg: VoiceRegistrationConfig = {
|
||||
baseUrl,
|
||||
apiKey,
|
||||
model: resolveTTSModel(providerId, body.ttsModelId),
|
||||
};
|
||||
|
||||
// Already registered → no-op (also avoids a redundant re-register when the
|
||||
// client offered a cached clip but the voice is still live on the backend).
|
||||
if (await adapter.voiceExists(cfg, voiceId)) {
|
||||
return apiSuccess({ voiceId, registered: true });
|
||||
}
|
||||
|
||||
// Not present, but the client has the cached reference clip → re-register it
|
||||
// (register-on-invalid; preserves the original timbre instead of re-synthesizing).
|
||||
if (body.referenceAudioBase64) {
|
||||
await adapter.registerVoice(cfg, {
|
||||
voiceId,
|
||||
referenceAudioBase64: body.referenceAudioBase64,
|
||||
mimeType: body.mimeType,
|
||||
});
|
||||
return apiSuccess({ voiceId, registered: true });
|
||||
}
|
||||
|
||||
// First use → bootstrap-synthesize the descriptor, register, return the clip.
|
||||
const clip = await adapter.bootstrapReferenceClip(cfg, {
|
||||
design: design!,
|
||||
language: body.language,
|
||||
});
|
||||
await adapter.registerVoice(cfg, {
|
||||
voiceId,
|
||||
referenceAudioBase64: clip.referenceAudioBase64,
|
||||
mimeType: clip.mimeType,
|
||||
});
|
||||
|
||||
log.info(`Registered auto voice ${voiceId} for provider ${providerId}`);
|
||||
return apiSuccess({
|
||||
voiceId,
|
||||
registered: true,
|
||||
referenceAudioBase64: clip.referenceAudioBase64,
|
||||
mimeType: clip.mimeType,
|
||||
});
|
||||
} catch (error) {
|
||||
log.error(
|
||||
`Voice registration failed [provider=${providerId ?? 'unknown'}, voiceId=${voiceId ?? 'unknown'}]:`,
|
||||
error,
|
||||
);
|
||||
return apiError(
|
||||
'GENERATION_FAILED',
|
||||
500,
|
||||
error instanceof Error ? error.message : String(error),
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -14,7 +14,8 @@ import { useSettingsStore } from '@/lib/store/settings';
|
||||
import { useAgentRegistry } from '@/lib/orchestration/registry/store';
|
||||
import { getEnabledProvidersWithVoices } from '@/lib/audio/voice-resolver';
|
||||
import { isTTSProviderEnabled } from '@/lib/audio/provider-enablement';
|
||||
import { getVoxCPMProviderOptions, useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { resolveAgentVoiceOptions, pickNarratorAgent } from '@/lib/audio/agent-voice';
|
||||
import { useI18n } from '@/lib/hooks/use-i18n';
|
||||
import {
|
||||
loadImageMapping,
|
||||
@@ -879,16 +880,14 @@ function GenerationPreviewContent() {
|
||||
)
|
||||
) {
|
||||
const ttsProviderConfig = settings.ttsProvidersConfig?.[settings.ttsProviderId];
|
||||
const providerOptions =
|
||||
settings.ttsProviderId === 'voxcpm-tts'
|
||||
? {
|
||||
...(ttsProviderConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(settings.ttsVoice, {
|
||||
role: 'teacher',
|
||||
language: languageDirective,
|
||||
})),
|
||||
}
|
||||
: undefined;
|
||||
// Narration uses the teacher agent's voice (single resolver → stable timbre).
|
||||
const teacherAgent = pickNarratorAgent(useAgentRegistry.getState().listAgents());
|
||||
const providerOptions = await resolveAgentVoiceOptions(teacherAgent, {
|
||||
providerId: settings.ttsProviderId,
|
||||
providerConfig: ttsProviderConfig,
|
||||
voiceId: settings.ttsVoice,
|
||||
language: languageDirective,
|
||||
});
|
||||
const speechActions = (data.scene.actions || []).filter(
|
||||
(a: { type: string; text?: string }) => a.type === 'speech' && a.text,
|
||||
);
|
||||
|
||||
@@ -10,7 +10,8 @@ import { useSettingsStore } from '@/lib/store/settings';
|
||||
import { useAgentRegistry } from '@/lib/orchestration/registry/store';
|
||||
import { resolveAgentVoice, getSelectableProvidersWithVoices } from '@/lib/audio/voice-resolver';
|
||||
import { playBrowserTTSPreview } from '@/lib/audio/browser-tts-preview';
|
||||
import { getVoxCPMProviderOptions, useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { resolveAgentVoiceOptions } from '@/lib/audio/agent-voice';
|
||||
import { VOXCPM_AUTO_VOICE_ID, VOXCPM_TTS_PROVIDER_ID } from '@/lib/audio/voxcpm';
|
||||
import {
|
||||
Sparkles,
|
||||
@@ -31,7 +32,11 @@ function matchesVoiceQuery(value: string | undefined, query: string): boolean {
|
||||
return !!value?.toLowerCase().includes(query);
|
||||
}
|
||||
|
||||
function getFilteredModelGroups(provider: ProviderWithVoices, query: string) {
|
||||
function getFilteredModelGroups(
|
||||
provider: ProviderWithVoices,
|
||||
query: string,
|
||||
autoVoiceLabel?: string,
|
||||
) {
|
||||
const normalizedQuery = query.trim().toLowerCase();
|
||||
if (!normalizedQuery) return provider.modelGroups;
|
||||
|
||||
@@ -47,7 +52,9 @@ function getFilteredModelGroups(provider: ProviderWithVoices, query: string) {
|
||||
groupMatches ||
|
||||
matchesVoiceQuery(voice.name, normalizedQuery) ||
|
||||
matchesVoiceQuery(voice.id, normalizedQuery) ||
|
||||
matchesVoiceQuery(voice.language, normalizedQuery),
|
||||
matchesVoiceQuery(voice.language, normalizedQuery) ||
|
||||
// Auto Voice is shown by its localized label, not voice.name — match it too.
|
||||
(voice.id === VOXCPM_AUTO_VOICE_ID && matchesVoiceQuery(autoVoiceLabel, normalizedQuery)),
|
||||
);
|
||||
return { ...group, voices };
|
||||
})
|
||||
@@ -82,7 +89,7 @@ function AgentVoicePill({
|
||||
const visibleProviderGroups = availableProviders
|
||||
.map((provider) => ({
|
||||
provider,
|
||||
groups: getFilteredModelGroups(provider, voiceQuery),
|
||||
groups: getFilteredModelGroups(provider, voiceQuery, t('settings.voxcpmAutoVoice')),
|
||||
}))
|
||||
.filter(({ groups }) => groups.length > 0);
|
||||
|
||||
@@ -139,18 +146,12 @@ function AgentVoicePill({
|
||||
const controller = new AbortController();
|
||||
previewAbortRef.current = controller;
|
||||
const providerConfig = ttsProvidersConfig[providerId];
|
||||
const providerOptions =
|
||||
providerId === 'voxcpm-tts'
|
||||
? {
|
||||
...(providerConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(voiceId, {
|
||||
agentName: agent.name,
|
||||
role: agent.role,
|
||||
persona: agent.persona,
|
||||
locale,
|
||||
})),
|
||||
}
|
||||
: undefined;
|
||||
const providerOptions = await resolveAgentVoiceOptions(agent, {
|
||||
providerId,
|
||||
providerConfig: { ...providerConfig, modelId: modelId || providerConfig?.modelId },
|
||||
voiceId,
|
||||
language: locale,
|
||||
});
|
||||
const res = await fetch('/api/generate/tts', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
@@ -367,7 +368,7 @@ function TeacherVoicePill({
|
||||
const visibleProviderGroups = availableProviders
|
||||
.map((provider) => ({
|
||||
provider,
|
||||
groups: getFilteredModelGroups(provider, voiceQuery),
|
||||
groups: getFilteredModelGroups(provider, voiceQuery, t('settings.voxcpmAutoVoice')),
|
||||
}))
|
||||
.filter(({ groups }) => groups.length > 0);
|
||||
|
||||
@@ -425,17 +426,12 @@ function TeacherVoicePill({
|
||||
const controller = new AbortController();
|
||||
previewAbortRef.current = controller;
|
||||
const providerConfig = ttsProvidersConfig[providerId];
|
||||
const providerOptions =
|
||||
providerId === 'voxcpm-tts'
|
||||
? {
|
||||
...(providerConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(voiceId, {
|
||||
agentName: 'Teacher',
|
||||
role: 'teacher',
|
||||
locale,
|
||||
})),
|
||||
}
|
||||
: undefined;
|
||||
const providerOptions = await resolveAgentVoiceOptions(undefined, {
|
||||
providerId,
|
||||
providerConfig: { ...providerConfig, modelId: modelId || providerConfig?.modelId },
|
||||
voiceId,
|
||||
language: locale,
|
||||
});
|
||||
const res = await fetch('/api/generate/tts', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
'use client';
|
||||
|
||||
/**
|
||||
* Single source of truth for "an agent's TTS voice".
|
||||
*
|
||||
* Every TTS path (lecture narration, multi-agent discussion, voice preview,
|
||||
* settings test) resolves voice options through `resolveAgentVoiceOptions`,
|
||||
* which reads the agent profile (`voiceConfig` + `voiceDesign`) and, for
|
||||
* registration-capable providers, ensures the auto voice is registered and
|
||||
* referenced by id (stable timbre) — otherwise falls back to the inline
|
||||
* voice-design prompt. There is no second code path to drift out of sync.
|
||||
*/
|
||||
|
||||
import type { AgentConfig } from '@/lib/orchestration/registry/types';
|
||||
import type { VoiceDesign } from '@/lib/audio/voice-design';
|
||||
import { getVoxCPMProviderOptions } from '@/lib/audio/voxcpm-voices';
|
||||
import {
|
||||
VOXCPM_AUTO_VOICE_ID,
|
||||
VOXCPM_TTS_PROVIDER_ID,
|
||||
normalizeVoxCPMBackend,
|
||||
} from '@/lib/audio/voxcpm';
|
||||
import { useSettingsStore } from '@/lib/store/settings';
|
||||
|
||||
interface TTSProviderConfigShape {
|
||||
apiKey?: string;
|
||||
baseUrl?: string;
|
||||
customDefaultBaseUrl?: string;
|
||||
modelId?: string;
|
||||
providerOptions?: Record<string, unknown>;
|
||||
}
|
||||
|
||||
export interface AgentVoiceResolveOptions {
|
||||
providerId: string;
|
||||
providerConfig?: TTSProviderConfigShape;
|
||||
voiceId: string;
|
||||
/** Course language / locale — only selects the one-time bootstrap sample sentence. */
|
||||
language?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Pick the agent whose voice narration should use (the teacher).
|
||||
*
|
||||
* The registry is always seeded with the DEFAULT agents, so a plain
|
||||
* `find(role === 'teacher')` returns the default teacher (no voiceDesign) even
|
||||
* when a generated classroom is active. Prefer a teacher that actually carries a
|
||||
* voiceDesign (the generated one) so narration registers a real voice instead of
|
||||
* drifting on the inline prompt; fall back to any teacher.
|
||||
*/
|
||||
export function pickNarratorAgent(agents: AgentConfig[]): AgentConfig | undefined {
|
||||
return (
|
||||
agents.find((a) => a.role === 'teacher' && a.voiceDesign) ??
|
||||
agents.find((a) => a.role === 'teacher')
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* The descriptor used to bootstrap an agent's voice: the real `voiceDesign`
|
||||
* when present (generated agents), otherwise the persona as a fallback seed.
|
||||
* Persona is not a vocal description, so the resulting voice is stable but
|
||||
* generic — good enough to register a consistent reference clip (no drift),
|
||||
* pending a real descriptor if quality matters for that agent.
|
||||
*/
|
||||
function effectiveVoiceDesign(agent: AgentConfig | undefined): VoiceDesign | undefined {
|
||||
if (agent?.voiceDesign) return agent.voiceDesign;
|
||||
const persona = agent?.persona?.trim();
|
||||
return persona ? { identity: persona, texture: '', delivery: '' } : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Produce the `ttsProviderOptions` to send to /api/generate/tts for `agent`
|
||||
* (pass the teacher agent for narration, the speaking agent for discussion,
|
||||
* or undefined when there is no agent). Returns undefined for providers with
|
||||
* no special options.
|
||||
*/
|
||||
export async function resolveAgentVoiceOptions(
|
||||
agent: AgentConfig | undefined,
|
||||
opts: AgentVoiceResolveOptions,
|
||||
): Promise<Record<string, unknown> | undefined> {
|
||||
if (opts.providerId !== VOXCPM_TTS_PROVIDER_ID) return undefined;
|
||||
return {
|
||||
...(opts.providerConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(
|
||||
opts.voiceId,
|
||||
{
|
||||
agentName: agent?.name,
|
||||
role: agent?.role ?? 'teacher',
|
||||
persona: agent?.persona,
|
||||
voiceDesign: effectiveVoiceDesign(agent),
|
||||
language: opts.language,
|
||||
backend: normalizeVoxCPMBackend(opts.providerConfig?.providerOptions?.backend),
|
||||
},
|
||||
{
|
||||
ttsApiKey: opts.providerConfig?.apiKey || undefined,
|
||||
ttsBaseUrl:
|
||||
opts.providerConfig?.baseUrl || opts.providerConfig?.customDefaultBaseUrl || undefined,
|
||||
ttsModelId: opts.providerConfig?.modelId,
|
||||
},
|
||||
)),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Eager warm-up: right after generated agents are saved, pre-register the
|
||||
* narrator's (teacher's) auto voice using the SAME idempotent ensure as the TTS
|
||||
* path, so the first spoken line is already stable. Only the narrator is warmed:
|
||||
* it always speaks (lecture narration), whereas discussion agents may never be
|
||||
* selected, so warming all of them would synthesize voices that go unused.
|
||||
* Fire-and-forget; the on-use ensure remains the correctness path for the rest.
|
||||
*/
|
||||
export function warmUpAgentVoices(agents: AgentConfig[]): void {
|
||||
const settings = useSettingsStore.getState();
|
||||
const providerId = settings.ttsProviderId;
|
||||
if (providerId !== VOXCPM_TTS_PROVIDER_ID) return;
|
||||
const providerConfig = settings.ttsProvidersConfig?.[providerId];
|
||||
|
||||
const narrator = pickNarratorAgent(agents);
|
||||
if (!narrator || !effectiveVoiceDesign(narrator)) return;
|
||||
void resolveAgentVoiceOptions(narrator, {
|
||||
providerId,
|
||||
providerConfig,
|
||||
voiceId: VOXCPM_AUTO_VOICE_ID,
|
||||
}).catch(() => undefined);
|
||||
}
|
||||
@@ -297,7 +297,9 @@ async function generateVoxCPMTTS(
|
||||
(config.voice && config.voice !== 'default' && config.voice !== VOXCPM_AUTO_VOICE_ID
|
||||
? config.voice
|
||||
: undefined);
|
||||
if (config.voice === VOXCPM_AUTO_VOICE_ID && !voicePrompt) {
|
||||
// A registered voice carries timbre by id, so no voice prompt is required.
|
||||
const registeredVoiceId = options.registeredVoiceId?.trim() || undefined;
|
||||
if (config.voice === VOXCPM_AUTO_VOICE_ID && !voicePrompt && !registeredVoiceId) {
|
||||
throw new Error('VoxCPM Auto Voice requires agent context');
|
||||
}
|
||||
const cfgValue = options.cfgValue ?? 2.0;
|
||||
@@ -308,6 +310,8 @@ async function generateVoxCPMTTS(
|
||||
|
||||
const request = {
|
||||
targetText: usePromptContinuation ? text : buildVoxCPMTargetText(text, voicePrompt),
|
||||
rawText: text,
|
||||
registeredVoiceId,
|
||||
voicePrompt,
|
||||
promptText: options.promptText,
|
||||
cfgValue,
|
||||
@@ -388,6 +392,8 @@ async function postVoxCPMVLLMOmni(
|
||||
baseUrl: string,
|
||||
params: {
|
||||
targetText: string;
|
||||
rawText?: string;
|
||||
registeredVoiceId?: string;
|
||||
promptText?: string;
|
||||
referenceAudioBase64?: string;
|
||||
referenceAudioMimeType?: string;
|
||||
@@ -398,13 +404,17 @@ async function postVoxCPMVLLMOmni(
|
||||
const payload: Record<string, unknown> = {
|
||||
model: getVLLMOmniModelId(config),
|
||||
input: params.targetText,
|
||||
// VoxCPM2's vLLM-Omni adapter currently ignores named voices; prompts/ref_audio carry voice identity.
|
||||
voice: 'default',
|
||||
response_format: 'wav',
|
||||
stream: false,
|
||||
};
|
||||
|
||||
if (params.referenceAudioBase64) {
|
||||
if (params.registeredVoiceId) {
|
||||
// A registered voice carries timbre by id (pre-encoded latents): reference it
|
||||
// directly and send the raw text — no inline voice-design prompt or ref_audio.
|
||||
payload.voice = params.registeredVoiceId;
|
||||
payload.input = params.rawText ?? params.targetText;
|
||||
} else if (params.referenceAudioBase64) {
|
||||
const referenceAudio = getVoxCPMDataAudioUrl(
|
||||
params.referenceAudioBase64,
|
||||
params.referenceAudioMimeType,
|
||||
|
||||
@@ -0,0 +1,84 @@
|
||||
/**
|
||||
* Provider-neutral per-agent voice design.
|
||||
*
|
||||
* A `VoiceDesign` describes an agent's vocal identity (not personality) as a
|
||||
* 3-layer recipe. It is consumed by any TTS provider: as an inline voice
|
||||
* prompt where supported, or as the seed for a registered/cloned voice
|
||||
* (see `voice-registration.ts`). Nothing here is VoxCPM-specific.
|
||||
*/
|
||||
|
||||
export interface VoiceDesign {
|
||||
identity: string; // gender / age / role
|
||||
texture: string; // pitch / vocal quality
|
||||
delivery: string; // emotion / pace
|
||||
}
|
||||
|
||||
const VOICE_DESIGN_PROMPT_MAX_CHARS = 200;
|
||||
|
||||
/** Prefix for deterministic auto-voice ids (provider-neutral, backend-name-safe). */
|
||||
export const AUTO_VOICE_ID_PREFIX = 'auto-' as const;
|
||||
|
||||
function sanitizeVoiceDesignPart(value?: string): string {
|
||||
return (
|
||||
(value || '')
|
||||
.replace(/[\p{C}]+/gu, ' ')
|
||||
// Strip parentheses: VoxCPM uses `(prompt)text` delimiters, so a paren in the
|
||||
// descriptor/persona would corrupt the bootstrap synthesis prompt.
|
||||
.replace(/[()()]/gu, ' ')
|
||||
.replace(/\s+/gu, ' ')
|
||||
.trim()
|
||||
.slice(0, VOICE_DESIGN_PROMPT_MAX_CHARS)
|
||||
.trim()
|
||||
);
|
||||
}
|
||||
|
||||
/** Compose the 3 layers into one comma-joined prompt, dropping blank layers. */
|
||||
export function buildVoiceDesignPrompt(design: VoiceDesign): string {
|
||||
return [design.identity, design.texture, design.delivery]
|
||||
.map((part) => sanitizeVoiceDesignPart(part))
|
||||
.filter(Boolean)
|
||||
.join(', ');
|
||||
}
|
||||
|
||||
/** Coerce an arbitrary (LLM-produced) value into a VoiceDesign, or undefined. */
|
||||
export function normalizeVoiceDesign(raw: unknown): VoiceDesign | undefined {
|
||||
if (!raw || typeof raw !== 'object') return undefined;
|
||||
const record = raw as Record<string, unknown>;
|
||||
const pick = (value: unknown) => (typeof value === 'string' ? value.trim() : '');
|
||||
const design = {
|
||||
identity: pick(record.identity),
|
||||
texture: pick(record.texture),
|
||||
delivery: pick(record.delivery),
|
||||
};
|
||||
if (!design.identity && !design.texture && !design.delivery) return undefined;
|
||||
return design;
|
||||
}
|
||||
|
||||
/**
|
||||
* Deterministic voice id derived from the descriptor (+ provider + model).
|
||||
* Stable across re-synthesis, recomputable anywhere from the descriptor on the
|
||||
* agent, and namespaced by provider so a shared registry can't collide.
|
||||
*
|
||||
* Note: language is intentionally NOT part of the id — the descriptor text is
|
||||
* already written in the course language, and language only selects the one-time
|
||||
* bootstrap sample sentence (which affects neither output language nor timbre).
|
||||
* Keeping it out means every TTS path (narration passes a directive, discussion
|
||||
* passes a locale) resolves to the SAME id for the same agent.
|
||||
*/
|
||||
export async function getDeterministicVoiceId(
|
||||
design: VoiceDesign,
|
||||
opts: { providerId?: string; model?: string } = {},
|
||||
): Promise<string> {
|
||||
const seed = [
|
||||
opts.providerId || '',
|
||||
design.identity,
|
||||
design.texture,
|
||||
design.delivery,
|
||||
opts.model || '',
|
||||
].join('|');
|
||||
const digest = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(seed));
|
||||
const hex = Array.from(new Uint8Array(digest))
|
||||
.map((byte) => byte.toString(16).padStart(2, '0'))
|
||||
.join('');
|
||||
return `${AUTO_VOICE_ID_PREFIX}${hex.slice(0, 16)}`;
|
||||
}
|
||||
@@ -0,0 +1,132 @@
|
||||
'use client';
|
||||
|
||||
/**
|
||||
* Provider-neutral client orchestrator for the auto-voice register-once flow.
|
||||
*
|
||||
* Given a provider id + voice design, it resolves a deterministic voice id,
|
||||
* ensures the voice is registered on the backend via `POST /api/generate/voice`
|
||||
* (which dispatches to the provider's adapter), and caches the reference clip
|
||||
* in IndexedDB so a GC'd voice can be re-registered. Callers decide *whether*
|
||||
* their provider supports registration; this module is provider-agnostic.
|
||||
*/
|
||||
|
||||
import { db } from '@/lib/utils/database';
|
||||
import { getDeterministicVoiceId, type VoiceDesign } from '@/lib/audio/voice-design';
|
||||
|
||||
export interface VoiceRegistrationRequestConfig {
|
||||
ttsApiKey?: string;
|
||||
ttsBaseUrl?: string;
|
||||
ttsModelId?: string;
|
||||
}
|
||||
|
||||
function base64ToBlob(base64: string, mimeType?: string): Blob {
|
||||
const binary = atob(base64);
|
||||
const bytes = new Uint8Array(binary.length);
|
||||
for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i);
|
||||
return new Blob([bytes], { type: mimeType || 'audio/wav' });
|
||||
}
|
||||
|
||||
async function blobToBase64(blob: Blob): Promise<string> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const reader = new FileReader();
|
||||
reader.onerror = () => reject(reader.error || new Error('Failed to read reference audio'));
|
||||
reader.onload = () => {
|
||||
const result = typeof reader.result === 'string' ? reader.result : '';
|
||||
const commaIndex = result.indexOf(',');
|
||||
resolve(commaIndex >= 0 ? result.slice(commaIndex + 1) : result);
|
||||
};
|
||||
reader.readAsDataURL(blob);
|
||||
});
|
||||
}
|
||||
|
||||
// Confirmed-registered + in-flight memos, keyed by (voiceId, backend, credential).
|
||||
// The same voiceId may be unregistered — or inaccessible — on a different backend
|
||||
// or under different credentials, so both the base URL and the API key are part
|
||||
// of the key. Otherwise switching the VoxCPM base URL or account mid-session
|
||||
// would skip re-registration and reuse an id from the old backend/credentials.
|
||||
const registeredThisSession = new Set<string>();
|
||||
const inFlight = new Map<string, Promise<string | undefined>>();
|
||||
|
||||
function memoKeyFor(voiceId: string, request: VoiceRegistrationRequestConfig): string {
|
||||
// In-memory only (never persisted or logged), so the raw key identity is fine.
|
||||
return `${voiceId}::${request.ttsBaseUrl ?? ''}::${request.ttsApiKey ?? ''}`;
|
||||
}
|
||||
|
||||
async function getCachedClip(
|
||||
voiceId: string,
|
||||
): Promise<{ base64: string; mimeType: string } | undefined> {
|
||||
const row = await db.autoVoiceCache.get(voiceId);
|
||||
if (!row) return undefined;
|
||||
return { base64: await blobToBase64(row.referenceAudio), mimeType: row.mimeType };
|
||||
}
|
||||
|
||||
/**
|
||||
* Ensure the agent's deterministic auto voice is registered for `providerId`,
|
||||
* returning its voice id (or undefined when unavailable, so callers fall back
|
||||
* to the inline voice-design prompt). Lazy + idempotent: memoized per session,
|
||||
* reference clip cached in IndexedDB. register-on-invalid is handled by the
|
||||
* endpoint's existence check, which re-registers a GC'd voice from the clip.
|
||||
*/
|
||||
export async function ensureRegisteredVoice(
|
||||
providerId: string,
|
||||
params: { voiceDesign?: VoiceDesign; language?: string },
|
||||
request: VoiceRegistrationRequestConfig,
|
||||
): Promise<string | undefined> {
|
||||
if (!params.voiceDesign) return undefined;
|
||||
|
||||
const voiceId = await getDeterministicVoiceId(params.voiceDesign, {
|
||||
providerId,
|
||||
model: request.ttsModelId,
|
||||
});
|
||||
const memoKey = memoKeyFor(voiceId, request);
|
||||
if (registeredThisSession.has(memoKey)) return voiceId;
|
||||
|
||||
// Coalesce concurrent calls for the same (voiceId, backend) into one request.
|
||||
const existing = inFlight.get(memoKey);
|
||||
if (existing) return existing;
|
||||
|
||||
const promise = registerOnce(providerId, voiceId, memoKey, params, request).finally(() =>
|
||||
inFlight.delete(memoKey),
|
||||
);
|
||||
inFlight.set(memoKey, promise);
|
||||
return promise;
|
||||
}
|
||||
|
||||
async function registerOnce(
|
||||
providerId: string,
|
||||
voiceId: string,
|
||||
memoKey: string,
|
||||
params: { voiceDesign?: VoiceDesign; language?: string },
|
||||
request: VoiceRegistrationRequestConfig,
|
||||
): Promise<string | undefined> {
|
||||
const cached = await getCachedClip(voiceId);
|
||||
const res = await fetch('/api/generate/voice', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
providerId,
|
||||
voiceId,
|
||||
descriptor: params.voiceDesign,
|
||||
language: params.language,
|
||||
referenceAudioBase64: cached?.base64,
|
||||
mimeType: cached?.mimeType,
|
||||
...request,
|
||||
}),
|
||||
});
|
||||
if (!res.ok) return undefined; // graceful fallback to the inline prompt path
|
||||
|
||||
const data = (await res.json().catch(() => ({}))) as {
|
||||
referenceAudioBase64?: string;
|
||||
mimeType?: string;
|
||||
};
|
||||
if (data.referenceAudioBase64 && !cached) {
|
||||
await db.autoVoiceCache.put({
|
||||
voiceId,
|
||||
referenceAudio: base64ToBlob(data.referenceAudioBase64, data.mimeType),
|
||||
mimeType: data.mimeType || 'audio/wav',
|
||||
updatedAt: Date.now(),
|
||||
});
|
||||
}
|
||||
registeredThisSession.add(memoKey);
|
||||
return voiceId;
|
||||
}
|
||||
@@ -0,0 +1,56 @@
|
||||
/**
|
||||
* Provider-neutral voice-registration seam.
|
||||
*
|
||||
* The "auto voice" timbre-stability pattern — synthesize a voice-design once,
|
||||
* register it, then reference it by id — is not VoxCPM-specific. Any TTS
|
||||
* backend that can register/clone a voice (VoxCPM/vLLM-Omni today; ElevenLabs,
|
||||
* MiniMax, Doubao, … later) plugs in by implementing `VoiceRegistrationAdapter`
|
||||
* and registering it below. Server routes and the client orchestrator dispatch
|
||||
* by `providerId` and never name a concrete provider.
|
||||
*/
|
||||
|
||||
import type { VoiceDesign } from '@/lib/audio/voice-design';
|
||||
import { voxcpmVoiceRegistrationAdapter } from '@/lib/audio/voxcpm-registration';
|
||||
|
||||
/** Resolved backend connection for a registration call (server-injected for managed providers). */
|
||||
export interface VoiceRegistrationConfig {
|
||||
baseUrl: string;
|
||||
apiKey?: string;
|
||||
model?: string;
|
||||
}
|
||||
|
||||
export interface VoiceRegistrationAdapter {
|
||||
/** Whether registration is available for this provider given its options (e.g. backend kind). */
|
||||
supportsRegistration(options?: Record<string, unknown>): boolean;
|
||||
/** Whether `voiceId` is already registered on the backend. */
|
||||
voiceExists(cfg: VoiceRegistrationConfig, voiceId: string): Promise<boolean>;
|
||||
/** Register (or idempotently re-register) a reference clip under `voiceId`; returns the id. */
|
||||
registerVoice(
|
||||
cfg: VoiceRegistrationConfig,
|
||||
params: { voiceId: string; referenceAudioBase64: string; mimeType?: string },
|
||||
): Promise<string>;
|
||||
/** Synthesize the voice design once into a reference clip. */
|
||||
bootstrapReferenceClip(
|
||||
cfg: VoiceRegistrationConfig,
|
||||
params: { design: VoiceDesign; language?: string },
|
||||
): Promise<{ referenceAudioBase64: string; mimeType: string }>;
|
||||
}
|
||||
|
||||
/** providerId → adapter. The only seam to touch when adding a provider. */
|
||||
const VOICE_REGISTRATION_ADAPTERS: Record<string, VoiceRegistrationAdapter> = {
|
||||
'voxcpm-tts': voxcpmVoiceRegistrationAdapter,
|
||||
};
|
||||
|
||||
export function getVoiceRegistrationAdapter(
|
||||
providerId: string,
|
||||
): VoiceRegistrationAdapter | undefined {
|
||||
return VOICE_REGISTRATION_ADAPTERS[providerId];
|
||||
}
|
||||
|
||||
/** Whether this provider supports register-once/reference-by-id for the given options. */
|
||||
export function supportsVoiceRegistration(
|
||||
providerId: string,
|
||||
options?: Record<string, unknown>,
|
||||
): boolean {
|
||||
return getVoiceRegistrationAdapter(providerId)?.supportsRegistration(options) ?? false;
|
||||
}
|
||||
@@ -0,0 +1,134 @@
|
||||
/**
|
||||
* VoxCPM voice-registration adapter (server-side) — the first concrete
|
||||
* implementation of the provider-neutral `VoiceRegistrationAdapter`.
|
||||
*
|
||||
* Drives the reference-by-id timbre-stability flow against vLLM-Omni
|
||||
* (`/v1/audio/voices`): synthesize a voice-design prompt once, register the
|
||||
* clip under a deterministic id, then later TTS references `voice=<id>`.
|
||||
*/
|
||||
|
||||
import { buildVoiceDesignPrompt, type VoiceDesign } from '@/lib/audio/voice-design';
|
||||
import {
|
||||
VOXCPM_VLLM_MODEL_ID,
|
||||
normalizeVoxCPMBackend,
|
||||
voxCPMBackendSupportsVoiceRegistration,
|
||||
} from '@/lib/audio/voxcpm';
|
||||
import type {
|
||||
VoiceRegistrationAdapter,
|
||||
VoiceRegistrationConfig,
|
||||
} from '@/lib/audio/voice-registration';
|
||||
|
||||
function v1(baseUrl: string): string {
|
||||
const clean = baseUrl.replace(/\/$/, '');
|
||||
return clean.endsWith('/v1') ? clean : `${clean}/v1`;
|
||||
}
|
||||
|
||||
function authHeaders(apiKey?: string): Record<string, string> {
|
||||
return apiKey?.trim() ? { Authorization: `Bearer ${apiKey.trim()}` } : {};
|
||||
}
|
||||
|
||||
function base64ToBlob(base64: string, mimeType?: string): Blob {
|
||||
const binary = atob(base64);
|
||||
const bytes = new Uint8Array(binary.length);
|
||||
for (let i = 0; i < binary.length; i++) bytes[i] = binary.charCodeAt(i);
|
||||
return new Blob([bytes], { type: mimeType || 'audio/wav' });
|
||||
}
|
||||
|
||||
function bytesToBase64(bytes: Uint8Array): string {
|
||||
let binary = '';
|
||||
for (let i = 0; i < bytes.length; i++) binary += String.fromCharCode(bytes[i]);
|
||||
return btoa(binary);
|
||||
}
|
||||
|
||||
/** vLLM-Omni requires a consent string on voice registration. */
|
||||
const VOXCPM_VOICE_CONSENT = 'I confirm I have the right to use this voice sample.';
|
||||
|
||||
/** A short neutral sentence used to synthesize the bootstrap reference clip. */
|
||||
const BOOTSTRAP_SENTENCE: Record<string, string> = {
|
||||
default: 'Hello, welcome to today’s lesson. Let us begin.',
|
||||
zh: '你好,欢迎来到今天的课程,我们开始吧。',
|
||||
};
|
||||
|
||||
function bootstrapSentence(language?: string): string {
|
||||
if (!language) return BOOTSTRAP_SENTENCE.default;
|
||||
const key = language.toLowerCase().split(/[-_]/)[0];
|
||||
return BOOTSTRAP_SENTENCE[key] || BOOTSTRAP_SENTENCE.default;
|
||||
}
|
||||
|
||||
/** Whether a voice id is already registered (vLLM-Omni has no per-name GET → list + membership). */
|
||||
export async function voxCPMVoiceExists(
|
||||
cfg: VoiceRegistrationConfig,
|
||||
voiceId: string,
|
||||
): Promise<boolean> {
|
||||
const res = await fetch(`${v1(cfg.baseUrl)}/audio/voices`, {
|
||||
method: 'GET',
|
||||
headers: authHeaders(cfg.apiKey),
|
||||
});
|
||||
if (!res.ok) return false;
|
||||
const data = (await res.json().catch(() => ({}))) as { voices?: unknown };
|
||||
return Array.isArray(data.voices) && data.voices.includes(voiceId);
|
||||
}
|
||||
|
||||
/** Register (or re-register, idempotently) a reference clip under `voiceId`. */
|
||||
export async function registerVoxCPMVoice(
|
||||
cfg: VoiceRegistrationConfig,
|
||||
params: { voiceId: string; referenceAudioBase64: string; mimeType?: string },
|
||||
): Promise<string> {
|
||||
const form = new FormData();
|
||||
form.set('name', params.voiceId);
|
||||
form.set('consent', VOXCPM_VOICE_CONSENT);
|
||||
form.set(
|
||||
'audio_sample',
|
||||
base64ToBlob(params.referenceAudioBase64, params.mimeType),
|
||||
`${params.voiceId}.wav`,
|
||||
);
|
||||
|
||||
const res = await fetch(`${v1(cfg.baseUrl)}/audio/voices`, {
|
||||
method: 'POST',
|
||||
headers: authHeaders(cfg.apiKey),
|
||||
body: form,
|
||||
});
|
||||
if (!res.ok) {
|
||||
throw new Error(`VoxCPM voice registration failed: ${res.status}`);
|
||||
}
|
||||
const data = (await res.json().catch(() => ({}))) as { voice?: { name?: string } };
|
||||
return data.voice?.name || params.voiceId;
|
||||
}
|
||||
|
||||
/** Synthesize the voice-design prompt once into a reference clip. */
|
||||
export async function bootstrapVoxCPMReferenceClip(
|
||||
cfg: VoiceRegistrationConfig,
|
||||
params: { design: VoiceDesign; language?: string },
|
||||
): Promise<{ referenceAudioBase64: string; mimeType: string }> {
|
||||
const prompt = buildVoiceDesignPrompt(params.design);
|
||||
const sample = bootstrapSentence(params.language);
|
||||
const res = await fetch(`${v1(cfg.baseUrl)}/audio/speech`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json; charset=utf-8', ...authHeaders(cfg.apiKey) },
|
||||
body: JSON.stringify({
|
||||
model: cfg.model || VOXCPM_VLLM_MODEL_ID,
|
||||
input: prompt ? `(${prompt})${sample}` : sample,
|
||||
voice: 'default',
|
||||
response_format: 'wav',
|
||||
stream: false,
|
||||
}),
|
||||
});
|
||||
if (!res.ok) {
|
||||
throw new Error(`VoxCPM bootstrap synthesis failed: ${res.status}`);
|
||||
}
|
||||
const bytes = new Uint8Array(await res.arrayBuffer());
|
||||
return {
|
||||
referenceAudioBase64: bytesToBase64(bytes),
|
||||
mimeType: res.headers.get('content-type') || 'audio/wav',
|
||||
};
|
||||
}
|
||||
|
||||
/** VoxCPM implementation of the provider-neutral registration adapter. */
|
||||
export const voxcpmVoiceRegistrationAdapter: VoiceRegistrationAdapter = {
|
||||
supportsRegistration(options) {
|
||||
return voxCPMBackendSupportsVoiceRegistration(normalizeVoxCPMBackend(options?.backend));
|
||||
},
|
||||
voiceExists: voxCPMVoiceExists,
|
||||
registerVoice: registerVoxCPMVoice,
|
||||
bootstrapReferenceClip: bootstrapVoxCPMReferenceClip,
|
||||
};
|
||||
@@ -10,9 +10,14 @@ import {
|
||||
buildAutoVoxCPMVoicePrompt,
|
||||
getVoxCPMProfileIdFromVoiceId,
|
||||
getVoxCPMProfileVoiceId,
|
||||
voxCPMBackendSupportsVoiceRegistration,
|
||||
type VoxCPMProviderOptions,
|
||||
type VoxCPMVoicePromptContext,
|
||||
} from '@/lib/audio/voxcpm';
|
||||
import {
|
||||
ensureRegisteredVoice,
|
||||
type VoiceRegistrationRequestConfig,
|
||||
} from '@/lib/audio/voice-registration-client';
|
||||
|
||||
export type VoxCPMVoiceProfile = VoiceProfileRecord;
|
||||
|
||||
@@ -276,11 +281,26 @@ export function useVoxCPMVoiceProfiles() {
|
||||
export async function getVoxCPMProviderOptions(
|
||||
voiceId: string,
|
||||
context?: VoxCPMVoicePromptContext,
|
||||
request?: VoiceRegistrationRequestConfig,
|
||||
): Promise<VoxCPMProviderOptions> {
|
||||
if (voiceId === VOXCPM_AUTO_VOICE_ID) {
|
||||
// Drive register-once only when this VoxCPM backend supports it; otherwise
|
||||
// (and on any failure) fall back to the inline voice-design prompt.
|
||||
const canRegister =
|
||||
!!request &&
|
||||
!!context?.voiceDesign &&
|
||||
voxCPMBackendSupportsVoiceRegistration(context.backend ?? 'vllm-omni');
|
||||
const registeredVoiceId = canRegister
|
||||
? await ensureRegisteredVoice(
|
||||
VOXCPM_TTS_PROVIDER_ID,
|
||||
{ voiceDesign: context!.voiceDesign, language: context!.language || context!.locale },
|
||||
request!,
|
||||
).catch(() => undefined)
|
||||
: undefined;
|
||||
return {
|
||||
voiceMode: 'auto',
|
||||
voicePrompt: buildAutoVoxCPMVoicePrompt(context),
|
||||
voicePrompt: buildAutoVoxCPMVoicePrompt(context), // inline fallback always set
|
||||
...(registeredVoiceId ? { registeredVoiceId } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
import type { TTSVoiceInfo } from '@/lib/audio/types';
|
||||
import { buildVoiceDesignPrompt, type VoiceDesign } from '@/lib/audio/voice-design';
|
||||
|
||||
export const VOXCPM_TTS_PROVIDER_ID = 'voxcpm-tts' as const;
|
||||
export const VOXCPM_MODEL_ID = 'VoxCPM2';
|
||||
@@ -38,6 +39,8 @@ export interface VoxCPMVoicePromptContext {
|
||||
persona?: string;
|
||||
language?: string;
|
||||
locale?: string;
|
||||
voiceDesign?: VoiceDesign;
|
||||
backend?: VoxCPMBackendType;
|
||||
}
|
||||
|
||||
export interface VoxCPMProviderOptions {
|
||||
@@ -52,6 +55,7 @@ export interface VoxCPMProviderOptions {
|
||||
inferenceTimesteps?: number;
|
||||
normalize?: boolean;
|
||||
denoise?: boolean;
|
||||
registeredVoiceId?: string;
|
||||
}
|
||||
|
||||
export const VOXCPM_AUTO_VOICE: TTSVoiceInfo = {
|
||||
@@ -102,7 +106,20 @@ function sanitizeAutoVoicePromptPart(value?: string): string {
|
||||
.trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Whether a VoxCPM backend exposes a runtime voice-registration API
|
||||
* (POST /v1/audio/voices) for reference-by-id timbre stability.
|
||||
*/
|
||||
export function voxCPMBackendSupportsVoiceRegistration(backend: VoxCPMBackendType): boolean {
|
||||
return backend === 'vllm-omni';
|
||||
}
|
||||
|
||||
export function buildAutoVoxCPMVoicePrompt(context: VoxCPMVoicePromptContext = {}): string {
|
||||
if (context.voiceDesign) {
|
||||
const designPrompt = sanitizeAutoVoicePromptPart(buildVoiceDesignPrompt(context.voiceDesign));
|
||||
if (designPrompt) return designPrompt;
|
||||
}
|
||||
|
||||
const persona = sanitizeAutoVoicePromptPart(context.persona);
|
||||
if (persona) return persona;
|
||||
|
||||
|
||||
@@ -9,7 +9,8 @@ import {
|
||||
type ResolvedVoice,
|
||||
} from '@/lib/audio/voice-resolver';
|
||||
import { isTTSProviderEnabled } from '@/lib/audio/provider-enablement';
|
||||
import { getVoxCPMProviderOptions, useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { useVoxCPMVoiceProfiles } from '@/lib/audio/voxcpm-voices';
|
||||
import { resolveAgentVoiceOptions } from '@/lib/audio/agent-voice';
|
||||
import type { AgentConfig } from '@/lib/orchestration/registry/types';
|
||||
import type { TTSProviderId } from '@/lib/audio/types';
|
||||
import type { AudioIndicatorState } from '@/components/roundtable/audio-indicator';
|
||||
@@ -180,18 +181,12 @@ export function useDiscussionTTS({ enabled, agents, onAudioStateChange }: Discus
|
||||
try {
|
||||
const providerConfig = ttsProvidersConfig[item.providerId];
|
||||
const agent = item.agentId ? agents.find((a) => a.id === item.agentId) : undefined;
|
||||
const providerOptions =
|
||||
item.providerId === 'voxcpm-tts'
|
||||
? {
|
||||
...(providerConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(item.voiceId, {
|
||||
agentName: agent?.name,
|
||||
role: agent?.role,
|
||||
persona: agent?.persona,
|
||||
locale,
|
||||
})),
|
||||
}
|
||||
: undefined;
|
||||
const providerOptions = await resolveAgentVoiceOptions(agent, {
|
||||
providerId: item.providerId,
|
||||
providerConfig: { ...providerConfig, modelId: item.modelId || providerConfig?.modelId },
|
||||
voiceId: item.voiceId,
|
||||
language: locale,
|
||||
});
|
||||
const res = await fetch('/api/generate/tts', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
|
||||
@@ -12,7 +12,8 @@ import type { Scene } from '@/lib/types/stage';
|
||||
import type { SpeechAction } from '@/lib/types/action';
|
||||
import { splitLongSpeechActions } from '@/lib/audio/tts-utils';
|
||||
import { isTTSProviderEnabled } from '@/lib/audio/provider-enablement';
|
||||
import { getVoxCPMProviderOptions } from '@/lib/audio/voxcpm-voices';
|
||||
import { resolveAgentVoiceOptions, pickNarratorAgent } from '@/lib/audio/agent-voice';
|
||||
import { useAgentRegistry } from '@/lib/orchestration/registry/store';
|
||||
import { generateMediaForOutlines } from '@/lib/media/media-orchestrator';
|
||||
import { createLogger } from '@/lib/logger';
|
||||
|
||||
@@ -147,13 +148,15 @@ export async function generateAndStoreTTS(
|
||||
return;
|
||||
|
||||
const ttsProviderConfig = settings.ttsProvidersConfig?.[settings.ttsProviderId];
|
||||
const providerOptions =
|
||||
settings.ttsProviderId === 'voxcpm-tts'
|
||||
? {
|
||||
...(ttsProviderConfig?.providerOptions || {}),
|
||||
...(await getVoxCPMProviderOptions(settings.ttsVoice, { role: 'teacher', language })),
|
||||
}
|
||||
: undefined;
|
||||
// Narration is the teacher's voice — resolve it from the teacher agent profile
|
||||
// through the single resolver (registers + references by id for stable timbre).
|
||||
const teacher = pickNarratorAgent(useAgentRegistry.getState().listAgents());
|
||||
const providerOptions = await resolveAgentVoiceOptions(teacher, {
|
||||
providerId: settings.ttsProviderId,
|
||||
providerConfig: ttsProviderConfig,
|
||||
voiceId: settings.ttsVoice,
|
||||
language,
|
||||
});
|
||||
const response = await fetch('/api/generate/tts', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
|
||||
@@ -8,6 +8,7 @@ import { persist } from 'zustand/middleware';
|
||||
import type { AgentConfig } from './types';
|
||||
import { getActionsForRole } from './types';
|
||||
import type { TTSProviderId } from '@/lib/audio/types';
|
||||
import type { VoiceDesign } from '@/lib/audio/voice-design';
|
||||
import { USER_AVATAR } from '@/lib/types/roundtable';
|
||||
import type { Participant, ParticipantRole } from '@/lib/types/roundtable';
|
||||
import { useUserProfileStore } from '@/lib/store/user-profile';
|
||||
@@ -383,6 +384,7 @@ export async function saveGeneratedAgents(
|
||||
color: string;
|
||||
priority: number;
|
||||
voiceConfig?: { providerId: string; voiceId: string };
|
||||
voiceDesign?: VoiceDesign;
|
||||
}>,
|
||||
): Promise<string[]> {
|
||||
const { db } = await import('@/lib/utils/database');
|
||||
@@ -422,5 +424,13 @@ export async function saveGeneratedAgents(
|
||||
});
|
||||
}
|
||||
|
||||
// Eager warm-up: pre-register each generated agent's auto voice so the first
|
||||
// spoken line is already stable. Same idempotent ensure as the TTS path;
|
||||
// fire-and-forget. Dynamic import keeps this client-only dep out of the
|
||||
// server-importable store module.
|
||||
void import('@/lib/audio/agent-voice')
|
||||
.then((m) => m.warmUpAgentVoices(registry.listAgents().filter((a) => a.isGenerated)))
|
||||
.catch(() => undefined);
|
||||
|
||||
return records.map((r) => r.id);
|
||||
}
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
*/
|
||||
|
||||
import type { TTSProviderId } from '@/lib/audio/types';
|
||||
import type { VoiceDesign } from '@/lib/audio/voice-design';
|
||||
|
||||
export interface AgentConfig {
|
||||
id: string; // Unique agent ID
|
||||
@@ -15,6 +16,7 @@ export interface AgentConfig {
|
||||
allowedActions: string[]; // Action types this agent can use
|
||||
priority: number; // Priority for director selection (1-10)
|
||||
voiceConfig?: { providerId: TTSProviderId; modelId?: string; voiceId: string }; // Per-agent TTS voice selection
|
||||
voiceDesign?: VoiceDesign; // 3-layer vocal descriptor for auto voice (provider-neutral)
|
||||
|
||||
// Metadata
|
||||
createdAt: Date;
|
||||
@@ -36,6 +38,7 @@ export interface AgentTemplate {
|
||||
allowedActions: string[];
|
||||
priority: number;
|
||||
voiceConfig?: { providerId: TTSProviderId; modelId?: string; voiceId: string }; // Per-agent TTS voice selection
|
||||
voiceDesign?: VoiceDesign; // 3-layer vocal descriptor for auto voice (provider-neutral)
|
||||
|
||||
// LLM-generated agent fields
|
||||
isGenerated?: boolean; // true for LLM-generated agents
|
||||
|
||||
@@ -445,6 +445,18 @@ export function resolveTTSBaseUrl(providerId: string, clientBaseUrl?: string): s
|
||||
return resolveSectionBaseUrl('tts', providerId, clientBaseUrl);
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the TTS model. A managed provider may pin its model server-side
|
||||
* (`${PREFIX}_MODELS`, first entry) — authoritative like its key/baseUrl, since
|
||||
* the managed-provider UI does not expose a model field. Otherwise the client
|
||||
* model wins.
|
||||
*/
|
||||
export function resolveTTSModel(providerId: string, clientModel?: string): string | undefined {
|
||||
const entry = getConfig().tts[providerId];
|
||||
if (entry?.models && entry.models.length > 0) return entry.models[0];
|
||||
return clientModel;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Public API — ASR
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
+32
-1
@@ -9,6 +9,7 @@ import type {
|
||||
ToolCallRequest,
|
||||
} from '@/lib/types/chat';
|
||||
import type { SceneOutline } from '@/lib/types/generation';
|
||||
import type { VoiceDesign } from '@/lib/audio/voice-design';
|
||||
import type { UIMessage } from 'ai';
|
||||
import { createLogger } from '@/lib/logger';
|
||||
|
||||
@@ -166,6 +167,7 @@ export interface GeneratedAgentRecord {
|
||||
avatar: string;
|
||||
color: string;
|
||||
priority: number;
|
||||
voiceDesign?: VoiceDesign; // 3-layer vocal descriptor for auto voice
|
||||
createdAt: number;
|
||||
}
|
||||
|
||||
@@ -186,6 +188,18 @@ export interface VoiceProfileRecord {
|
||||
updatedAt: number;
|
||||
}
|
||||
|
||||
/**
|
||||
* Cached reference clip for a registered auto voice (any TTS provider). The
|
||||
* clip is the source of truth; the deterministic `voiceId` is its key, enabling
|
||||
* register-on-invalid re-registration after backend GC/restart.
|
||||
*/
|
||||
export interface AutoVoiceCacheRecord {
|
||||
voiceId: string;
|
||||
referenceAudio: Blob;
|
||||
mimeType: string;
|
||||
updatedAt: number;
|
||||
}
|
||||
|
||||
/** Build the compound primary key for mediaFiles: `${stageId}:${elementId}` */
|
||||
export function mediaFileKey(stageId: string, elementId: string): string {
|
||||
return `${stageId}:${elementId}`;
|
||||
@@ -194,7 +208,7 @@ export function mediaFileKey(stageId: string, elementId: string): string {
|
||||
// ==================== Database Definition ====================
|
||||
|
||||
const DATABASE_NAME = 'MAIC-Database';
|
||||
const _DATABASE_VERSION = 10;
|
||||
const _DATABASE_VERSION = 11;
|
||||
|
||||
/**
|
||||
* MAIC Database Instance
|
||||
@@ -212,6 +226,7 @@ class MAICDatabase extends Dexie {
|
||||
mediaFiles!: EntityTable<MediaFileRecord, 'id'>;
|
||||
generatedAgents!: EntityTable<GeneratedAgentRecord, 'id'>;
|
||||
voiceProfiles!: EntityTable<VoiceProfileRecord, 'id'>;
|
||||
autoVoiceCache!: EntityTable<AutoVoiceCacheRecord, 'voiceId'>;
|
||||
|
||||
constructor() {
|
||||
super(DATABASE_NAME);
|
||||
@@ -378,6 +393,22 @@ class MAICDatabase extends Dexie {
|
||||
generatedAgents: 'id, stageId',
|
||||
voiceProfiles: 'id, providerId, kind, updatedAt',
|
||||
});
|
||||
|
||||
// Version 11: Add auto-voice reference-clip cache (provider-neutral register-by-id).
|
||||
this.version(11).stores({
|
||||
stages: 'id, updatedAt',
|
||||
scenes: 'id, stageId, order, [stageId+order]',
|
||||
audioFiles: 'id, createdAt',
|
||||
imageFiles: 'id, createdAt',
|
||||
snapshots: '++id',
|
||||
chatSessions: 'id, stageId, [stageId+createdAt]',
|
||||
playbackState: 'stageId',
|
||||
stageOutlines: 'stageId',
|
||||
mediaFiles: 'id, stageId, [stageId+type]',
|
||||
generatedAgents: 'id, stageId',
|
||||
voiceProfiles: 'id, providerId, kind, updatedAt',
|
||||
autoVoiceCache: 'voiceId, updatedAt',
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
|
||||
// agent-voice pulls browser-only deps transitively (IndexedDB, settings store);
|
||||
// stub them so we can unit-test the pure narrator-selection logic in node.
|
||||
vi.mock('@/lib/audio/voxcpm-voices', () => ({ getVoxCPMProviderOptions: vi.fn() }));
|
||||
vi.mock('@/lib/store/settings', () => ({ useSettingsStore: { getState: () => ({}) } }));
|
||||
|
||||
import { pickNarratorAgent, resolveAgentVoiceOptions } from '@/lib/audio/agent-voice';
|
||||
import { getVoxCPMProviderOptions } from '@/lib/audio/voxcpm-voices';
|
||||
import type { AgentConfig } from '@/lib/orchestration/registry/types';
|
||||
|
||||
const mockGetOptions = getVoxCPMProviderOptions as unknown as ReturnType<typeof vi.fn>;
|
||||
|
||||
function agent(partial: Partial<AgentConfig>): AgentConfig {
|
||||
return {
|
||||
id: partial.id ?? 'a',
|
||||
name: partial.name ?? 'A',
|
||||
role: partial.role ?? 'student',
|
||||
persona: '',
|
||||
avatar: '',
|
||||
color: '',
|
||||
allowedActions: [],
|
||||
priority: 5,
|
||||
createdAt: new Date(0),
|
||||
updatedAt: new Date(0),
|
||||
isDefault: false,
|
||||
...partial,
|
||||
};
|
||||
}
|
||||
|
||||
const DESIGN = { identity: 'middle-aged female teacher', texture: 'warm', delivery: 'calm' };
|
||||
|
||||
describe('pickNarratorAgent', () => {
|
||||
it('prefers the teacher WITH a voiceDesign over a default teacher (the registry-seeding bug)', () => {
|
||||
// Registry is always seeded with DEFAULT_AGENTS first, so the default teacher
|
||||
// (no voiceDesign) precedes the generated one. Narration must pick the latter.
|
||||
const agents = [
|
||||
agent({ id: 'default-1', role: 'teacher', isDefault: true }), // no voiceDesign
|
||||
agent({ id: 'gen-1', role: 'teacher', voiceDesign: DESIGN, isGenerated: true }),
|
||||
];
|
||||
expect(pickNarratorAgent(agents)?.id).toBe('gen-1');
|
||||
});
|
||||
|
||||
it('still prefers the voiceDesign teacher regardless of array order', () => {
|
||||
const agents = [
|
||||
agent({ id: 'gen-1', role: 'teacher', voiceDesign: DESIGN }),
|
||||
agent({ id: 'default-1', role: 'teacher', isDefault: true }),
|
||||
];
|
||||
expect(pickNarratorAgent(agents)?.id).toBe('gen-1');
|
||||
});
|
||||
|
||||
it('falls back to any teacher when none has a voiceDesign', () => {
|
||||
const agents = [agent({ id: 'default-1', role: 'teacher', isDefault: true })];
|
||||
expect(pickNarratorAgent(agents)?.id).toBe('default-1');
|
||||
});
|
||||
|
||||
it('returns undefined when there is no teacher', () => {
|
||||
expect(pickNarratorAgent([agent({ role: 'student' })])).toBeUndefined();
|
||||
expect(pickNarratorAgent([])).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe('resolveAgentVoiceOptions — voice-design source', () => {
|
||||
beforeEach(() => {
|
||||
mockGetOptions.mockReset();
|
||||
mockGetOptions.mockResolvedValue({});
|
||||
});
|
||||
const opts = { providerId: 'voxcpm-tts', voiceId: 'voxcpm:auto', providerConfig: {} };
|
||||
|
||||
it('uses the real voiceDesign when the agent has one (generated agents)', async () => {
|
||||
await resolveAgentVoiceOptions(agent({ role: 'teacher', voiceDesign: DESIGN }), opts);
|
||||
expect(mockGetOptions.mock.calls[0][1].voiceDesign).toEqual(DESIGN);
|
||||
});
|
||||
|
||||
it('falls back to the persona as the descriptor when there is no voiceDesign (preset agents)', async () => {
|
||||
await resolveAgentVoiceOptions(agent({ role: 'teacher', persona: 'patient mentor' }), opts);
|
||||
expect(mockGetOptions.mock.calls[0][1].voiceDesign).toEqual({
|
||||
identity: 'patient mentor',
|
||||
texture: '',
|
||||
delivery: '',
|
||||
});
|
||||
});
|
||||
|
||||
it('has no descriptor when the agent has neither voiceDesign nor persona', async () => {
|
||||
await resolveAgentVoiceOptions(agent({ role: 'teacher', persona: '' }), opts);
|
||||
expect(mockGetOptions.mock.calls[0][1].voiceDesign).toBeUndefined();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,89 @@
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import {
|
||||
buildVoiceDesignPrompt,
|
||||
normalizeVoiceDesign,
|
||||
getDeterministicVoiceId,
|
||||
type VoiceDesign,
|
||||
} from '@/lib/audio/voice-design';
|
||||
|
||||
const design: VoiceDesign = {
|
||||
identity: 'middle-aged male teacher',
|
||||
texture: 'warm low-pitched resonant',
|
||||
delivery: 'calm measured encouraging',
|
||||
};
|
||||
|
||||
describe('buildVoiceDesignPrompt', () => {
|
||||
it('composes the three layers into one comma-joined prompt', () => {
|
||||
expect(buildVoiceDesignPrompt(design)).toBe(
|
||||
'middle-aged male teacher, warm low-pitched resonant, calm measured encouraging',
|
||||
);
|
||||
});
|
||||
it('drops blank layers and collapses whitespace', () => {
|
||||
expect(
|
||||
buildVoiceDesignPrompt({ identity: ' male teacher ', texture: '', delivery: 'slow' }),
|
||||
).toBe('male teacher, slow');
|
||||
});
|
||||
it('strips parentheses so they cannot break the (prompt)text delimiter', () => {
|
||||
expect(
|
||||
buildVoiceDesignPrompt({
|
||||
identity: 'male teacher (deep)',
|
||||
texture: '英文(带口音)',
|
||||
delivery: 'calm',
|
||||
}),
|
||||
).toBe('male teacher deep, 英文 带口音, calm');
|
||||
});
|
||||
});
|
||||
|
||||
describe('normalizeVoiceDesign', () => {
|
||||
it('returns a clean design from a well-formed object', () => {
|
||||
expect(normalizeVoiceDesign({ identity: 'a', texture: 'b', delivery: 'c' })).toEqual({
|
||||
identity: 'a',
|
||||
texture: 'b',
|
||||
delivery: 'c',
|
||||
});
|
||||
});
|
||||
it('returns undefined when all layers are empty/missing', () => {
|
||||
expect(normalizeVoiceDesign({})).toBeUndefined();
|
||||
expect(normalizeVoiceDesign(null)).toBeUndefined();
|
||||
expect(normalizeVoiceDesign('nope')).toBeUndefined();
|
||||
});
|
||||
it('keeps a partial design (some layers present)', () => {
|
||||
expect(normalizeVoiceDesign({ identity: 'a' })).toEqual({
|
||||
identity: 'a',
|
||||
texture: '',
|
||||
delivery: '',
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe('getDeterministicVoiceId', () => {
|
||||
it('is stable for the same descriptor+provider+model, with the neutral prefix', async () => {
|
||||
const opts = { providerId: 'voxcpm-tts', model: 'VoxCPM2' };
|
||||
const a = await getDeterministicVoiceId(design, opts);
|
||||
const b = await getDeterministicVoiceId(design, opts);
|
||||
expect(a).toBe(b);
|
||||
expect(a).toMatch(/^auto-[0-9a-f]{16}$/);
|
||||
});
|
||||
it('changes when descriptor, model, or provider changes', async () => {
|
||||
const base = await getDeterministicVoiceId(design, { providerId: 'voxcpm-tts', model: 'm' });
|
||||
const tex = await getDeterministicVoiceId(
|
||||
{ ...design, texture: 'bright' },
|
||||
{ providerId: 'voxcpm-tts', model: 'm' },
|
||||
);
|
||||
const model = await getDeterministicVoiceId(design, { providerId: 'voxcpm-tts', model: 'm2' });
|
||||
const prov = await getDeterministicVoiceId(design, {
|
||||
providerId: 'elevenlabs-tts',
|
||||
model: 'm',
|
||||
});
|
||||
expect(tex).not.toBe(base);
|
||||
expect(model).not.toBe(base);
|
||||
expect(prov).not.toBe(base);
|
||||
});
|
||||
it('is independent of language (descriptor already encodes it)', async () => {
|
||||
// language is not a parameter of the id — same descriptor → same id regardless
|
||||
// of which TTS path (narration directive vs discussion locale) resolves it.
|
||||
const a = await getDeterministicVoiceId(design, { providerId: 'voxcpm-tts', model: 'm' });
|
||||
const b = await getDeterministicVoiceId(design, { providerId: 'voxcpm-tts', model: 'm' });
|
||||
expect(a).toBe(b);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,82 @@
|
||||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
|
||||
// db is browser-only (Dexie); stub it so the client module loads in node.
|
||||
vi.mock('@/lib/utils/database', () => ({
|
||||
db: {
|
||||
autoVoiceCache: {
|
||||
get: vi.fn(async () => undefined),
|
||||
put: vi.fn(async () => undefined),
|
||||
},
|
||||
},
|
||||
}));
|
||||
|
||||
import { ensureRegisteredVoice } from '@/lib/audio/voice-registration-client';
|
||||
|
||||
function okFetch() {
|
||||
const f = vi.fn(
|
||||
async () => new Response(JSON.stringify({ voiceId: 'x', registered: true }), { status: 200 }),
|
||||
);
|
||||
vi.stubGlobal('fetch', f);
|
||||
return f;
|
||||
}
|
||||
|
||||
describe('ensureRegisteredVoice memoization', () => {
|
||||
beforeEach(() => vi.unstubAllGlobals());
|
||||
|
||||
it('re-registers when the backend base URL changes (memo keyed by backend, not voiceId alone)', async () => {
|
||||
const f = okFetch();
|
||||
// Distinct descriptor per test so the module-level memo from other tests can't collide.
|
||||
const voiceDesign = { identity: 'backend-switch teacher', texture: 'warm', delivery: 'calm' };
|
||||
|
||||
await ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, { ttsBaseUrl: 'https://a.test/v1' });
|
||||
// Same backend again → memoized, no second round-trip.
|
||||
await ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, { ttsBaseUrl: 'https://a.test/v1' });
|
||||
expect(f).toHaveBeenCalledTimes(1);
|
||||
|
||||
// Different backend → must NOT be skipped by the memo; re-register there.
|
||||
await ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, { ttsBaseUrl: 'https://b.test/v1' });
|
||||
expect(f).toHaveBeenCalledTimes(2);
|
||||
});
|
||||
|
||||
it('coalesces concurrent calls for the same (voiceId, backend) into one request', async () => {
|
||||
const f = okFetch();
|
||||
const voiceDesign = { identity: 'concurrent teacher', texture: 'warm', delivery: 'calm' };
|
||||
const req = { ttsBaseUrl: 'https://c.test/v1' };
|
||||
|
||||
await Promise.all([
|
||||
ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, req),
|
||||
ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, req),
|
||||
ensureRegisteredVoice('voxcpm-tts', { voiceDesign }, req),
|
||||
]);
|
||||
expect(f).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it('re-registers when the API key changes on the same base URL (auth-scoped)', async () => {
|
||||
const f = okFetch();
|
||||
const voiceDesign = {
|
||||
identity: 'credential-switch teacher',
|
||||
texture: 'warm',
|
||||
delivery: 'calm',
|
||||
};
|
||||
const base = 'https://d.test/v1';
|
||||
|
||||
await ensureRegisteredVoice(
|
||||
'voxcpm-tts',
|
||||
{ voiceDesign },
|
||||
{ ttsBaseUrl: base, ttsApiKey: 'k1' },
|
||||
);
|
||||
await ensureRegisteredVoice(
|
||||
'voxcpm-tts',
|
||||
{ voiceDesign },
|
||||
{ ttsBaseUrl: base, ttsApiKey: 'k1' },
|
||||
);
|
||||
expect(f).toHaveBeenCalledTimes(1); // same creds → memoized
|
||||
|
||||
await ensureRegisteredVoice(
|
||||
'voxcpm-tts',
|
||||
{ voiceDesign },
|
||||
{ ttsBaseUrl: base, ttsApiKey: 'k2' },
|
||||
);
|
||||
expect(f).toHaveBeenCalledTimes(2); // different creds → re-validate
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,23 @@
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import {
|
||||
getVoiceRegistrationAdapter,
|
||||
supportsVoiceRegistration,
|
||||
} from '@/lib/audio/voice-registration';
|
||||
|
||||
describe('voice-registration provider dispatch', () => {
|
||||
it('resolves an adapter for a registration-capable provider', () => {
|
||||
expect(getVoiceRegistrationAdapter('voxcpm-tts')).toBeDefined();
|
||||
});
|
||||
|
||||
it('returns undefined for providers without an adapter', () => {
|
||||
expect(getVoiceRegistrationAdapter('openai-tts')).toBeUndefined();
|
||||
expect(supportsVoiceRegistration('openai-tts')).toBe(false);
|
||||
});
|
||||
|
||||
it('honors per-provider capability (voxcpm only with the vllm-omni backend)', () => {
|
||||
expect(supportsVoiceRegistration('voxcpm-tts', { backend: 'vllm-omni' })).toBe(true);
|
||||
expect(supportsVoiceRegistration('voxcpm-tts', { backend: 'nano-vllm' })).toBe(false);
|
||||
// default backend (vllm-omni) when unspecified
|
||||
expect(supportsVoiceRegistration('voxcpm-tts')).toBe(true);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,58 @@
|
||||
import { describe, it, expect, vi, afterEach } from 'vitest';
|
||||
import { generateTTS } from '@/lib/audio/tts-providers';
|
||||
import { VOXCPM_AUTO_VOICE_ID } from '@/lib/audio/voxcpm';
|
||||
import type { TTSModelConfig } from '@/lib/audio/types';
|
||||
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
|
||||
function stubSpeech() {
|
||||
const f = vi.fn(
|
||||
async () =>
|
||||
new Response(new Uint8Array([82, 73, 70, 70]), {
|
||||
status: 200,
|
||||
headers: { 'content-type': 'audio/wav' },
|
||||
}),
|
||||
);
|
||||
vi.stubGlobal('fetch', f);
|
||||
return f;
|
||||
}
|
||||
|
||||
function lastPayload(f: ReturnType<typeof stubSpeech>) {
|
||||
const [, init] = f.mock.calls[0] as unknown as [string, RequestInit];
|
||||
return JSON.parse(String(init.body));
|
||||
}
|
||||
|
||||
describe('VoxCPM vLLM-Omni registered voice', () => {
|
||||
it('references voice=registeredVoiceId and skips ref_audio / prompt prefix', async () => {
|
||||
const f = stubSpeech();
|
||||
const config: TTSModelConfig = {
|
||||
providerId: 'voxcpm-tts',
|
||||
voice: VOXCPM_AUTO_VOICE_ID,
|
||||
baseUrl: 'https://voxcpm.test/v1',
|
||||
providerOptions: { backend: 'vllm-omni', registeredVoiceId: 'voxcpm:voice:abc' },
|
||||
};
|
||||
|
||||
await generateTTS(config, 'Hello class');
|
||||
|
||||
const payload = lastPayload(f);
|
||||
expect(payload.voice).toBe('voxcpm:voice:abc');
|
||||
expect(payload.input).toBe('Hello class'); // no "(prompt)" prefix
|
||||
expect(payload.ref_audio).toBeUndefined();
|
||||
});
|
||||
|
||||
it('keeps the inline prompt path (voice=default) when no registeredVoiceId', async () => {
|
||||
const f = stubSpeech();
|
||||
const config: TTSModelConfig = {
|
||||
providerId: 'voxcpm-tts',
|
||||
voice: VOXCPM_AUTO_VOICE_ID,
|
||||
baseUrl: 'https://voxcpm.test/v1',
|
||||
providerOptions: { backend: 'vllm-omni', voicePrompt: 'warm male teacher' },
|
||||
};
|
||||
|
||||
await generateTTS(config, 'Hello class');
|
||||
|
||||
const payload = lastPayload(f);
|
||||
expect(payload.voice).toBe('default');
|
||||
expect(payload.input).toBe('(warm male teacher)Hello class');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,97 @@
|
||||
import { describe, it, expect, vi, afterEach } from 'vitest';
|
||||
import {
|
||||
voxCPMVoiceExists,
|
||||
registerVoxCPMVoice,
|
||||
bootstrapVoxCPMReferenceClip,
|
||||
} from '@/lib/audio/voxcpm-registration';
|
||||
|
||||
const cfg = { baseUrl: 'https://voxcpm.test/v1', apiKey: 'k', model: 'voxcpm2' };
|
||||
|
||||
afterEach(() => vi.unstubAllGlobals());
|
||||
|
||||
describe('voxCPMVoiceExists', () => {
|
||||
it('lists voices and checks membership', async () => {
|
||||
vi.stubGlobal(
|
||||
'fetch',
|
||||
vi.fn(
|
||||
async () =>
|
||||
new Response(JSON.stringify({ voices: ['default', 'voxcpm:voice:abc'] }), {
|
||||
status: 200,
|
||||
}),
|
||||
),
|
||||
);
|
||||
expect(await voxCPMVoiceExists(cfg, 'voxcpm:voice:abc')).toBe(true);
|
||||
expect(await voxCPMVoiceExists(cfg, 'voxcpm:voice:missing')).toBe(false);
|
||||
});
|
||||
|
||||
it('returns false when the list call fails', async () => {
|
||||
vi.stubGlobal(
|
||||
'fetch',
|
||||
vi.fn(async () => new Response('', { status: 500 })),
|
||||
);
|
||||
expect(await voxCPMVoiceExists(cfg, 'voxcpm:voice:abc')).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('registerVoxCPMVoice', () => {
|
||||
it('POSTs multipart name + consent + audio_sample with Bearer auth', async () => {
|
||||
const f = vi.fn(
|
||||
async () =>
|
||||
new Response(JSON.stringify({ success: true, voice: { name: 'voxcpm:voice:abc' } }), {
|
||||
status: 200,
|
||||
}),
|
||||
);
|
||||
vi.stubGlobal('fetch', f);
|
||||
|
||||
const id = await registerVoxCPMVoice(cfg, {
|
||||
voiceId: 'voxcpm:voice:abc',
|
||||
referenceAudioBase64: btoa('RIFFdata'),
|
||||
mimeType: 'audio/wav',
|
||||
});
|
||||
|
||||
expect(id).toBe('voxcpm:voice:abc');
|
||||
const [url, init] = f.mock.calls[0] as unknown as [string, RequestInit];
|
||||
expect(String(url)).toBe('https://voxcpm.test/v1/audio/voices');
|
||||
expect(init.method).toBe('POST');
|
||||
const form = init.body as FormData;
|
||||
expect(form).toBeInstanceOf(FormData);
|
||||
expect(form.get('name')).toBe('voxcpm:voice:abc');
|
||||
expect(typeof form.get('consent')).toBe('string');
|
||||
expect(form.get('audio_sample')).toBeInstanceOf(Blob);
|
||||
expect((init.headers as Record<string, string>).Authorization).toBe('Bearer k');
|
||||
});
|
||||
|
||||
it('throws on non-ok response', async () => {
|
||||
vi.stubGlobal(
|
||||
'fetch',
|
||||
vi.fn(async () => new Response('nope', { status: 400 })),
|
||||
);
|
||||
await expect(
|
||||
registerVoxCPMVoice(cfg, { voiceId: 'v', referenceAudioBase64: btoa('x') }),
|
||||
).rejects.toThrow();
|
||||
});
|
||||
});
|
||||
|
||||
describe('bootstrapVoxCPMReferenceClip', () => {
|
||||
it('synthesizes the descriptor prompt into base64 wav', async () => {
|
||||
const wav = new Uint8Array([82, 73, 70, 70]); // "RIFF"
|
||||
const f = vi.fn(
|
||||
async () => new Response(wav, { status: 200, headers: { 'content-type': 'audio/wav' } }),
|
||||
);
|
||||
vi.stubGlobal('fetch', f);
|
||||
|
||||
const out = await bootstrapVoxCPMReferenceClip(cfg, {
|
||||
design: { identity: 'male teacher', texture: 'warm', delivery: 'calm' },
|
||||
language: 'en',
|
||||
});
|
||||
|
||||
expect(out.mimeType).toContain('wav');
|
||||
expect(typeof out.referenceAudioBase64).toBe('string');
|
||||
expect(out.referenceAudioBase64.length).toBeGreaterThan(0);
|
||||
|
||||
const [url, init] = f.mock.calls[0] as unknown as [string, RequestInit];
|
||||
expect(String(url)).toBe('https://voxcpm.test/v1/audio/speech');
|
||||
const payload = JSON.parse(String(init.body));
|
||||
expect(payload.input).toContain('(male teacher, warm, calm)');
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,33 @@
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import {
|
||||
buildAutoVoxCPMVoicePrompt,
|
||||
voxCPMBackendSupportsVoiceRegistration,
|
||||
} from '@/lib/audio/voxcpm';
|
||||
import { buildVoiceDesignPrompt, type VoiceDesign } from '@/lib/audio/voice-design';
|
||||
|
||||
const design: VoiceDesign = {
|
||||
identity: 'middle-aged male teacher',
|
||||
texture: 'warm low-pitched resonant',
|
||||
delivery: 'calm measured encouraging',
|
||||
};
|
||||
|
||||
describe('buildAutoVoxCPMVoicePrompt fallback chain', () => {
|
||||
it('prefers voiceDesign over persona', () => {
|
||||
expect(buildAutoVoxCPMVoicePrompt({ voiceDesign: design, persona: 'loves cats' })).toBe(
|
||||
buildVoiceDesignPrompt(design),
|
||||
);
|
||||
});
|
||||
it('falls back to persona, then role/name, then default', () => {
|
||||
expect(buildAutoVoxCPMVoicePrompt({ persona: 'patient mentor' })).toBe('patient mentor');
|
||||
expect(buildAutoVoxCPMVoicePrompt({ role: 'teacher', agentName: 'Lin' })).toBe('teacher Lin');
|
||||
expect(buildAutoVoxCPMVoicePrompt({})).toBe('natural classroom voice');
|
||||
});
|
||||
});
|
||||
|
||||
describe('voxCPMBackendSupportsVoiceRegistration', () => {
|
||||
it('is true only for vllm-omni', () => {
|
||||
expect(voxCPMBackendSupportsVoiceRegistration('vllm-omni')).toBe(true);
|
||||
expect(voxCPMBackendSupportsVoiceRegistration('nano-vllm')).toBe(false);
|
||||
expect(voxCPMBackendSupportsVoiceRegistration('python-api')).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,86 @@
|
||||
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
||||
import { NextRequest } from 'next/server';
|
||||
|
||||
const callLLM = vi.fn();
|
||||
|
||||
vi.mock('@/lib/ai/llm', () => ({
|
||||
callLLM: (...args: unknown[]) => callLLM(...args),
|
||||
}));
|
||||
|
||||
vi.mock('@/lib/server/resolve-model', () => ({
|
||||
resolveModelFromRequest: async () => ({
|
||||
model: {},
|
||||
modelString: 'test-model',
|
||||
thinkingConfig: undefined,
|
||||
}),
|
||||
}));
|
||||
|
||||
import { POST } from '@/app/api/generate/agent-profiles/route';
|
||||
|
||||
function makeRequest(): NextRequest {
|
||||
return new NextRequest('http://localhost/api/generate/agent-profiles', {
|
||||
method: 'POST',
|
||||
headers: { 'content-type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
stageInfo: { name: 'Intro to Algebra' },
|
||||
languageDirective: 'Respond in English.',
|
||||
availableAvatars: ['/a.png', '/b.png'],
|
||||
}),
|
||||
});
|
||||
}
|
||||
|
||||
function llmAgents(extra: Record<string, unknown>) {
|
||||
return JSON.stringify({
|
||||
agents: [
|
||||
{
|
||||
name: 'Prof. Lin',
|
||||
role: 'teacher',
|
||||
persona: 'A patient mentor.',
|
||||
avatar: '/a.png',
|
||||
color: '#111111',
|
||||
priority: 10,
|
||||
...extra,
|
||||
},
|
||||
{
|
||||
name: 'Sam',
|
||||
role: 'student',
|
||||
persona: 'Curious learner.',
|
||||
avatar: '/b.png',
|
||||
color: '#222222',
|
||||
priority: 5,
|
||||
},
|
||||
],
|
||||
});
|
||||
}
|
||||
|
||||
describe('agent-profiles route — voiceDesign', () => {
|
||||
beforeEach(() => callLLM.mockReset());
|
||||
|
||||
it('attaches a normalized voiceDesign when the LLM emits one', async () => {
|
||||
callLLM.mockResolvedValue({
|
||||
text: llmAgents({
|
||||
voiceDesign: { identity: 'older male teacher', texture: 'warm low', delivery: 'calm' },
|
||||
}),
|
||||
});
|
||||
|
||||
const res = await POST(makeRequest());
|
||||
const body = await res.json();
|
||||
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.agents[0].voiceDesign).toEqual({
|
||||
identity: 'older male teacher',
|
||||
texture: 'warm low',
|
||||
delivery: 'calm',
|
||||
});
|
||||
});
|
||||
|
||||
it('omits voiceDesign when the LLM does not emit one', async () => {
|
||||
callLLM.mockResolvedValue({ text: llmAgents({}) });
|
||||
|
||||
const res = await POST(makeRequest());
|
||||
const body = await res.json();
|
||||
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.agents[0]).not.toHaveProperty('voiceDesign');
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user