feat(generate): server-side media & TTS generation (#75)

* feat(generate): server-side media & TTS generation in classroom pipeline

When `enableImageGeneration`, `enableVideoGeneration`, or `enableTTS` flags
are passed to /api/generate-classroom, the server now actually generates
media files and TTS audio, persists them to disk, and replaces placeholders
in the classroom JSON with serving URLs.

- Add media serving API route (GET /api/classroom-media/[id]/[...path])
- Add server-side media generation utility (image, video, TTS)
- Wire generation phases into classroom pipeline with progress steps
- Update AudioPlayer to support URL-based playback (server-generated TTS)
- Add audioUrl field to SpeechAction type
- Extend /api/health with capabilities detection
- Pass through feature flags in generate-classroom route
- Update skill docs for optional feature configuration

Closes #49

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix prettier formatting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(generate): add agentMode support for server-side classroom generation

Allow callers to choose between built-in default agents ('default') or
LLM-generated custom agent profiles ('generate') via a new agentMode
field on GenerateClassroomInput. The resolved agents are injected into
outline, scene content, and scene action prompts via teacherContext.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(media): harden download, parallelize media gen, fix TTS concatenation

- Add timeout (2min) and max size (100MB) guard to downloadToBuffer
- Run image and video generation in parallel via Promise.all
- Only split TTS text for MP3 (frame-based); skip for WAV/OGG/AAC
  whose container headers break on naive byte concatenation
- Remove unnecessary type assertion for speechAction.audioUrl

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix prettier formatting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(media): harden media serving and TTS generation

- Add realpath check to prevent symlink-based path traversal in media route
- Stream files via createReadStream instead of readFile to avoid OOM on large videos
- Truncate TTS text for non-concatenable formats (WAV/OGG/AAC) when exceeding provider limit
- Add 'researching' step to distinguish web search from outline generation progress

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tts): split long speech actions instead of concatenating audio bytes

Adopt the same strategy as client-side: split a long speech action into
multiple shorter actions before TTS generation, so each sub-action gets
its own audio file. This avoids the broken byte-concatenation approach
that corrupted non-MP3 formats (WAV/OGG/AAC container headers).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(tts): extract shared TTS splitting logic into lib/audio/tts-utils

Both client-side (use-scene-generator) and server-side (classroom-media-generation)
had duplicated TTS_MAX_TEXT_LENGTH, text splitting, and action splitting logic.
Extract into a shared module to keep them in sync.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: remove unused TTS_MAX_TEXT_LENGTH import

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
This commit is contained in:
杨慎
2026-03-19 14:40:04 +08:00
committed by GitHub
co-authored by Claude Opus 4.6 wyuc
parent 43a0c8e39e
commit fbe8399678
17 changed files with 769 additions and 96 deletions
@@ -0,0 +1,88 @@
import { promises as fs, createReadStream } from 'fs';
import path from 'path';
import { NextRequest, NextResponse } from 'next/server';
import { CLASSROOMS_DIR, isValidClassroomId } from '@/lib/server/classroom-storage';
const MIME_TYPES: Record<string, string> = {
'.png': 'image/png',
'.jpg': 'image/jpeg',
'.jpeg': 'image/jpeg',
'.webp': 'image/webp',
'.gif': 'image/gif',
'.mp4': 'video/mp4',
'.webm': 'video/webm',
'.mp3': 'audio/mpeg',
'.wav': 'audio/wav',
'.ogg': 'audio/ogg',
'.aac': 'audio/aac',
};
export async function GET(
_req: NextRequest,
{ params }: { params: Promise<{ classroomId: string; path: string[] }> },
) {
const { classroomId, path: pathSegments } = await params;
// Validate classroomId
if (!isValidClassroomId(classroomId)) {
return NextResponse.json({ error: 'Invalid classroom ID' }, { status: 400 });
}
// Validate path segments — no traversal
const joined = pathSegments.join('/');
if (joined.includes('..') || pathSegments.some((s) => s.includes('\0'))) {
return NextResponse.json({ error: 'Invalid path' }, { status: 400 });
}
// Only allow media/ and audio/ subdirectories
const subDir = pathSegments[0];
if (subDir !== 'media' && subDir !== 'audio') {
return NextResponse.json({ error: 'Invalid path' }, { status: 404 });
}
const filePath = path.join(CLASSROOMS_DIR, classroomId, ...pathSegments);
const resolvedBase = path.resolve(CLASSROOMS_DIR, classroomId);
try {
// Resolve symlinks and verify the real path stays within the classroom dir
const realPath = await fs.realpath(filePath);
if (!realPath.startsWith(resolvedBase + path.sep) && realPath !== resolvedBase) {
return NextResponse.json({ error: 'Not found' }, { status: 404 });
}
const stat = await fs.stat(realPath);
if (!stat.isFile()) {
return NextResponse.json({ error: 'Not found' }, { status: 404 });
}
const ext = path.extname(realPath).toLowerCase();
const contentType = MIME_TYPES[ext] || 'application/octet-stream';
// Stream the file to avoid loading large videos into memory
const stream = createReadStream(realPath);
const webStream = new ReadableStream({
start(controller) {
stream.on('data', (chunk: Buffer | string) => controller.enqueue(chunk));
stream.on('end', () => controller.close());
stream.on('error', (err) => controller.error(err));
},
cancel() {
stream.destroy();
},
});
return new NextResponse(webStream, {
status: 200,
headers: {
'Content-Type': contentType,
'Content-Length': String(stat.size),
'Cache-Control': 'public, max-age=86400, immutable',
},
});
} catch (error) {
if ((error as NodeJS.ErrnoException).code === 'ENOENT') {
return NextResponse.json({ error: 'Not found' }, { status: 404 });
}
return NextResponse.json({ error: 'Internal error' }, { status: 500 });
}
}
+9
View File
@@ -15,6 +15,15 @@ export async function POST(req: NextRequest) {
requirement: rawBody.requirement || '',
...(rawBody.pdfContent ? { pdfContent: rawBody.pdfContent } : {}),
...(rawBody.language ? { language: rawBody.language } : {}),
...(rawBody.enableWebSearch != null ? { enableWebSearch: rawBody.enableWebSearch } : {}),
...(rawBody.enableImageGeneration != null
? { enableImageGeneration: rawBody.enableImageGeneration }
: {}),
...(rawBody.enableVideoGeneration != null
? { enableVideoGeneration: rawBody.enableVideoGeneration }
: {}),
...(rawBody.enableTTS != null ? { enableTTS: rawBody.enableTTS } : {}),
...(rawBody.agentMode ? { agentMode: rawBody.agentMode } : {}),
};
const { requirement } = body;
+16 -1
View File
@@ -1,7 +1,22 @@
import { apiSuccess } from '@/lib/server/api-response';
import {
getServerWebSearchProviders,
getServerImageProviders,
getServerVideoProviders,
getServerTTSProviders,
} from '@/lib/server/provider-config';
const version = process.env.npm_package_version || '0.1.0';
export async function GET() {
return apiSuccess({ status: 'ok', version });
return apiSuccess({
status: 'ok',
version,
capabilities: {
webSearch: Object.keys(getServerWebSearchProviders()).length > 0,
imageGeneration: Object.keys(getServerImageProviders()).length > 0,
videoGeneration: Object.keys(getServerVideoProviders()).length > 0,
tts: Object.keys(getServerTTSProviders()).length > 0,
},
});
}
+1 -1
View File
@@ -169,7 +169,7 @@ export class ActionEngine {
return new Promise<void>((resolve) => {
this.audioPlayer!.onEnded(() => resolve());
this.audioPlayer!.play(action.audioId || '')
this.audioPlayer!.play(action.audioId || '', action.audioUrl)
.then((audioStarted) => {
if (!audioStarted) resolve();
})
+106
View File
@@ -0,0 +1,106 @@
/**
* Shared TTS utilities used by both client-side and server-side generation.
*/
import type { TTSProviderId } from './types';
import type { Action, SpeechAction } from '@/lib/types/action';
import { createLogger } from '@/lib/logger';
const log = createLogger('TTS');
/** Provider-specific max text length limits. */
export const TTS_MAX_TEXT_LENGTH: Partial<Record<TTSProviderId, number>> = {
'glm-tts': 1024,
};
/**
* Split long text into chunks that respect sentence boundaries.
* Tries splitting at sentence-ending punctuation first, then clause-level
* punctuation, and finally hard-splits at maxLength as a last resort.
*/
export function splitLongSpeechText(text: string, maxLength: number): string[] {
const normalized = text.trim();
if (!normalized || normalized.length <= maxLength) return [normalized];
const units = normalized
.split(/(?<=[。!?!?;;::\n])/u)
.map((part) => part.trim())
.filter(Boolean);
const chunks: string[] = [];
let current = '';
const pushChunk = (value: string) => {
const trimmed = value.trim();
if (trimmed) chunks.push(trimmed);
};
const appendUnit = (unit: string) => {
if (!current) {
current = unit;
return;
}
if ((current + unit).length <= maxLength) {
current += unit;
return;
}
pushChunk(current);
current = unit;
};
const hardSplitUnit = (unit: string) => {
const parts = unit.split(/(?<=[,,、])/u).filter(Boolean);
if (parts.length > 1) {
for (const part of parts) {
if (part.length <= maxLength) appendUnit(part);
else hardSplitUnit(part);
}
return;
}
let start = 0;
while (start < unit.length) {
appendUnit(unit.slice(start, start + maxLength));
start += maxLength;
}
};
for (const unit of units.length > 0 ? units : [normalized]) {
if (unit.length <= maxLength) appendUnit(unit);
else hardSplitUnit(unit);
}
pushChunk(current);
return chunks;
}
/**
* Split long speech actions into multiple shorter actions so each stays
* within the TTS provider's text length limit. Each sub-action gets its
* own independent audio file — no byte concatenation needed.
*/
export function splitLongSpeechActions(actions: Action[], providerId: TTSProviderId): Action[] {
const maxLength = TTS_MAX_TEXT_LENGTH[providerId];
if (!maxLength) return actions;
let didSplit = false;
const nextActions: Action[] = actions.flatMap((action) => {
if (action.type !== 'speech' || !action.text || action.text.length <= maxLength)
return [action];
const chunks = splitLongSpeechText(action.text, maxLength);
if (chunks.length <= 1) return [action];
didSplit = true;
const { audioId: _audioId, ...baseAction } = action as SpeechAction;
log.info(
`Split speech for ${providerId}: action=${action.id}, len=${action.text.length}, chunks=${chunks.length}`,
);
return chunks.map((chunk, i) => ({
...baseAction,
id: `${action.id}_tts_${i + 1}`,
text: chunk,
}));
});
return didSplit ? nextActions : actions;
}
+6
View File
@@ -34,6 +34,8 @@ export async function generateSceneOutlinesFromRequirements(
imageMapping?: ImageMapping;
imageGenerationEnabled?: boolean;
videoGenerationEnabled?: boolean;
researchContext?: string;
teacherContext?: string;
},
): Promise<GenerationResult<SceneOutline[]>> {
// Build available images description for the prompt
@@ -105,6 +107,10 @@ export async function generateSceneOutlinesFromRequirements(
availableImages: availableImagesText,
userProfile: userProfileText,
mediaGenerationPolicy,
researchContext:
options?.researchContext || (requirements.language === 'zh-CN' ? '无' : 'None'),
// Server-side generation populates this via options; client-side populates via formatTeacherPersonaForPrompt
teacherContext: options?.teacherContext || '',
});
if (!prompts) {
+1 -85
View File
@@ -10,13 +10,11 @@ import type { AgentInfo } from '@/lib/generation/generation-pipeline';
import type { Scene } from '@/lib/types/stage';
import type { Action, SpeechAction } from '@/lib/types/action';
import type { TTSProviderId } from '@/lib/audio/types';
import { splitLongSpeechActions } from '@/lib/audio/tts-utils';
import { generateMediaForOutlines } from '@/lib/media/media-orchestrator';
import { createLogger } from '@/lib/logger';
const log = createLogger('SceneGenerator');
const TTS_MAX_TEXT_LENGTH: Partial<Record<TTSProviderId, number>> = {
'glm-tts': 1024,
};
interface SceneContentResult {
success: boolean;
@@ -32,88 +30,6 @@ interface SceneActionsResult {
error?: string;
}
export function splitLongSpeechText(text: string, maxLength: number): string[] {
const normalized = text.trim();
if (!normalized || normalized.length <= maxLength) return [normalized];
const units = normalized
.split(/(?<=[。!?!?;;::\n])/u)
.map((part) => part.trim())
.filter(Boolean);
const chunks: string[] = [];
let current = '';
const pushChunk = (value: string) => {
const trimmed = value.trim();
if (trimmed) chunks.push(trimmed);
};
const appendUnit = (unit: string) => {
if (!current) {
current = unit;
return;
}
if ((current + unit).length <= maxLength) {
current += unit;
return;
}
pushChunk(current);
current = unit;
};
const hardSplitUnit = (unit: string) => {
const parts = unit.split(/(?<=[,,、])/u).filter(Boolean);
if (parts.length > 1) {
for (const part of parts) {
if (part.length <= maxLength) appendUnit(part);
else hardSplitUnit(part);
}
return;
}
let start = 0;
while (start < unit.length) {
appendUnit(unit.slice(start, start + maxLength));
start += maxLength;
}
};
for (const unit of units.length > 0 ? units : [normalized]) {
if (unit.length <= maxLength) appendUnit(unit);
else hardSplitUnit(unit);
}
pushChunk(current);
return chunks;
}
function splitLongSpeechActions(actions: Action[], providerId: TTSProviderId): Action[] {
const maxLength = TTS_MAX_TEXT_LENGTH[providerId];
if (!maxLength) return actions;
let didSplit = false;
const nextActions: Action[] = actions.flatMap((action) => {
if (action.type !== 'speech' || !action.text || action.text.length <= maxLength)
return [action];
const chunks = splitLongSpeechText(action.text, maxLength);
if (chunks.length <= 1) return [action];
didSplit = true;
const { audioId: _audioId, ...baseAction } = action;
log.info(
`Split speech for ${providerId}: action=${action.id}, len=${action.text.length}, chunks=${chunks.length}`,
);
return chunks.map((chunk, i) => ({
...baseAction,
id: `${action.id}_tts_${i + 1}`,
text: chunk,
}));
});
return didSplit ? nextActions : actions;
}
function getApiHeaders(): HeadersInit {
const config = getCurrentModelConfig();
const settings = useSettingsStore.getState();
+14
View File
@@ -10,6 +10,7 @@ import { getActionsForRole } from './types';
import { USER_AVATAR } from '@/lib/types/roundtable';
import type { Participant, ParticipantRole } from '@/lib/types/roundtable';
import { useUserProfileStore } from '@/lib/store/user-profile';
import type { AgentInfo } from '@/lib/generation/pipeline-types';
interface AgentRegistryState {
agents: Record<string, AgentConfig>; // Map of agentId -> config
@@ -186,6 +187,19 @@ Tone: Thoughtful, measured, intellectually curious. You pause before speaking. Y
},
};
/**
* Return the built-in default agents as lightweight AgentInfo objects
* suitable for the generation pipeline (no UI-only fields like avatar/color).
*/
export function getDefaultAgents(): AgentInfo[] {
return Object.values(DEFAULT_AGENTS).map((a) => ({
id: a.id,
name: a.name,
role: a.role,
persona: a.persona,
}));
}
export const useAgentRegistry = create<AgentRegistryState>()(
persist(
(set, get) => ({
+1 -1
View File
@@ -465,7 +465,7 @@ export class PlaybackEngine {
};
this.audioPlayer
.play(speechAction.audioId || '')
.play(speechAction.audioId || '', speechAction.audioUrl)
.then((audioStarted) => {
if (!audioStarted) {
// No pre-generated audio — try browser-native TTS if selected
+177 -4
View File
@@ -12,11 +12,20 @@ import {
generateSceneContent,
} from '@/lib/generation/scene-generator';
import type { AICallFn } from '@/lib/generation/pipeline-types';
import type { AgentInfo } from '@/lib/generation/pipeline-types';
import { formatTeacherPersonaForPrompt } from '@/lib/generation/prompt-formatters';
import { getDefaultAgents } from '@/lib/orchestration/registry/store';
import { createLogger } from '@/lib/logger';
import { parseModelString } from '@/lib/ai/providers';
import { resolveApiKey } from '@/lib/server/provider-config';
import { resolveApiKey, resolveWebSearchApiKey } from '@/lib/server/provider-config';
import { resolveModel } from '@/lib/server/resolve-model';
import { searchWithTavily, formatSearchResultsAsContext } from '@/lib/web-search/tavily';
import { persistClassroom } from '@/lib/server/classroom-storage';
import {
generateMediaForClassroom,
replaceMediaPlaceholders,
generateTTSForClassroom,
} from '@/lib/server/classroom-media-generation';
import type { UserRequirements } from '@/lib/types/generation';
import type { Scene, Stage } from '@/lib/types/stage';
@@ -26,12 +35,20 @@ export interface GenerateClassroomInput {
requirement: string;
pdfContent?: { text: string; images: string[] };
language?: string;
enableWebSearch?: boolean;
enableImageGeneration?: boolean;
enableVideoGeneration?: boolean;
enableTTS?: boolean;
agentMode?: 'default' | 'generate';
}
export type ClassroomGenerationStep =
| 'initializing'
| 'researching'
| 'generating_outlines'
| 'generating_scenes'
| 'generating_media'
| 'generating_tts'
| 'persisting'
| 'completed';
@@ -83,6 +100,65 @@ function normalizeLanguage(language?: string): 'zh-CN' | 'en-US' {
return language === 'en-US' ? 'en-US' : 'zh-CN';
}
function stripCodeFences(text: string): string {
let cleaned = text.trim();
if (cleaned.startsWith('```')) {
cleaned = cleaned.replace(/^```(?:json)?\s*\n?/, '').replace(/\n?```\s*$/, '');
}
return cleaned.trim();
}
async function generateAgentProfiles(
requirement: string,
language: string,
aiCall: AICallFn,
): Promise<AgentInfo[]> {
const systemPrompt =
'You are an expert instructional designer. Generate agent profiles for a multi-agent classroom simulation. Return ONLY valid JSON, no markdown or explanation.';
const userPrompt = `Generate agent profiles for a course with this requirement:
${requirement}
Requirements:
- Decide the appropriate number of agents based on the course content (typically 3-5)
- Exactly 1 agent must have role "teacher", the rest can be "assistant" or "student"
- Each agent needs: name, role, persona (2-3 sentences describing personality and teaching/learning style)
- Names and personas must be in language: ${language}
Return a JSON object with this exact structure:
{
"agents": [
{
"name": "string",
"role": "teacher" | "assistant" | "student",
"persona": "string (2-3 sentences)"
}
]
}`;
const response = await aiCall(systemPrompt, userPrompt);
const rawText = stripCodeFences(response);
const parsed = JSON.parse(rawText) as {
agents: Array<{ name: string; role: string; persona: string }>;
};
if (!parsed.agents || !Array.isArray(parsed.agents) || parsed.agents.length < 2) {
throw new Error(`Expected at least 2 agents, got ${parsed.agents?.length ?? 0}`);
}
const teacherCount = parsed.agents.filter((a) => a.role === 'teacher').length;
if (teacherCount !== 1) {
throw new Error(`Expected exactly 1 teacher, got ${teacherCount}`);
}
return parsed.agents.map((a, i) => ({
id: `gen-server-${i}`,
name: a.name,
role: a.role,
persona: a.persona,
}));
}
export async function generateClassroom(
input: GenerateClassroomInput,
options: {
@@ -134,6 +210,50 @@ export async function generateClassroom(
};
const pdfText = pdfContent?.text || undefined;
// Resolve agents based on agentMode
let agents: AgentInfo[];
const agentMode = input.agentMode || 'default';
if (agentMode === 'generate') {
log.info('Generating custom agent profiles via LLM...');
try {
agents = await generateAgentProfiles(requirement, lang, aiCall);
log.info(`Generated ${agents.length} agent profiles`);
} catch (e) {
log.warn('Agent profile generation failed, falling back to defaults:', e);
agents = getDefaultAgents();
}
} else {
agents = getDefaultAgents();
}
const teacherContext = formatTeacherPersonaForPrompt(agents);
await options.onProgress?.({
step: 'researching',
progress: 10,
message: 'Researching topic',
scenesGenerated: 0,
});
// Web search (optional, graceful degradation)
let researchContext: string | undefined;
if (input.enableWebSearch) {
const tavilyKey = resolveWebSearchApiKey();
if (tavilyKey) {
try {
log.info('Running web search for requirement context...');
const searchResult = await searchWithTavily({ query: requirement, apiKey: tavilyKey });
researchContext = formatSearchResultsAsContext(searchResult);
if (researchContext) {
log.info(`Web search returned ${searchResult.sources.length} sources`);
}
} catch (e) {
log.warn('Web search failed, continuing without search context:', e);
}
} else {
log.warn('enableWebSearch is true but no Tavily API key configured, skipping web search');
}
}
await options.onProgress?.({
step: 'generating_outlines',
progress: 15,
@@ -146,6 +266,13 @@ export async function generateClassroom(
pdfText,
undefined,
aiCall,
undefined,
{
imageGenerationEnabled: input.enableImageGeneration,
videoGenerationEnabled: input.enableVideoGeneration,
researchContext,
teacherContext,
},
);
if (!outlinesResult.success || !outlinesResult.data) {
@@ -193,13 +320,22 @@ export async function generateClassroom(
totalScenes: outlines.length,
});
const content = await generateSceneContent(safeOutline, aiCall);
const content = await generateSceneContent(
safeOutline,
aiCall,
undefined,
undefined,
undefined,
undefined,
undefined,
agents,
);
if (!content) {
log.warn(`Skipping scene "${safeOutline.title}" — content generation failed`);
continue;
}
const actions = await generateSceneActions(safeOutline, content, aiCall);
const actions = await generateSceneActions(safeOutline, content, aiCall, undefined, agents);
log.info(`Scene "${safeOutline.title}": ${actions.length} actions`);
const sceneId = createSceneWithActions(safeOutline, content, actions, api);
@@ -226,9 +362,46 @@ export async function generateClassroom(
throw new Error('No scenes were generated');
}
// Phase: Media generation (after all scenes generated)
if (input.enableImageGeneration || input.enableVideoGeneration) {
await options.onProgress?.({
step: 'generating_media',
progress: 90,
message: 'Generating media files',
scenesGenerated: scenes.length,
totalScenes: outlines.length,
});
try {
const mediaMap = await generateMediaForClassroom(outlines, stageId, options.baseUrl);
replaceMediaPlaceholders(scenes, mediaMap);
log.info(`Media generation complete: ${Object.keys(mediaMap).length} files`);
} catch (err) {
log.warn('Media generation phase failed, continuing:', err);
}
}
// Phase: TTS generation
if (input.enableTTS) {
await options.onProgress?.({
step: 'generating_tts',
progress: 94,
message: 'Generating TTS audio',
scenesGenerated: scenes.length,
totalScenes: outlines.length,
});
try {
await generateTTSForClassroom(scenes, stageId, options.baseUrl);
log.info('TTS generation complete');
} catch (err) {
log.warn('TTS generation phase failed, continuing:', err);
}
}
await options.onProgress?.({
step: 'persisting',
progress: 95,
progress: 98,
message: 'Persisting classroom data',
scenesGenerated: scenes.length,
totalScenes: outlines.length,
+260
View File
@@ -0,0 +1,260 @@
/**
* Server-side media and TTS generation for classrooms.
*
* Generates image/video files and TTS audio for a classroom,
* writes them to disk, and returns serving URL mappings.
*/
import { promises as fs } from 'fs';
import path from 'path';
import { createLogger } from '@/lib/logger';
import { CLASSROOMS_DIR } from '@/lib/server/classroom-storage';
import { generateImage } from '@/lib/media/image-providers';
import { generateVideo, normalizeVideoOptions } from '@/lib/media/video-providers';
import { generateTTS } from '@/lib/audio/tts-providers';
import { DEFAULT_TTS_VOICES, TTS_PROVIDERS } from '@/lib/audio/constants';
import { IMAGE_PROVIDERS } from '@/lib/media/image-providers';
import { VIDEO_PROVIDERS } from '@/lib/media/video-providers';
import { isMediaPlaceholder } from '@/lib/store/media-generation';
import {
getServerImageProviders,
getServerVideoProviders,
getServerTTSProviders,
resolveImageApiKey,
resolveImageBaseUrl,
resolveVideoApiKey,
resolveVideoBaseUrl,
resolveTTSApiKey,
resolveTTSBaseUrl,
} from '@/lib/server/provider-config';
import type { SceneOutline } from '@/lib/types/generation';
import type { Scene } from '@/lib/types/stage';
import type { SpeechAction } from '@/lib/types/action';
import type { ImageProviderId } from '@/lib/media/types';
import type { VideoProviderId } from '@/lib/media/types';
import type { TTSProviderId } from '@/lib/audio/types';
import { splitLongSpeechActions } from '@/lib/audio/tts-utils';
const log = createLogger('ClassroomMedia');
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
async function ensureDir(dir: string) {
await fs.mkdir(dir, { recursive: true });
}
const DOWNLOAD_TIMEOUT_MS = 120_000; // 2 minutes
const DOWNLOAD_MAX_SIZE = 100 * 1024 * 1024; // 100 MB
async function downloadToBuffer(url: string): Promise<Buffer> {
const resp = await fetch(url, { signal: AbortSignal.timeout(DOWNLOAD_TIMEOUT_MS) });
if (!resp.ok) throw new Error(`Download failed: ${resp.status} ${resp.statusText}`);
const contentLength = Number(resp.headers.get('content-length') || 0);
if (contentLength > DOWNLOAD_MAX_SIZE) {
throw new Error(`File too large: ${contentLength} bytes (max ${DOWNLOAD_MAX_SIZE})`);
}
return Buffer.from(await resp.arrayBuffer());
}
function mediaServingUrl(baseUrl: string, classroomId: string, subPath: string): string {
return `${baseUrl}/api/classroom-media/${classroomId}/${subPath}`;
}
// ---------------------------------------------------------------------------
// Image / Video generation
// ---------------------------------------------------------------------------
export async function generateMediaForClassroom(
outlines: SceneOutline[],
classroomId: string,
baseUrl: string,
): Promise<Record<string, string>> {
const mediaDir = path.join(CLASSROOMS_DIR, classroomId, 'media');
await ensureDir(mediaDir);
// Collect all media generation requests from outlines
const requests = outlines.flatMap((o) => o.mediaGenerations ?? []);
if (requests.length === 0) return {};
// Resolve providers
const imageProviderIds = Object.keys(getServerImageProviders());
const videoProviderIds = Object.keys(getServerVideoProviders());
const mediaMap: Record<string, string> = {};
// Separate image and video requests, generate each type sequentially
// but run the two types in parallel (providers often have limited concurrency).
const imageRequests = requests.filter((r) => r.type === 'image' && imageProviderIds.length > 0);
const videoRequests = requests.filter((r) => r.type === 'video' && videoProviderIds.length > 0);
const generateImages = async () => {
for (const req of imageRequests) {
try {
const providerId = imageProviderIds[0] as ImageProviderId;
const apiKey = resolveImageApiKey(providerId);
if (!apiKey) {
log.warn(`No API key for image provider "${providerId}", skipping ${req.elementId}`);
continue;
}
const providerConfig = IMAGE_PROVIDERS[providerId];
const model = providerConfig?.models?.[0]?.id;
const result = await generateImage(
{ providerId, apiKey, baseUrl: resolveImageBaseUrl(providerId), model },
{ prompt: req.prompt, aspectRatio: req.aspectRatio || '16:9' },
);
let buf: Buffer;
let ext: string;
if (result.base64) {
buf = Buffer.from(result.base64, 'base64');
ext = 'png';
} else if (result.url) {
buf = await downloadToBuffer(result.url);
const urlExt = path.extname(new URL(result.url).pathname).replace('.', '');
ext = ['png', 'jpg', 'jpeg', 'webp'].includes(urlExt) ? urlExt : 'png';
} else {
log.warn(`Image generation returned no data for ${req.elementId}`);
continue;
}
const filename = `${req.elementId}.${ext}`;
await fs.writeFile(path.join(mediaDir, filename), buf);
mediaMap[req.elementId] = mediaServingUrl(baseUrl, classroomId, `media/${filename}`);
log.info(`Generated image: ${filename}`);
} catch (err) {
log.warn(`Image generation failed for ${req.elementId}:`, err);
}
}
};
const generateVideos = async () => {
for (const req of videoRequests) {
try {
const providerId = videoProviderIds[0] as VideoProviderId;
const apiKey = resolveVideoApiKey(providerId);
if (!apiKey) {
log.warn(`No API key for video provider "${providerId}", skipping ${req.elementId}`);
continue;
}
const providerConfig = VIDEO_PROVIDERS[providerId];
const model = providerConfig?.models?.[0]?.id;
const normalized = normalizeVideoOptions(providerId, {
prompt: req.prompt,
aspectRatio: (req.aspectRatio as '16:9' | '4:3' | '1:1' | '9:16') || '16:9',
});
const result = await generateVideo(
{ providerId, apiKey, baseUrl: resolveVideoBaseUrl(providerId), model },
normalized,
);
const buf = await downloadToBuffer(result.url);
const filename = `${req.elementId}.mp4`;
await fs.writeFile(path.join(mediaDir, filename), buf);
mediaMap[req.elementId] = mediaServingUrl(baseUrl, classroomId, `media/${filename}`);
log.info(`Generated video: ${filename}`);
} catch (err) {
log.warn(`Video generation failed for ${req.elementId}:`, err);
}
}
};
await Promise.all([generateImages(), generateVideos()]);
return mediaMap;
}
// ---------------------------------------------------------------------------
// Placeholder replacement in scene content
// ---------------------------------------------------------------------------
export function replaceMediaPlaceholders(scenes: Scene[], mediaMap: Record<string, string>): void {
if (Object.keys(mediaMap).length === 0) return;
for (const scene of scenes) {
if (scene.type !== 'slide') continue;
const canvas = (
scene.content as {
canvas?: { elements?: Array<{ id: string; src?: string; type?: string }> };
}
)?.canvas;
if (!canvas?.elements) continue;
for (const el of canvas.elements) {
if (
(el.type === 'image' || el.type === 'video') &&
typeof el.src === 'string' &&
isMediaPlaceholder(el.src) &&
mediaMap[el.src]
) {
el.src = mediaMap[el.src];
}
}
}
}
// ---------------------------------------------------------------------------
// TTS generation
// ---------------------------------------------------------------------------
export async function generateTTSForClassroom(
scenes: Scene[],
classroomId: string,
baseUrl: string,
): Promise<void> {
const audioDir = path.join(CLASSROOMS_DIR, classroomId, 'audio');
await ensureDir(audioDir);
// Resolve TTS provider (exclude browser-native-tts)
const ttsProviderIds = Object.keys(getServerTTSProviders()).filter(
(id) => id !== 'browser-native-tts',
);
if (ttsProviderIds.length === 0) {
log.warn('No server TTS provider configured, skipping TTS generation');
return;
}
const providerId = ttsProviderIds[0] as TTSProviderId;
const apiKey = resolveTTSApiKey(providerId);
if (!apiKey) {
log.warn(`No API key for TTS provider "${providerId}", skipping TTS generation`);
return;
}
const ttsBaseUrl = resolveTTSBaseUrl(providerId) || TTS_PROVIDERS[providerId]?.defaultBaseUrl;
const voice = DEFAULT_TTS_VOICES[providerId] || 'default';
const format = TTS_PROVIDERS[providerId]?.supportedFormats?.[0] || 'mp3';
for (const scene of scenes) {
if (!scene.actions) continue;
// Split long speech actions into multiple shorter ones before TTS generation,
// mirroring the client-side approach. Each sub-action gets its own audio file.
scene.actions = splitLongSpeechActions(scene.actions, providerId);
for (const action of scene.actions) {
if (action.type !== 'speech' || !(action as SpeechAction).text) continue;
const speechAction = action as SpeechAction;
const audioId = `tts_${action.id}`;
try {
const result = await generateTTS(
{ providerId, apiKey, baseUrl: ttsBaseUrl, voice, speed: speechAction.speed },
speechAction.text,
);
const filename = `${audioId}.${format}`;
await fs.writeFile(path.join(audioDir, filename), result.audio);
speechAction.audioId = audioId;
speechAction.audioUrl = mediaServingUrl(baseUrl, classroomId, `audio/${filename}`);
log.info(`Generated TTS: ${filename} (${result.audio.length} bytes)`);
} catch (err) {
log.warn(`TTS generation failed for action ${action.id}:`, err);
}
}
}
}
+1
View File
@@ -40,6 +40,7 @@ export interface SpeechAction extends ActionBase {
type: 'speech';
text: string;
audioId?: string;
audioUrl?: string; // Server-generated TTS audio URL
voice?: string;
speed?: number; // default 1.0
}
+21 -3
View File
@@ -22,13 +22,31 @@ export class AudioPlayer {
private playbackRate: number = 1;
/**
* Play audio (from IndexedDB pre-generated cache)
* Play audio (from URL or IndexedDB pre-generated cache)
* @param audioId Audio ID
* @param audioUrl Optional server-generated audio URL (takes priority over IndexedDB)
* @returns true if audio started playing, false if no audio (TTS disabled or not generated)
*/
public async play(audioId: string): Promise<boolean> {
public async play(audioId: string, audioUrl?: string): Promise<boolean> {
try {
// Get audio from database
// 1. Try audioUrl first (server-generated TTS)
if (audioUrl) {
this.stop();
this.audio = new Audio();
this.audio.src = audioUrl;
if (this.muted) this.audio.volume = 0;
else this.audio.volume = this.volume;
this.audio.defaultPlaybackRate = this.playbackRate;
this.audio.playbackRate = this.playbackRate;
this.audio.addEventListener('ended', () => {
this.onEndedCallback?.();
});
await this.audio.play();
this.audio.playbackRate = this.playbackRate;
return true;
}
// 2. Fall back to IndexedDB (client-generated TTS)
const audioRecord = await db.audioFiles.get(audioId);
if (!audioRecord) {
+2
View File
@@ -81,6 +81,8 @@ Load [references/provider-keys.md](references/provider-keys.md).
Use this before starting classroom generation. Recommend a provider path and tell the user exactly which config file to edit themselves. If generation later fails due to provider/model/auth issues, return to this phase and direct the user to update the same server-side config files.
After the core LLM key is configured, ask the user if they want to enable optional features (web search, image generation, video generation, TTS). Each requires its own provider key — see the "Optional Features" section in provider-keys.md.
### 4. Start And Verify OpenMAIC
After the user has chosen a startup mode and configured keys, start OpenMAIC using the chosen method, then verify the service with `GET {url}/api/health`.
+29 -1
View File
@@ -32,9 +32,37 @@ Request body:
Only send supported content fields:
- `requirement`
- `requirement` (required)
- optional `pdfContent`
- optional `language`
- optional `enableWebSearch` (boolean) — include web search context in outline generation
- optional `enableImageGeneration` (boolean) — allow image generation metadata in outlines
- optional `enableVideoGeneration` (boolean) — allow video generation metadata in outlines
- optional `enableTTS` (boolean) — reserved for future server-side TTS generation
- optional `agentMode` (`"default"` | `"generate"`) — controls agent profile strategy:
- `"default"` (or omitted): uses built-in default agents
- `"generate"`: uses LLM to generate custom agent profiles tailored to the course content
All optional boolean fields default to `false` when omitted. Omitting them preserves backward compatibility.
### Feature Detection
Before sending optional feature flags, query `GET {url}/api/health` and check the `capabilities` object:
```json
{
"status": "ok",
"version": "...",
"capabilities": {
"webSearch": true,
"imageGeneration": false,
"videoGeneration": false,
"tts": false
}
}
```
Only set a feature flag to `true` if the corresponding capability is `true`. If the server does not return `capabilities` (older version), do not send the new fields.
Do not rely on request-time model or provider override parameters.
@@ -24,6 +24,10 @@ Follow the same generation flow as [generate-flow.md](generate-flow.md) with the
- **Authorization**: Include header `Authorization: Bearer <access-code>` on all API requests
- **Classroom URL**: `https://open.maic.chat/classroom/{id}`
### Feature Detection in Hosted Mode
Before generating, query `GET https://open.maic.chat/api/health` (with auth header) to check `capabilities`. Automatically include optional feature flags (`enableWebSearch`, `enableImageGeneration`, etc.) based on what the server supports. Do not send new fields if the server does not return `capabilities` (older version). This ensures forward compatibility — the hosted instance may update on a different schedule than the local codebase.
## Quota
- 10 generations per day, independent of web UI quota
@@ -144,3 +144,36 @@ Avoid as the first move:
- Wait for the user to confirm they finished editing before continuing.
- Do not request the literal key.
- If provider/model/auth errors happen later, tell the user exactly which config entry to fix and wait for confirmation before retrying.
## Optional Features
These features require additional provider keys beyond the core LLM provider. Ask the user if they want to enable any of these after the core LLM key is configured.
| Feature | Env Variable(s) | Description |
|---------|-----------------|-------------|
| Web Search | `TAVILY_API_KEY` | Enriches outlines with real-time web research |
| Image Generation | `IMAGE_SEEDREAM_API_KEY`, `IMAGE_QWEN_IMAGE_API_KEY`, `IMAGE_NANO_BANANA_API_KEY` | Generates images for slides (any one suffices) |
| Video Generation | `VIDEO_SEEDANCE_API_KEY`, `VIDEO_KLING_API_KEY`, `VIDEO_VEO_API_KEY`, `VIDEO_SORA_API_KEY` | Generates short videos (any one suffices) |
| TTS | `TTS_OPENAI_API_KEY`, `TTS_AZURE_API_KEY`, `TTS_GLM_API_KEY`, `TTS_QWEN_API_KEY` | Text-to-speech narration (any one suffices) |
These are all optional. The classroom generation works without them — they only unlock richer content.
Alternatively, configure via `server-providers.yml`:
```yaml
web-search:
tavily:
apiKey: tvly-...
image:
seedream:
apiKey: ...
video:
seedance:
apiKey: ...
tts:
openai-tts:
apiKey: sk-...
```