mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 09:24:43 +08:00
* feat(model): optional per-stage LLM model routing (#745) Add a config-only stage -> model map consulted during model resolution, falling back to DEFAULT_MODEL when unset (zero behavior change unless opted in). - lib/server/model-routes.ts: parse/validate/cache MODEL_ROUTES env JSON; LLM_STAGES registry of routable stages; getStageModel(stage). - resolveModel: resolution order x-model > stage route > DEFAULT_MODEL > builtin; thread optional `stage` through resolveModelFromHeaders/FromRequest. - Wire each route's resolveModel call site to its stage; classroom-generation resolves generate-classroom and web-search-query-rewrite independently. - .env.example: document MODEL_ROUTES; unit tests for both modules. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(model): address cross-review findings for per-stage routing (#745) - classroom-generation: resolve the web-search-query-rewrite model lazily, only when explicitly routed and inside the web-search branch, with try/catch fallback to the classroom model. Fixes a regression where a misconfigured optional route (keyless/invalid provider) threw and aborted ALL classroom generation even with web search disabled; also avoids wasted resolution on the common no-search path and preserves classroom-model inheritance when unrouted. - resolve-model: type the stage param as LlmStage across resolveModel/ resolveModelFromHeaders/resolveModelFromRequest so a mistyped stage literal is a compile error instead of silently falling through to DEFAULT_MODEL. - model-routes: simplify the cache to a process singleton (matching provider-config), dropping the raw-string-keyed cache and redundant STAGE_SET. - .env.example: clarify x-model precedence and that a route to an unconfigured provider fails at request time (no startup validation), like a bad DEFAULT_MODEL. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(model): stage route takes precedence over client x-model (#745) Cross-review (codex P2 + Claude tracer) found that the browser UI always sends its saved model as x-model, so the previous x-model > stage order made MODEL_ROUTES inert for normal UI traffic on exactly the heavy stages the RFC targets (scene-content, quiz, pbl-chat, chat). Flip resolution to stage route > x-model > DEFAULT_MODEL: a configured route is the operator's deliberate per-stage choice and wins, while unrouted stages still honor the client's x-model. Update docs and tests accordingly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(model): prettier-format changed files for #745 CI Run the repo formatter (pnpm check / prettier) on the changed routes, the new model-routes module, and the new tests so `prettier . --check` (CI) passes; also correct a stale precedence comment in the resolve-model test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(model): isolate routed model from client connection params (#745) When a stage route overrides the client's x-model and points at a different provider, the client-sent apiKey/baseUrl/providerType (for the client's model) must not bleed onto the routed provider — otherwise a routed Anthropic model would be constructed with the client's OpenAI providerType/key and fail. A routed model now resolves its connection params purely from server config, as if no x-model was sent. Unrouted stages still honor the client params. Adds tests covering both the drop (routed) and keep (unrouted) cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(model): correct graceful-degradation comment in classroom (#745) resolveModel does not throw on a missing key; clarify that a keyless route surfaces later in callLLM (degraded by the outer try/catch), while only a resolution-time failure is caught at the rewrite re-resolve. Comment-only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(model): route scene-content per scene type (#745) Extend MODEL_ROUTES with composite keys scene-content:<type> for the four core scene types (slide/quiz/interactive/pbl). getStageModel now resolves composite `a:b` stages most-specific-first, trimming `:` segments, so scene-content:<type> falls back to the base scene-content route, then x-model, then DEFAULT_MODEL. The scene-content route derives its stage from outline.type. Backward compatible: with no composite key configured, behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(model): per-stage thinking effort + drop hardcoded gpt fallback (#745) MODEL_ROUTES values may now be {model, effort} (not just a model string) to pin the thinking effort per stage. Arbitration mirrors model routing: a routed stage with effort set wins over the client's thinking (effort "none" disables); routed with no effort uses the model's default and drops client thinking; unrouted stages keep the client thinking. resolveModel is the single arbiter — body thinking is threaded in via resolveModelFromRequest. classroom-generation passes the resolved thinking into its callLLM calls so generate-classroom and web-search-query-rewrite honor route effort too. Also: removed the hardcoded `|| 'gpt-5.4-mini'` fallback in resolveModel — if no model resolves (no route / x-model / DEFAULT_MODEL) it now throws instead of silently picking a vendor default. model-routes: getStageRoute returns {model, effort?}; getStageModel delegates to it. Docs + tests updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(model): route value carries full ThinkingConfig (#745) Per-stage routing accepts the full unified ThinkingConfig in the route value ({model, thinking:{mode,effort,level,enabled,budgetTokens,excludeReasoningOutput}}) instead of just an effort string. The route's thinking is passed through resolveModel and normalized per the model's capability by callLLM, so budgetTokens (qwen), level (Gemini), enabled/mode, etc. all work. (qwen3.7-plus/max thinking capability now comes from upstream #753, so this no longer registers them.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(model): honor routed thinking for chat-adapter; doc fixes (#745) Cross-review (codex P2 + Claude) found chat-adapter was the only routable stage that ignored the route's thinking: it took the routed model but built its own thinkingConfig from the request body, so operator-pinned thinking never applied and client thinking still leaked onto a routed model. Now pass the client thinking into resolveModel and use the resolved thinkingConfig (route-pinned for a routed stage, client otherwise; defaults to disabled for low-latency chat). Also: clarify DEFAULT_MODEL is now required for server-side stages (the hardcoded gpt-5.4-mini fallback was removed → resolveModel throws if nothing resolves), and fix a stale {model,effort}→{model,thinking} doc comment. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
+36
-1
@@ -203,12 +203,47 @@ WEB_SEARCH_MINIMAX_BASE_URL=https://api.minimaxi.com
|
||||
|
||||
# --- Misc ---------------------------------------------------------------------
|
||||
|
||||
# Optional server-side default model for API routes like /api/generate-classroom
|
||||
# Server-side default model for API routes like /api/generate-classroom.
|
||||
# Required for server-side stages (those that don't receive a client x-model):
|
||||
# resolveModel throws if a stage resolves to no model (no MODEL_ROUTES entry, no
|
||||
# x-model, no DEFAULT_MODEL) — there is intentionally no hardcoded vendor fallback.
|
||||
# Example: anthropic:claude-3-5-haiku-20241022 or google:gemini-3-flash-preview
|
||||
# OpenAI example: openai:gpt-5.5
|
||||
# MiniMax example: minimax:MiniMax-M2.7-highspeed
|
||||
DEFAULT_MODEL=
|
||||
|
||||
# Optional per-stage model routing (#745). A JSON object mapping a generation
|
||||
# stage to a model string (`provider:model`). Resolution order for a stage:
|
||||
# stage route > x-model (client) > DEFAULT_MODEL. A configured route is the
|
||||
# operator's deliberate choice and wins even when the browser sends its saved
|
||||
# model as x-model. Unset = identical to today (everything uses DEFAULT_MODEL).
|
||||
# Unlisted stages fall back to x-model then DEFAULT_MODEL, so you can override
|
||||
# just one or two and leave the rest to the client/default.
|
||||
# Point a stage only at a server-configured provider (its key resolvable); a
|
||||
# route to an unconfigured/invalid model fails at request time for that stage,
|
||||
# the same way a bad DEFAULT_MODEL would (no startup validation).
|
||||
# Routable stages: scene-outlines-stream, scene-content, scene-actions,
|
||||
# agent-profiles, quiz-grade, pbl-chat, chat-adapter, generate-classroom,
|
||||
# web-search-query-rewrite.
|
||||
# scene-content can also be routed per scene type with composite keys:
|
||||
# scene-content:slide, scene-content:quiz, scene-content:interactive,
|
||||
# scene-content:pbl. A type falls back to the base scene-content route when it
|
||||
# has no key of its own (so scene-content:<type> > scene-content > x-model >
|
||||
# DEFAULT_MODEL).
|
||||
# A route value can be a model string, OR an object {"model","thinking"} where
|
||||
# `thinking` is the full ThinkingConfig: mode (default|disabled|enabled|auto),
|
||||
# effort (none|minimal|low|medium|high|xhigh|max), level (minimal|low|medium|
|
||||
# high, Gemini), enabled (bool), budgetTokens (number), excludeReasoningOutput
|
||||
# (bool). It is normalized per the model's capability. When a stage is routed:
|
||||
# a set `thinking` wins over the client's thinking; with no `thinking` the routed
|
||||
# model uses its own default and the client's thinking is dropped. Unrouted
|
||||
# stages keep the client thinking.
|
||||
# Example: cheap default, stronger model only for the heavy/conversational stages:
|
||||
# MODEL_ROUTES='{"scene-content":"openai:gpt-5.4","scene-actions":"openai:gpt-5.4","pbl-chat":"anthropic:claude-sonnet-4","chat-adapter":"anthropic:claude-sonnet-4"}'
|
||||
# Example: per scene type + pinned thinking (qwen budget, deepseek off):
|
||||
# MODEL_ROUTES='{"scene-content:interactive":{"model":"qwen:qwen3.7-plus","thinking":{"enabled":true,"budgetTokens":8000}},"scene-content:quiz":{"model":"deepseek:deepseek-v4-pro","thinking":{"enabled":false}}}'
|
||||
# MODEL_ROUTES=
|
||||
|
||||
# LOG_LEVEL=info
|
||||
# LOG_FORMAT=pretty
|
||||
# LLM_THINKING_DISABLED=false
|
||||
|
||||
+11
-4
@@ -68,11 +68,16 @@ export async function POST(req: NextRequest) {
|
||||
model: languageModel,
|
||||
apiKey: resolvedApiKey,
|
||||
providerId,
|
||||
thinkingConfig: resolvedThinking,
|
||||
} = await resolveModel({
|
||||
modelString: body.model,
|
||||
stage: 'chat-adapter',
|
||||
apiKey: body.apiKey,
|
||||
baseUrl: body.baseUrl,
|
||||
providerType: body.providerType,
|
||||
// Let resolveModel arbitrate thinking too: a routed chat-adapter's thinking
|
||||
// wins, an unrouted one honors this client thinking (see resolve-model.ts).
|
||||
thinkingConfig: body.thinkingConfig ?? body.thinking,
|
||||
});
|
||||
|
||||
if (isProviderKeyRequired(providerId) && !resolvedApiKey) {
|
||||
@@ -117,10 +122,12 @@ export async function POST(req: NextRequest) {
|
||||
try {
|
||||
startHeartbeat();
|
||||
|
||||
// Default: thinking disabled for low-latency chat. UI requests send
|
||||
// `thinkingConfig`; eval harnesses can still opt in via `thinking`.
|
||||
const thinkingConfig: ThinkingConfig = body.thinkingConfig ??
|
||||
body.thinking ?? { mode: 'disabled', enabled: false };
|
||||
// Use the resolved thinking (route-pinned for a routed chat-adapter,
|
||||
// else the client's). Default to disabled for low-latency chat.
|
||||
const thinkingConfig: ThinkingConfig = resolvedThinking ?? {
|
||||
mode: 'disabled',
|
||||
enabled: false,
|
||||
};
|
||||
|
||||
const generator = statelessGenerate(
|
||||
{
|
||||
|
||||
@@ -76,7 +76,7 @@ export async function POST(req: NextRequest) {
|
||||
model: languageModel,
|
||||
modelString: _modelString,
|
||||
thinkingConfig,
|
||||
} = await resolveModelFromRequest(req, body);
|
||||
} = await resolveModelFromRequest(req, body, 'agent-profiles');
|
||||
modelString = _modelString;
|
||||
|
||||
// ── Build prompt ──
|
||||
|
||||
@@ -84,7 +84,7 @@ export async function POST(req: NextRequest) {
|
||||
modelInfo,
|
||||
modelString,
|
||||
thinkingConfig,
|
||||
} = await resolveModelFromRequest(req, body);
|
||||
} = await resolveModelFromRequest(req, body, 'scene-actions');
|
||||
outlineTitle = outline?.title;
|
||||
resolvedModelString = modelString;
|
||||
|
||||
|
||||
@@ -73,12 +73,15 @@ export async function POST(req: NextRequest) {
|
||||
const outline: SceneOutline = { ...rawOutline };
|
||||
|
||||
// ── Model resolution from request headers/body ──
|
||||
// Route per scene-content type (e.g. `scene-content:quiz`); getStageModel
|
||||
// falls back to the base `scene-content` route when the type is unrouted.
|
||||
const stage = outline.type ? (`scene-content:${outline.type}` as const) : 'scene-content';
|
||||
const {
|
||||
model: languageModel,
|
||||
modelInfo,
|
||||
modelString,
|
||||
thinkingConfig,
|
||||
} = await resolveModelFromRequest(req, body);
|
||||
} = await resolveModelFromRequest(req, body, stage);
|
||||
outlineTitle = rawOutline?.title;
|
||||
resolvedModelString = modelString;
|
||||
|
||||
|
||||
@@ -252,7 +252,7 @@ export async function POST(req: NextRequest) {
|
||||
modelInfo,
|
||||
modelString,
|
||||
thinkingConfig,
|
||||
} = await resolveModelFromRequest(req, body);
|
||||
} = await resolveModelFromRequest(req, body, 'scene-outlines-stream');
|
||||
resolvedModelString = modelString;
|
||||
|
||||
if (!body.requirements) {
|
||||
|
||||
@@ -36,7 +36,7 @@ export async function POST(req: NextRequest) {
|
||||
}
|
||||
|
||||
// Get model config from request headers/body
|
||||
const { model, thinkingConfig } = await resolveModelFromRequest(req, body);
|
||||
const { model, thinkingConfig } = await resolveModelFromRequest(req, body, 'pbl-chat');
|
||||
|
||||
// Build context for the agent, differentiating question vs judge
|
||||
let issueContext = '';
|
||||
|
||||
@@ -44,7 +44,11 @@ export async function POST(req: NextRequest) {
|
||||
}
|
||||
|
||||
// Resolve model from request headers/body
|
||||
const { model: languageModel, thinkingConfig } = await resolveModelFromRequest(req, body);
|
||||
const { model: languageModel, thinkingConfig } = await resolveModelFromRequest(
|
||||
req,
|
||||
body,
|
||||
'quiz-grade',
|
||||
);
|
||||
|
||||
const isZh = language === 'zh-CN';
|
||||
|
||||
|
||||
@@ -78,7 +78,11 @@ export async function POST(req: NextRequest) {
|
||||
|
||||
let aiCall: AICallFn | undefined;
|
||||
try {
|
||||
const { model: languageModel, thinkingConfig } = await resolveModelFromRequest(req, body);
|
||||
const { model: languageModel, thinkingConfig } = await resolveModelFromRequest(
|
||||
req,
|
||||
body,
|
||||
'web-search-query-rewrite',
|
||||
);
|
||||
aiCall = async (systemPrompt, userPrompt) => {
|
||||
const result = await callLLM(
|
||||
{
|
||||
|
||||
@@ -18,6 +18,7 @@ import { createLogger } from '@/lib/logger';
|
||||
import { isProviderKeyRequired } from '@/lib/ai/providers';
|
||||
import { resolveClassroomWebSearchConfig } from '@/lib/server/web-search-config';
|
||||
import { resolveModel } from '@/lib/server/resolve-model';
|
||||
import { getStageModel } from '@/lib/server/model-routes';
|
||||
import { resolveVocationalActive } from '@/lib/config/feature-flags';
|
||||
import { buildSearchQuery } from '@/lib/server/search-query-builder';
|
||||
import { formatSearchResultsAsContext, searchWeb } from '@/lib/web-search';
|
||||
@@ -184,7 +185,8 @@ export async function generateClassroom(
|
||||
modelString,
|
||||
providerId,
|
||||
apiKey,
|
||||
} = await resolveModel({});
|
||||
thinkingConfig: classroomThinking,
|
||||
} = await resolveModel({ stage: 'generate-classroom' });
|
||||
log.info(`Using server-configured model: ${modelString}`);
|
||||
|
||||
// Fail fast if the resolved provider has no API key configured
|
||||
@@ -195,6 +197,14 @@ export async function generateClassroom(
|
||||
);
|
||||
}
|
||||
|
||||
// The web-search query rewrite is a light, separable stage operators may route
|
||||
// to a cheaper model. It defaults to the classroom model and is only
|
||||
// re-resolved lazily (inside the web-search branch, and only when a route is
|
||||
// configured). This keeps a misconfigured optional route from aborting all
|
||||
// classroom generation, and skips the extra resolution when web search is off.
|
||||
let searchQueryModel = languageModel;
|
||||
let searchQueryThinking = classroomThinking;
|
||||
|
||||
const aiCall: AICallFn = async (systemPrompt, userPrompt, _images) => {
|
||||
const result = await callLLM(
|
||||
{
|
||||
@@ -206,6 +216,8 @@ export async function generateClassroom(
|
||||
maxOutputTokens: modelInfo?.outputWindow,
|
||||
},
|
||||
'generate-classroom',
|
||||
undefined,
|
||||
classroomThinking,
|
||||
);
|
||||
return result.text;
|
||||
};
|
||||
@@ -213,7 +225,7 @@ export async function generateClassroom(
|
||||
const searchQueryAiCall: AICallFn = async (systemPrompt, userPrompt, _images) => {
|
||||
const result = await callLLM(
|
||||
{
|
||||
model: languageModel,
|
||||
model: searchQueryModel,
|
||||
messages: [
|
||||
{ role: 'system', content: systemPrompt },
|
||||
{ role: 'user', content: userPrompt },
|
||||
@@ -221,6 +233,8 @@ export async function generateClassroom(
|
||||
maxOutputTokens: 256,
|
||||
},
|
||||
'web-search-query-rewrite',
|
||||
undefined,
|
||||
searchQueryThinking,
|
||||
);
|
||||
return result.text;
|
||||
};
|
||||
@@ -243,6 +257,24 @@ export async function generateClassroom(
|
||||
if (input.enableWebSearch) {
|
||||
const webSearchConfig = resolveClassroomWebSearchConfig(input);
|
||||
if (webSearchConfig) {
|
||||
// Re-resolve the query-rewrite model only when explicitly routed. If
|
||||
// resolution itself fails (e.g. unknown provider in the route), fall back
|
||||
// to the classroom model here; a route with a missing key resolves fine
|
||||
// and surfaces only later in callLLM, which the outer try/catch below
|
||||
// degrades gracefully — either way the pipeline still works.
|
||||
const rewriteRoute = getStageModel('web-search-query-rewrite');
|
||||
if (rewriteRoute) {
|
||||
try {
|
||||
const rewriteResolved = await resolveModel({ stage: 'web-search-query-rewrite' });
|
||||
searchQueryModel = rewriteResolved.model;
|
||||
searchQueryThinking = rewriteResolved.thinkingConfig;
|
||||
} catch (err) {
|
||||
log.warn(
|
||||
`web-search-query-rewrite route "${rewriteRoute}" unavailable; using classroom model for query rewrite`,
|
||||
err,
|
||||
);
|
||||
}
|
||||
}
|
||||
try {
|
||||
const searchQuery = await buildSearchQuery(requirement, pdfText, searchQueryAiCall);
|
||||
|
||||
|
||||
@@ -0,0 +1,207 @@
|
||||
/**
|
||||
* Per-stage LLM model routing (issue #745).
|
||||
*
|
||||
* Optional, config-only overrides that map a generation *stage* to a specific
|
||||
* model string. Consulted during model resolution and falling back to today's
|
||||
* behavior (`DEFAULT_MODEL`) when unset — zero behavior change unless opted in.
|
||||
*
|
||||
* Surface: a single JSON env var `MODEL_ROUTES`. Each value is a model string in
|
||||
* the canonical `provider:model` format (see parseModelString), OR an object
|
||||
* `{model, thinking}` where `thinking` is the full ThinkingConfig abstraction
|
||||
* (mode/effort/level/enabled/budgetTokens/excludeReasoningOutput) — normalized
|
||||
* per the model's capability by callLLM. e.g.
|
||||
*
|
||||
* DEFAULT_MODEL=openai:gpt-5.4-mini
|
||||
* MODEL_ROUTES='{"scene-content":"openai:gpt-5.4","pbl-chat":{"model":"anthropic:claude-sonnet-4","thinking":{"enabled":false}}}'
|
||||
*
|
||||
* Only the *routable* stages below are valid keys — each is backed by a real
|
||||
* `resolveModel` call site. Downstream sub-calls (e.g. `pbl-generate`,
|
||||
* `chat-adapter-stream`) inherit their parent stage's resolved model.
|
||||
*/
|
||||
|
||||
import { createLogger } from '@/lib/logger';
|
||||
import type {
|
||||
ThinkingConfig,
|
||||
ThinkingEffort,
|
||||
ThinkingLevel,
|
||||
ThinkingMode,
|
||||
} from '@/lib/types/provider';
|
||||
|
||||
const log = createLogger('model-routes');
|
||||
|
||||
const VALID_MODES: readonly ThinkingMode[] = ['default', 'disabled', 'enabled', 'auto'];
|
||||
const VALID_EFFORTS: readonly ThinkingEffort[] = [
|
||||
'none',
|
||||
'minimal',
|
||||
'low',
|
||||
'medium',
|
||||
'high',
|
||||
'xhigh',
|
||||
'max',
|
||||
];
|
||||
const VALID_LEVELS: readonly ThinkingLevel[] = ['minimal', 'low', 'medium', 'high'];
|
||||
|
||||
/** A resolved route entry: the model string plus an optional full thinking config. */
|
||||
export interface StageRoute {
|
||||
model: string;
|
||||
/**
|
||||
* Full thinking config for this stage (the unified ThinkingConfig abstraction:
|
||||
* mode / effort / level / enabled / budgetTokens / excludeReasoningOutput).
|
||||
* Passed through to callLLM, which normalizes it against the model's capability.
|
||||
*/
|
||||
thinking?: ThinkingConfig;
|
||||
}
|
||||
|
||||
/** Validate/sanitize a route's `thinking` object into a ThinkingConfig (drops bad fields with a warn). */
|
||||
function parseThinking(key: string, raw: unknown): ThinkingConfig | undefined {
|
||||
if (!raw || typeof raw !== 'object' || Array.isArray(raw)) {
|
||||
log.warn(`"thinking" for stage "${key}" must be an object in MODEL_ROUTES; ignored.`);
|
||||
return undefined;
|
||||
}
|
||||
const o = raw as Record<string, unknown>;
|
||||
const out: ThinkingConfig = {};
|
||||
const checkEnum = <T>(field: string, val: unknown, valid: readonly T[]): T | undefined => {
|
||||
if (val === undefined) return undefined;
|
||||
if (typeof val === 'string' && (valid as readonly string[]).includes(val)) return val as T;
|
||||
log.warn(
|
||||
`Invalid ${field} "${String(val)}" for stage "${key}" ignored. Valid: ${valid.join(', ')}`,
|
||||
);
|
||||
return undefined;
|
||||
};
|
||||
const mode = checkEnum<ThinkingMode>('mode', o.mode, VALID_MODES);
|
||||
if (mode) out.mode = mode;
|
||||
const effort = checkEnum<ThinkingEffort>('effort', o.effort, VALID_EFFORTS);
|
||||
if (effort) out.effort = effort;
|
||||
const level = checkEnum<ThinkingLevel>('level', o.level, VALID_LEVELS);
|
||||
if (level) out.level = level;
|
||||
if (o.enabled !== undefined) {
|
||||
if (typeof o.enabled === 'boolean') out.enabled = o.enabled;
|
||||
else
|
||||
log.warn(
|
||||
`Invalid enabled "${String(o.enabled)}" for stage "${key}" ignored (must be boolean).`,
|
||||
);
|
||||
}
|
||||
if (o.budgetTokens !== undefined) {
|
||||
if (typeof o.budgetTokens === 'number') out.budgetTokens = o.budgetTokens;
|
||||
else
|
||||
log.warn(
|
||||
`Invalid budgetTokens "${String(o.budgetTokens)}" for stage "${key}" ignored (must be number).`,
|
||||
);
|
||||
}
|
||||
if (o.excludeReasoningOutput !== undefined) {
|
||||
if (typeof o.excludeReasoningOutput === 'boolean')
|
||||
out.excludeReasoningOutput = o.excludeReasoningOutput;
|
||||
else log.warn(`Invalid excludeReasoningOutput for stage "${key}" ignored (must be boolean).`);
|
||||
}
|
||||
return Object.keys(out).length ? out : undefined;
|
||||
}
|
||||
|
||||
/**
|
||||
* Stages that can be independently routed to a model. Each value is a valid
|
||||
* `MODEL_ROUTES` key; the base entries also mirror a `callLLM` source label.
|
||||
*
|
||||
* `scene-content:<type>` are finer-grained composite keys: when a scene-content
|
||||
* request carries an `outline.type`, it routes via the composite key and falls
|
||||
* back to the base `scene-content` route (see getStageModel). Only the four
|
||||
* core scene types are routable; interactive widget sub-types are not split.
|
||||
*/
|
||||
export const LLM_STAGES = [
|
||||
'scene-outlines-stream',
|
||||
'scene-content',
|
||||
'scene-content:slide',
|
||||
'scene-content:quiz',
|
||||
'scene-content:interactive',
|
||||
'scene-content:pbl',
|
||||
'scene-actions',
|
||||
'agent-profiles',
|
||||
'quiz-grade',
|
||||
'pbl-chat',
|
||||
'chat-adapter',
|
||||
'generate-classroom',
|
||||
'web-search-query-rewrite',
|
||||
] as const;
|
||||
|
||||
export type LlmStage = (typeof LLM_STAGES)[number];
|
||||
|
||||
/** Parsed once per process (env is read at startup; tests reset via vi.resetModules). */
|
||||
let _routes: Record<string, StageRoute> | null = null;
|
||||
|
||||
/** Parse one MODEL_ROUTES value (string model, or {model, thinking}) into a StageRoute. */
|
||||
function parseRouteValue(key: string, value: unknown): StageRoute | undefined {
|
||||
if (typeof value === 'string') {
|
||||
return value.trim() ? { model: value.trim() } : undefined;
|
||||
}
|
||||
if (value && typeof value === 'object' && !Array.isArray(value)) {
|
||||
const obj = value as Record<string, unknown>;
|
||||
const model = typeof obj.model === 'string' ? obj.model.trim() : '';
|
||||
if (!model) {
|
||||
log.warn(`Route for stage "${key}" has no model string in MODEL_ROUTES; ignored.`);
|
||||
return undefined;
|
||||
}
|
||||
const route: StageRoute = { model };
|
||||
if (obj.thinking !== undefined) {
|
||||
const thinking = parseThinking(key, obj.thinking);
|
||||
if (thinking) route.thinking = thinking;
|
||||
}
|
||||
return route;
|
||||
}
|
||||
log.warn(`Invalid route value for stage "${key}" in MODEL_ROUTES ignored.`);
|
||||
return undefined;
|
||||
}
|
||||
|
||||
function loadRoutes(): Record<string, StageRoute> {
|
||||
if (_routes) return _routes;
|
||||
|
||||
const routes: Record<string, StageRoute> = {};
|
||||
const raw = process.env.MODEL_ROUTES?.trim();
|
||||
if (raw) {
|
||||
try {
|
||||
const parsed = JSON.parse(raw) as unknown;
|
||||
if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) {
|
||||
for (const [key, value] of Object.entries(parsed as Record<string, unknown>)) {
|
||||
if (!(LLM_STAGES as readonly string[]).includes(key)) {
|
||||
log.warn(
|
||||
`Unknown stage "${key}" in MODEL_ROUTES ignored. Valid stages: ${LLM_STAGES.join(', ')}`,
|
||||
);
|
||||
continue;
|
||||
}
|
||||
const route = parseRouteValue(key, value);
|
||||
if (route) routes[key] = route;
|
||||
}
|
||||
} else {
|
||||
log.error('MODEL_ROUTES must be a JSON object of stage -> model; ignoring.');
|
||||
}
|
||||
} catch (err) {
|
||||
log.error('Invalid MODEL_ROUTES JSON, ignoring (falling back to DEFAULT_MODEL).', err);
|
||||
}
|
||||
}
|
||||
|
||||
_routes = routes;
|
||||
return _routes;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the configured model string for a stage, or `undefined` when the
|
||||
* stage is unset/unconfigured (callers fall back to `DEFAULT_MODEL`).
|
||||
*
|
||||
* Composite `a:b` stages resolve most-specific-first: the full key is tried,
|
||||
* then successively shorter prefixes (e.g. `scene-content:quiz` →
|
||||
* `scene-content`). Plain stages (no colon) are a single exact lookup.
|
||||
*/
|
||||
export function getStageRoute(stage?: string): StageRoute | undefined {
|
||||
if (!stage) return undefined;
|
||||
const routes = loadRoutes();
|
||||
let key: string | undefined = stage;
|
||||
while (key) {
|
||||
const route = routes[key];
|
||||
if (route) return route;
|
||||
const lastColon = key.lastIndexOf(':');
|
||||
key = lastColon > 0 ? key.slice(0, lastColon) : undefined;
|
||||
}
|
||||
return undefined;
|
||||
}
|
||||
|
||||
/** Convenience: the resolved model string for a stage (route's `model`). */
|
||||
export function getStageModel(stage?: string): string | undefined {
|
||||
return getStageRoute(stage)?.model;
|
||||
}
|
||||
+60
-11
@@ -15,6 +15,7 @@ import {
|
||||
resolveProxy,
|
||||
} from '@/lib/server/provider-config';
|
||||
import { validateUrlForSSRF } from '@/lib/server/ssrf-guard';
|
||||
import { getStageRoute, type LlmStage } from '@/lib/server/model-routes';
|
||||
|
||||
export interface ResolvedModel extends ModelWithInfo {
|
||||
/** Original model string (e.g. "openai/gpt-4o-mini") */
|
||||
@@ -38,20 +39,52 @@ export interface ResolvedModel extends ModelWithInfo {
|
||||
*/
|
||||
export async function resolveModel(params: {
|
||||
modelString?: string;
|
||||
/**
|
||||
* Optional generation stage (a `callLLM` source label, e.g. 'scene-content').
|
||||
* When set and a route is configured via `MODEL_ROUTES`, the route wins for
|
||||
* this call — even over a client-sent `modelString` (x-model). Unrouted
|
||||
* stages fall back to `modelString` then `DEFAULT_MODEL`. See
|
||||
* lib/server/model-routes.ts.
|
||||
*/
|
||||
stage?: LlmStage;
|
||||
apiKey?: string;
|
||||
baseUrl?: string;
|
||||
providerType?: string;
|
||||
thinkingConfig?: ThinkingConfig;
|
||||
}): Promise<ResolvedModel> {
|
||||
const modelString = params.modelString || process.env.DEFAULT_MODEL || 'gpt-5.4-mini';
|
||||
// Resolution order: stage route > x-model > DEFAULT_MODEL.
|
||||
// A configured stage route is the operator's deliberate per-stage choice and
|
||||
// wins even over a client-sent x-model (otherwise the browser UI, which always
|
||||
// sends its saved model, would shadow every route). Unrouted stages fall back
|
||||
// to the client x-model, then DEFAULT_MODEL. There is intentionally no hardcoded
|
||||
// model fallback — if nothing resolves we fail loud rather than silently pick a
|
||||
// vendor default.
|
||||
const stageRoute = getStageRoute(params.stage);
|
||||
const stageModel = stageRoute?.model;
|
||||
const modelString = stageModel || params.modelString || process.env.DEFAULT_MODEL;
|
||||
if (!modelString) {
|
||||
throw new Error(
|
||||
'No model could be resolved. Configure DEFAULT_MODEL (and/or a MODEL_ROUTES entry for this stage), or send a model via x-model.',
|
||||
);
|
||||
}
|
||||
const { providerId, modelId } = parseModelString(modelString);
|
||||
|
||||
// When a stage route overrides the client's model, the client-sent connection
|
||||
// params (apiKey/baseUrl/providerType) belong to the client's *other* model
|
||||
// and must not bleed onto the routed provider — otherwise e.g. a routed
|
||||
// Anthropic model would be built with the client's OpenAI providerType/key.
|
||||
// A routed model resolves purely from server config, as if no x-model was sent.
|
||||
const routed = Boolean(stageModel);
|
||||
const clientApiKey = routed ? undefined : params.apiKey;
|
||||
const clientProviderType = routed ? undefined : params.providerType;
|
||||
const clientBaseUrlParam = routed ? undefined : params.baseUrl;
|
||||
|
||||
// Server-managed providers are admin-owned: the operator's key and base URL
|
||||
// are authoritative and any client-sent override is ignored. SSRF validation
|
||||
// therefore applies only to unmanaged providers, where the base URL really is
|
||||
// client-supplied. (Server-configured URLs are trusted by the operator.)
|
||||
const managed = isServerConfiguredProvider('providers', providerId);
|
||||
const clientBaseUrl = managed ? undefined : params.baseUrl || undefined;
|
||||
const clientBaseUrl = managed ? undefined : clientBaseUrlParam || undefined;
|
||||
if (clientBaseUrl && process.env.NODE_ENV === 'production') {
|
||||
const ssrfError = await validateUrlForSSRF(clientBaseUrl);
|
||||
if (ssrfError) {
|
||||
@@ -59,7 +92,7 @@ export async function resolveModel(params: {
|
||||
}
|
||||
}
|
||||
|
||||
const apiKey = resolveApiKey(providerId, params.apiKey || '');
|
||||
const apiKey = resolveApiKey(providerId, clientApiKey || '');
|
||||
const baseUrl = resolveBaseUrl(providerId, clientBaseUrl);
|
||||
const proxy = resolveProxy(providerId);
|
||||
const { model, modelInfo } = getModel({
|
||||
@@ -68,9 +101,20 @@ export async function resolveModel(params: {
|
||||
apiKey,
|
||||
baseUrl,
|
||||
proxy,
|
||||
providerType: params.providerType as 'openai' | 'anthropic' | 'google' | undefined,
|
||||
providerType: clientProviderType as 'openai' | 'anthropic' | 'google' | undefined,
|
||||
});
|
||||
|
||||
// Thinking arbitration mirrors model routing — the route carries a full
|
||||
// ThinkingConfig (mode/effort/level/enabled/budgetTokens/…) which callLLM
|
||||
// normalizes against the model's capability:
|
||||
// - routed + thinking set → the route's thinking wins (over client thinking).
|
||||
// - routed + no thinking → routed model uses its own default; client thinking
|
||||
// is dropped (it belonged to the client's other model).
|
||||
// - unrouted → honor the client's thinking config.
|
||||
const thinkingConfig: ThinkingConfig | undefined = routed
|
||||
? stageRoute?.thinking
|
||||
: params.thinkingConfig;
|
||||
|
||||
return {
|
||||
model,
|
||||
modelInfo,
|
||||
@@ -79,7 +123,7 @@ export async function resolveModel(params: {
|
||||
modelId,
|
||||
apiKey,
|
||||
baseUrl,
|
||||
thinkingConfig: params.thinkingConfig,
|
||||
thinkingConfig,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -97,12 +141,18 @@ function getThinkingConfigFromBody(body: unknown): ThinkingConfig | undefined {
|
||||
* Note: requiresApiKey is derived server-side from the provider registry,
|
||||
* never from client headers, to prevent auth bypass.
|
||||
*/
|
||||
export async function resolveModelFromHeaders(req: NextRequest): Promise<ResolvedModel> {
|
||||
export async function resolveModelFromHeaders(
|
||||
req: NextRequest,
|
||||
stage?: LlmStage,
|
||||
thinkingConfig?: ThinkingConfig,
|
||||
): Promise<ResolvedModel> {
|
||||
return resolveModel({
|
||||
modelString: req.headers.get('x-model') || undefined,
|
||||
stage,
|
||||
apiKey: req.headers.get('x-api-key') || undefined,
|
||||
baseUrl: req.headers.get('x-base-url') || undefined,
|
||||
providerType: req.headers.get('x-provider-type') || undefined,
|
||||
thinkingConfig,
|
||||
});
|
||||
}
|
||||
|
||||
@@ -115,10 +165,9 @@ export async function resolveModelFromHeaders(req: NextRequest): Promise<Resolve
|
||||
export async function resolveModelFromRequest(
|
||||
req: NextRequest,
|
||||
body: unknown,
|
||||
stage?: LlmStage,
|
||||
): Promise<ResolvedModel> {
|
||||
const resolved = await resolveModelFromHeaders(req);
|
||||
return {
|
||||
...resolved,
|
||||
thinkingConfig: getThinkingConfigFromBody(body) ?? resolved.thinkingConfig,
|
||||
};
|
||||
// Pass the client's body thinking into resolveModel so the single arbiter
|
||||
// there decides (a routed stage may override or drop it). See resolveModel.
|
||||
return resolveModelFromHeaders(req, stage, getThinkingConfigFromBody(body));
|
||||
}
|
||||
|
||||
@@ -0,0 +1,216 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
|
||||
|
||||
// model-routes reads process.env.MODEL_ROUTES once and caches the parsed map.
|
||||
// Tests reset the module registry between cases via vi.resetModules() so each
|
||||
// case re-reads a fresh env, mirroring the provider-config test convention.
|
||||
|
||||
describe('model-routes', () => {
|
||||
beforeEach(() => {
|
||||
vi.resetModules();
|
||||
delete process.env.MODEL_ROUTES;
|
||||
vi.restoreAllMocks();
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
delete process.env.MODEL_ROUTES;
|
||||
});
|
||||
|
||||
it('returns undefined for any stage when MODEL_ROUTES is unset', async () => {
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content')).toBeUndefined();
|
||||
expect(getStageModel('pbl-chat')).toBeUndefined();
|
||||
});
|
||||
|
||||
it('returns undefined when no stage is provided', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel(undefined)).toBeUndefined();
|
||||
});
|
||||
|
||||
it('returns the mapped model for a configured routable stage', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': 'openai:gpt-5.4',
|
||||
'pbl-chat': 'anthropic:claude-sonnet-4',
|
||||
});
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content')).toBe('openai:gpt-5.4');
|
||||
expect(getStageModel('pbl-chat')).toBe('anthropic:claude-sonnet-4');
|
||||
});
|
||||
|
||||
it('returns undefined for a routable stage that is not listed', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-actions')).toBeUndefined();
|
||||
});
|
||||
|
||||
it('ignores unknown stage keys with a warning but keeps valid ones', async () => {
|
||||
const warn = vi.fn();
|
||||
vi.doMock('@/lib/logger', () => ({
|
||||
createLogger: () => ({ warn, info: vi.fn(), error: vi.fn(), debug: vi.fn() }),
|
||||
}));
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'not-a-stage': 'openai:gpt-5.4',
|
||||
'scene-content': 'openai:gpt-5.4',
|
||||
});
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('not-a-stage')).toBeUndefined();
|
||||
expect(getStageModel('scene-content')).toBe('openai:gpt-5.4');
|
||||
expect(warn).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('returns undefined for everything when MODEL_ROUTES is invalid JSON (no throw)', async () => {
|
||||
const error = vi.fn();
|
||||
vi.doMock('@/lib/logger', () => ({
|
||||
createLogger: () => ({ error, info: vi.fn(), warn: vi.fn(), debug: vi.fn() }),
|
||||
}));
|
||||
process.env.MODEL_ROUTES = '{not valid json';
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content')).toBeUndefined();
|
||||
expect(error).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('ignores non-string route values', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': 123,
|
||||
'pbl-chat': 'anthropic:claude-sonnet-4',
|
||||
});
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content')).toBeUndefined();
|
||||
expect(getStageModel('pbl-chat')).toBe('anthropic:claude-sonnet-4');
|
||||
});
|
||||
|
||||
it('resolves a composite scene-content:<type> key', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content:quiz': 'openai:gpt-5.4' });
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content:quiz')).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('falls back from a composite key to the base stage', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4-mini' });
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content:quiz')).toBe('openai:gpt-5.4-mini');
|
||||
});
|
||||
|
||||
it('prefers the composite key over the base stage', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': 'openai:gpt-5.4-mini',
|
||||
'scene-content:quiz': 'openai:gpt-5.4',
|
||||
});
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content:quiz')).toBe('openai:gpt-5.4');
|
||||
// a type without its own route falls back to the base scene-content route
|
||||
expect(getStageModel('scene-content:slide')).toBe('openai:gpt-5.4-mini');
|
||||
});
|
||||
|
||||
it('ignores an unknown scene-content:<type> key with a warning', async () => {
|
||||
const warn = vi.fn();
|
||||
vi.doMock('@/lib/logger', () => ({
|
||||
createLogger: () => ({ warn, info: vi.fn(), error: vi.fn(), debug: vi.fn() }),
|
||||
}));
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content:bogus': 'openai:gpt-5.4' });
|
||||
const { getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageModel('scene-content:bogus')).toBeUndefined();
|
||||
expect(warn).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('parses an object route value {model, thinking} via getStageRoute', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': { model: 'openai:gpt-5.4', thinking: { effort: 'high' } },
|
||||
});
|
||||
const { getStageRoute, getStageModel } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('scene-content')).toEqual({
|
||||
model: 'openai:gpt-5.4',
|
||||
thinking: { effort: 'high' },
|
||||
});
|
||||
expect(getStageModel('scene-content')).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('supports the full thinking config (budgetTokens/enabled/level/mode)', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content:interactive': {
|
||||
model: 'qwen:qwen3.7-plus',
|
||||
thinking: { enabled: true, budgetTokens: 8000 },
|
||||
},
|
||||
'scene-content:slide': { model: 'google:gemini-3-flash-preview', thinking: { level: 'low' } },
|
||||
'scene-content:quiz': {
|
||||
model: 'deepseek:deepseek-v4-pro',
|
||||
thinking: { mode: 'disabled', enabled: false },
|
||||
},
|
||||
});
|
||||
const { getStageRoute } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('scene-content:interactive')!.thinking).toEqual({
|
||||
enabled: true,
|
||||
budgetTokens: 8000,
|
||||
});
|
||||
expect(getStageRoute('scene-content:slide')!.thinking).toEqual({ level: 'low' });
|
||||
expect(getStageRoute('scene-content:quiz')!.thinking).toEqual({
|
||||
mode: 'disabled',
|
||||
enabled: false,
|
||||
});
|
||||
});
|
||||
|
||||
it('treats a string route value as a model with no thinking', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'pbl-chat': 'anthropic:claude-sonnet-4' });
|
||||
const { getStageRoute } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('pbl-chat')).toEqual({ model: 'anthropic:claude-sonnet-4' });
|
||||
});
|
||||
|
||||
it('drops invalid thinking fields with a warning but keeps the model', async () => {
|
||||
const warn = vi.fn();
|
||||
vi.doMock('@/lib/logger', () => ({
|
||||
createLogger: () => ({ warn, info: vi.fn(), error: vi.fn(), debug: vi.fn() }),
|
||||
}));
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': {
|
||||
model: 'openai:gpt-5.4',
|
||||
thinking: { effort: 'bogus', budgetTokens: 'x' },
|
||||
},
|
||||
});
|
||||
const { getStageRoute } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('scene-content')).toEqual({ model: 'openai:gpt-5.4' });
|
||||
expect(warn).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('ignores an object route value with no model string', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': { thinking: { effort: 'high' } },
|
||||
});
|
||||
const { getStageRoute } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('scene-content')).toBeUndefined();
|
||||
});
|
||||
|
||||
it('getStageRoute returns the matched composite entry (model+thinking) as a unit', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': { model: 'openai:gpt-5.4-mini' },
|
||||
'scene-content:quiz': { model: 'openai:gpt-5.4', thinking: { effort: 'high' } },
|
||||
});
|
||||
const { getStageRoute } = await import('@/lib/server/model-routes');
|
||||
expect(getStageRoute('scene-content:quiz')).toEqual({
|
||||
model: 'openai:gpt-5.4',
|
||||
thinking: { effort: 'high' },
|
||||
});
|
||||
// unrouted type falls back to the base entry (no thinking there)
|
||||
expect(getStageRoute('scene-content:slide')).toEqual({ model: 'openai:gpt-5.4-mini' });
|
||||
});
|
||||
|
||||
it('exposes the routable stage registry', async () => {
|
||||
const { LLM_STAGES } = await import('@/lib/server/model-routes');
|
||||
expect(LLM_STAGES).toEqual(
|
||||
expect.arrayContaining([
|
||||
'scene-content:slide',
|
||||
'scene-content:quiz',
|
||||
'scene-content:interactive',
|
||||
'scene-content:pbl',
|
||||
'scene-outlines-stream',
|
||||
'scene-content',
|
||||
'scene-actions',
|
||||
'agent-profiles',
|
||||
'quiz-grade',
|
||||
'pbl-chat',
|
||||
'chat-adapter',
|
||||
'generate-classroom',
|
||||
'web-search-query-rewrite',
|
||||
]),
|
||||
);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,207 @@
|
||||
import { beforeEach, describe, expect, it, vi } from 'vitest';
|
||||
|
||||
// Mock the heavy downstream of resolveModel so the test isolates the model
|
||||
// string *resolution order*: stage route > x-model > DEFAULT_MODEL > builtin.
|
||||
// model-routes is left real (it just reads MODEL_ROUTES) so we exercise the
|
||||
// real integration point.
|
||||
// Use the real parseModelString (canonical `provider:model` colon format) so
|
||||
// the test exercises actual separator handling; only stub getModel (recording
|
||||
// its args) so no real provider client is constructed. provider-config stubs
|
||||
// echo the client-supplied key/baseUrl so a test can assert they are dropped
|
||||
// when a stage route overrides the client model.
|
||||
const mocks = vi.hoisted(() => ({ getModelCalls: [] as Array<Record<string, unknown>> }));
|
||||
|
||||
vi.mock('@/lib/ai/providers', async (importOriginal) => {
|
||||
const actual = await importOriginal<typeof import('@/lib/ai/providers')>();
|
||||
return {
|
||||
...actual,
|
||||
getModel: (args: Record<string, unknown>) => {
|
||||
mocks.getModelCalls.push(args);
|
||||
return { model: { id: args.modelId }, modelInfo: undefined };
|
||||
},
|
||||
};
|
||||
});
|
||||
|
||||
vi.mock('@/lib/server/provider-config', () => ({
|
||||
isServerConfiguredProvider: () => false,
|
||||
resolveApiKey: (_id: string, clientKey: string) => clientKey || 'server-key',
|
||||
resolveBaseUrl: (_id: string, clientBaseUrl?: string) => clientBaseUrl,
|
||||
resolveProxy: () => undefined,
|
||||
}));
|
||||
|
||||
vi.mock('@/lib/server/ssrf-guard', () => ({
|
||||
validateUrlForSSRF: async () => null,
|
||||
}));
|
||||
|
||||
describe('resolveModel — per-stage resolution order', () => {
|
||||
beforeEach(() => {
|
||||
vi.resetModules();
|
||||
mocks.getModelCalls.length = 0;
|
||||
delete process.env.MODEL_ROUTES;
|
||||
delete process.env.DEFAULT_MODEL;
|
||||
});
|
||||
|
||||
it('throws (no hardcoded fallback) when nothing is configured', async () => {
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
await expect(resolveModel({ stage: 'scene-content' })).rejects.toThrow(
|
||||
/No model could be resolved/,
|
||||
);
|
||||
});
|
||||
|
||||
it('uses DEFAULT_MODEL when no stage route matches', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'scene-content' });
|
||||
expect(r.modelString).toBe('openai:gpt-5.4-mini');
|
||||
});
|
||||
|
||||
it('uses the stage route over DEFAULT_MODEL', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'scene-content' });
|
||||
expect(r.modelString).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('uses DEFAULT_MODEL for stages not listed in MODEL_ROUTES', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'quiz-grade' });
|
||||
expect(r.modelString).toBe('openai:gpt-5.4-mini');
|
||||
});
|
||||
|
||||
it('lets a configured stage route win over an explicit modelString (x-model)', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({
|
||||
stage: 'scene-content',
|
||||
modelString: 'anthropic:claude-sonnet-4',
|
||||
});
|
||||
expect(r.modelString).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('falls back to x-model for a stage that is not routed', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'quiz-grade', modelString: 'anthropic:claude-sonnet-4' });
|
||||
expect(r.modelString).toBe('anthropic:claude-sonnet-4');
|
||||
});
|
||||
|
||||
it('drops client apiKey/baseUrl/providerType when a stage route overrides the client model', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'pbl-chat': 'anthropic:claude-sonnet-4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
await resolveModel({
|
||||
stage: 'pbl-chat',
|
||||
modelString: 'openai:gpt-5.4-mini',
|
||||
apiKey: 'client-openai-key',
|
||||
baseUrl: 'https://client.example/v1',
|
||||
providerType: 'openai',
|
||||
});
|
||||
const call = mocks.getModelCalls.at(-1)!;
|
||||
expect(call.providerId).toBe('anthropic');
|
||||
expect(call.modelId).toBe('claude-sonnet-4');
|
||||
// None of the client-sent connection params for the OLD provider leak onto
|
||||
// the routed provider — they resolve from server config instead.
|
||||
expect(call.providerType).toBeUndefined();
|
||||
expect(call.baseUrl).toBeUndefined();
|
||||
expect(call.apiKey).toBe('server-key');
|
||||
});
|
||||
|
||||
it('keeps client apiKey/baseUrl/providerType for an unrouted stage (x-model honored)', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
await resolveModel({
|
||||
stage: 'quiz-grade',
|
||||
modelString: 'openai:gpt-5.4-mini',
|
||||
apiKey: 'client-key',
|
||||
baseUrl: 'https://client.example/v1',
|
||||
providerType: 'openai',
|
||||
});
|
||||
const call = mocks.getModelCalls.at(-1)!;
|
||||
expect(call.providerType).toBe('openai');
|
||||
expect(call.baseUrl).toBe('https://client.example/v1');
|
||||
expect(call.apiKey).toBe('client-key');
|
||||
});
|
||||
|
||||
it('uses a scene-content:<type> route over the base route and x-model', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content': 'openai:gpt-5.4-mini',
|
||||
'scene-content:quiz': 'openai:gpt-5.4',
|
||||
});
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({
|
||||
stage: 'scene-content:quiz',
|
||||
modelString: 'anthropic:claude-sonnet-4',
|
||||
});
|
||||
expect(r.modelString).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('falls back to the base scene-content route for an unrouted type', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'scene-content:slide' });
|
||||
expect(r.modelString).toBe('openai:gpt-5.4');
|
||||
});
|
||||
|
||||
it('resolves the stage route provider for cross-provider routing', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'pbl-chat': 'anthropic:claude-sonnet-4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'pbl-chat' });
|
||||
expect(r.modelString).toBe('anthropic:claude-sonnet-4');
|
||||
expect(r.providerId).toBe('anthropic');
|
||||
expect(r.modelId).toBe('claude-sonnet-4');
|
||||
});
|
||||
|
||||
it('route thinking wins over client thinking when the stage is routed', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'pbl-chat': { model: 'anthropic:claude-sonnet-4', thinking: { effort: 'high' } },
|
||||
});
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'pbl-chat', thinkingConfig: { effort: 'low' } });
|
||||
expect(r.thinkingConfig).toEqual({ effort: 'high' });
|
||||
});
|
||||
|
||||
it('route can pass a full thinking config (enabled + budgetTokens)', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({
|
||||
'scene-content:interactive': {
|
||||
model: 'qwen:qwen3.7-plus',
|
||||
thinking: { enabled: true, budgetTokens: 8000 },
|
||||
},
|
||||
});
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'scene-content:interactive' });
|
||||
expect(r.thinkingConfig).toEqual({ enabled: true, budgetTokens: 8000 });
|
||||
});
|
||||
|
||||
it('routed-without-thinking drops client thinking (routed model uses its default)', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'deepseek:deepseek-v4-pro' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({ stage: 'scene-content', thinkingConfig: { effort: 'high' } });
|
||||
expect(r.thinkingConfig).toBeUndefined();
|
||||
});
|
||||
|
||||
it('unrouted stage keeps the client thinking config', async () => {
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'deepseek:deepseek-v4-pro' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({
|
||||
stage: 'quiz-grade',
|
||||
modelString: 'openai:gpt-5.4-mini',
|
||||
thinkingConfig: { effort: 'medium' },
|
||||
});
|
||||
expect(r.thinkingConfig).toEqual({ effort: 'medium' });
|
||||
});
|
||||
|
||||
it('ignores stage routing entirely when no stage is passed', async () => {
|
||||
process.env.DEFAULT_MODEL = 'openai:gpt-5.4-mini';
|
||||
process.env.MODEL_ROUTES = JSON.stringify({ 'scene-content': 'openai:gpt-5.4' });
|
||||
const { resolveModel } = await import('@/lib/server/resolve-model');
|
||||
const r = await resolveModel({});
|
||||
expect(r.modelString).toBe('openai:gpt-5.4-mini');
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user