Files
小魔andwyuc a41bbac239 feat(llm): retry once on a fallback model for retryable failures (#1614)
* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* Delete omai-upload-tree directory

* fix: ensure trailing newline in locale files

fix: ensure trailing newline in locale files

* chore: english comments for upstream review

chore: english comments for upstream review

* Delete components/settings/model-schedule-settings.tsx

* Delete lib/server/model-schedule.ts

* Delete app/api/model-schedule/route.ts

* Delete assets/model-schedule/ledger-with-entry.png

* Delete assets/model-schedule/settings-panel.png

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Update .env.example

* Enhance access control comments in .env.example

Added additional context and warnings for access control configuration.

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* run ci

* Add files via upload

* Add files via upload

* Add files via upload

* Supplement. env.example

Supplement. env.example

* Add files via upload

* Update .env.example

* Update fallback notes in .env.example

Clarified fallback behavior in callLLM layer documentation.

* Update llm.ts

* Update llm-fallback.test.ts

* fix(llm): round-4 fallback hardening

fix(llm): round-4 fallback hardening — server-managed gate, content-filter refusal, APICallError unwrap stop, outlines fullStream errors

* fix(llm): rebase round-4 onto main v1.1.1

fix(llm): rebase round-4 onto main v1.1.1

* fix(llm): round4c-prettier formatting+retry test 

fix(llm): round4c-prettier formatting+retry test

* test(llm): round-4d-align outline-stream

test(llm): align outline-stream and title-generator tests with round-4 call signatures

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-28 17:24:55 +08:00
..