mirror of
https://github.com/nashsu/llm_wiki.git
synced 2026-10-02 02:44:34 +08:00
The Ollama provider built its OpenAI-compatible /v1/chat/completions body
without translating the caller's `reasoning: { mode }` onto any wire field.
Unlike the DeepSeek/Qwen/MiMo branches, `mode: "off"` was silently dropped —
so a thinking-capable Ollama model (e.g. gemma4:12b, whose capabilities
include "thinking") spent its whole token budget on chain-of-thought and ended
the stream with empty `content`. That surfaced two ways:
- large ingest prompts: "produced N chars of reasoning, but no actual
response content" (>200-char threshold), and
- the connection test (max_tokens: 32): the model burned all 32 tokens
thinking (finish_reason=length, 0 content) -> "Model connected but
returned empty content".
Map reasoning mode onto Ollama's `reasoning_effort` ("high"|"medium"|"low"|
"none"; "none" disables thinking), per docs.ollama.com/api/openai-compatibility.
"max" has no Ollama analogue, so it maps to "high"; "auto"/"custom" leave the
field unset for the model default. Verified on a live gemma4:12b: with
reasoning_effort:"none" it returns clean content; without it, 0 content +
finish_reason=length.
Also add scripts/debug_ollama_tokens.py — a stdlib-only probe that replays the
app's wiki-generation request against a raw Ollama endpoint (model context
inspection, prompt-size sweep, OpenAI-compat vs native num_ctx comparison),
used to confirm the "too many tokens" report was a downstream symptom of the
reasoning runaway, not a real context-window overflow.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>