Files
llm_wiki/scripts
Kostadis RousossandClaude Opus 4.8 a543fc54f7 fix: disable thinking on Ollama path so reasoning models can answer
The Ollama provider built its OpenAI-compatible /v1/chat/completions body
without translating the caller's `reasoning: { mode }` onto any wire field.
Unlike the DeepSeek/Qwen/MiMo branches, `mode: "off"` was silently dropped —
so a thinking-capable Ollama model (e.g. gemma4:12b, whose capabilities
include "thinking") spent its whole token budget on chain-of-thought and ended
the stream with empty `content`. That surfaced two ways:

  - large ingest prompts: "produced N chars of reasoning, but no actual
    response content" (>200-char threshold), and
  - the connection test (max_tokens: 32): the model burned all 32 tokens
    thinking (finish_reason=length, 0 content) -> "Model connected but
    returned empty content".

Map reasoning mode onto Ollama's `reasoning_effort` ("high"|"medium"|"low"|
"none"; "none" disables thinking), per docs.ollama.com/api/openai-compatibility.
"max" has no Ollama analogue, so it maps to "high"; "auto"/"custom" leave the
field unset for the model default. Verified on a live gemma4:12b: with
reasoning_effort:"none" it returns clean content; without it, 0 content +
finish_reason=length.

Also add scripts/debug_ollama_tokens.py — a stdlib-only probe that replays the
app's wiki-generation request against a raw Ollama endpoint (model context
inspection, prompt-size sweep, OpenAI-compat vs native num_ctx comparison),
used to confirm the "too many tokens" report was a downstream symptom of the
reasoning runaway, not a real context-window overflow.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 10:17:20 -07:00
..