mirror of
https://github.com/debpalash/VoiceStudio.git
synced 2026-10-02 01:26:35 +08:00
Some chat templates write the opening think tag into the prompt
(Spark-X2.5, Qwen3 thinking variants, DeepSeek-R1-0528 served without a
reasoning parser), so the model streams only …SAY:. ReplyParser had no
opening tag to match, settled on plain mode at the first non-format
characters, and spoke every reasoning sentence — to the person on the
phone, in the user's voice — plus the stray tags.
Implements the reporter's recommended parser-side fix (FenjuFu,
option 1):
- _body now drops everything up to a closing tag found anywhere, or up
to a line-start SAY: when no closing tag comes either. Both searches
are stable on the append-only raw text, so _emitted stays aligned.
- _decide_mode no longer settles plain early: the system prompt
requires replies to start with SAY: or {, so a body with neither is
held while it streams (it may be prefilled thinking) and only
finish() speaks it as plain text. Format-following replies keep
first-sentence streaming.
Regression tests fail before the parser change; the reporter's exact
repro now returns only the SAY: text in both variants.