Files
teamai-cli/agents
3bc148863f fix(search): deduplicate recall results by stable content identity (#298)
The dedup key included `score` and `tokens.length`. Neither identifies a
document, so results were collapsed in both directions.

Identical content survived deduplication: `score` carries the vote bonus
(+0.5/vote), so two copies of one re-shared learning drift apart as they
collect votes independently and both consume a `limit` slot.

Distinct content was collapsed: `tokens.length` only counts tokens, so two
entries sharing type, title, date, author, token count and score were merged
and one was silently dropped.

Key on the fields a re-share copies verbatim instead — type, domain, title,
date, author and the full token set. Tags and body reach the key through the
token set, already normalized. The tokens are sorted because tag order follows
hand-authored frontmatter, which a re-share may reorder. `domain` is listed
separately: it comes from frontmatter, never reaches the token set, and was
previously kept apart only by the domain multiplier inside `score`.

The index schema and public CLI behavior are unchanged.

Co-authored-by: Oreo9 <x9276@qq.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 18:37:53 +08:00
..