mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
Extends the in-repo retrieval eval into an A/B evaluation that runs the same question set through a baseline and candidate config and reports a delta over the triple: - accuracy: hit@k / recall@k (existing); - latency: per-query timing, p50/p95; - context-tokens: sum of returned hit title+snippet, chars/4 estimate, mean/median. Config knobs: --embeddings none|local, --reranker (now configurable, was force-removed), --server-env, and --query-arg passthrough into memory_query (ready for pin_first/include_superseded/etc.), each with a --candidate-* twin. Delta report in markdown+JSON preserving provenance (commit, dataset sha, hw). Single-config path unchanged. Also fixes a latent harness deadlock (child stderr pipe now drained on a background thread) that surfaced with two servers. Dev tooling only (evals is publish=false); numbers land in docs/benchmarks. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm