mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
A 5-fixture side-by-side runner so we can validate "does Ollama
qwen3:32b degrade quality vs OpenRouter kimi-k2.6" empirically
instead of by gut feeling.
## What it does
cargo run -p ai-memory-eval -- \
--baseline-provider openai-compat --baseline-model ... \
--candidate-provider openai-compat --candidate-model ...
For each `evals/fixtures/*.json`:
1. Build the EXACT ChatRequest the live consolidator builds via
the now-public `ai_memory_consolidate::build_batch_request`.
2. Send it concurrently to both providers through
`ai_memory_llm::complete_structured` (same schema-validate
path the production system uses).
3. Persist `<fixture>.{json, md, meta.json}` under
`evals/runs/<timestamp>/{baseline, candidate}/`.
4. Print per-fixture latency + parse status, then a tail summary.
Quality judgement is left to the human reading the markdown
outputs side-by-side — the runner only reports objective deltas
(latency, parse_ok, update_count). Anything subtler (faithfulness,
hallucination, scoping) is for eyeball review.
## Five fixtures
01-rust-bug-fix — extracts both decision + gotcha pages
02-architecture-decision — ADR-style page distinct from session
03-gotcha-with-rule — rule-classification + auto-routing
04-low-signal-session — RESIST manufacturing pages
05-multi-topic-session — separate pages per topic
Each fixture exercises a different hard case in the
`BATCH_SYSTEM_PROMPT` design space.
## Layout
evals/
Cargo.toml publish=false, workspace member
src/main.rs clap + tokio runner
fixtures/*.json 5 synthetic session logs
runs/ gitignored output dir
README.md how to read + how to add fixtures
## Out of scope (deliberately)
- Automatic quality scoring. Side-by-side only.
- Embedding A/B (separate pipeline, not added here).
- Persisting via the real wiki layer — pure prompt→response.
The harness is a workspace member purely so it shares deps and
compiles with the rest of the workspace; it's never bundled in
the docker image and never run by CI.
Two small `consolidator.rs` changes to support this:
- `build_batch_request` is now pub (was fn-level private)
- `BATCH_SYSTEM_PROMPT` is now pub (was const-level private)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>