Files
AkitaOnRailsandClaude Opus 4.7 df7671dd56 evals: live A/B harness comparing two LLM providers on the consolidation prompt
A 5-fixture side-by-side runner so we can validate "does Ollama
qwen3:32b degrade quality vs OpenRouter kimi-k2.6" empirically
instead of by gut feeling.

## What it does

  cargo run -p ai-memory-eval -- \
    --baseline-provider  openai-compat --baseline-model  ... \
    --candidate-provider openai-compat --candidate-model ...

For each `evals/fixtures/*.json`:
  1. Build the EXACT ChatRequest the live consolidator builds via
     the now-public `ai_memory_consolidate::build_batch_request`.
  2. Send it concurrently to both providers through
     `ai_memory_llm::complete_structured` (same schema-validate
     path the production system uses).
  3. Persist `<fixture>.{json, md, meta.json}` under
     `evals/runs/<timestamp>/{baseline, candidate}/`.
  4. Print per-fixture latency + parse status, then a tail summary.

Quality judgement is left to the human reading the markdown
outputs side-by-side — the runner only reports objective deltas
(latency, parse_ok, update_count). Anything subtler (faithfulness,
hallucination, scoping) is for eyeball review.

## Five fixtures

  01-rust-bug-fix         — extracts both decision + gotcha pages
  02-architecture-decision — ADR-style page distinct from session
  03-gotcha-with-rule      — rule-classification + auto-routing
  04-low-signal-session    — RESIST manufacturing pages
  05-multi-topic-session   — separate pages per topic

Each fixture exercises a different hard case in the
`BATCH_SYSTEM_PROMPT` design space.

## Layout

  evals/
    Cargo.toml          publish=false, workspace member
    src/main.rs         clap + tokio runner
    fixtures/*.json     5 synthetic session logs
    runs/               gitignored output dir
    README.md           how to read + how to add fixtures

## Out of scope (deliberately)

  - Automatic quality scoring. Side-by-side only.
  - Embedding A/B (separate pipeline, not added here).
  - Persisting via the real wiki layer — pure prompt→response.

The harness is a workspace member purely so it shares deps and
compiles with the rest of the workspace; it's never bundled in
the docker image and never run by CI.

Two small `consolidator.rs` changes to support this:
  - `build_batch_request` is now pub (was fn-level private)
  - `BATCH_SYSTEM_PROMPT` is now pub (was const-level private)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-22 15:46:24 -03:00
..