Files
ai-memory/crates/ai-memory-store/tests/suite/access_breadth.rs
T
gb 3efb5b0b89 perf(dev): a self-contained build and a two-tier test loop on every platform
The edit-to-result loop was ~380s for the workspace on macOS and needed an
environment variable on every command. This makes `cargo t` the whole story
on macOS, Linux, and Windows, with numbers measured along the way.

Build
- `[profile.dev]` keeps only line tables (full debuginfo put ~190 MB of
  DWARF in each test binary and made the build linker-bound); dependencies
  build at opt-level 1 with no debuginfo; proc macros and build scripts at
  opt-level 3, since they are run once per dependent crate.
- Test binaries: 78 to 11 in the everyday loop (13 under `--workspace`).
  Each one is a link and, on macOS (Gatekeeper) and Windows (Defender), a
  first-run malware scan of the whole file, paid serially by nextest's list
  phase before the first test starts. Integration tests now live in
  `tests/suite/` and compile into their crate's own test harness (`mod.rs`,
  included from `src/lib.rs` under `#[cfg(test)]`, with `extern crate self`
  so they keep addressing the public API by crate name). Only the CLI keeps
  a separate `suite` target, because its tests run the built executable.
  The evals harness leaves `default-members`, so a bare `cargo t` skips its
  two binaries while `--workspace` (CI, the hook, `cargo tf`) still builds
  them. A repo-layout test fails on an undeclared suite file, a stray
  top-level `tests/*.rs`, or a `mod.rs` that `lib.rs` never includes.
- `ai-memory-cli` gains a lib target; `main.rs` is a shim. 806 tests that
  lived in the bin are reachable, and `--lib` runs skip the 127 MB binary.
- The web crate's vendored `static/tailwind.css` is the default on every
  build, so nothing needs `TAILWIND_SKIP=1` any more: every release, Docker,
  and CI path already used the vendored file, and the download branch only
  ever ran for developers who forgot the flag (and then rewrote the source
  tree as a side effect). `TAILWIND_BUILD=1 cargo build -p ai-memory-web`
  regenerates it explicitly. CI runs that on Linux and fails if the
  committed file is stale, a check that did not exist before; the committed
  file reproduces byte for byte today.
- `tokenizers` aligned on one version instead of the 0.21 pin plus the 0.22
  candle pulled in.

Test tiers
- `.config/nextest.toml`: the `default` profile skips any test whose module
  path has a segment starting with `slow` or `stress` (`packaging::slow::*`
  drives real wrapper scripts and fake container engines at 10-20s each;
  `stress_*` modules hammer concurrency), reports every failure in one run,
  and marks anything over 5s in its summary so a new slow test is visible
  the day it lands. `full` runs everything. `ci` keeps its retries and
  writes JUnit.
- `.cargo/config.toml` holds two aliases and nothing else: `cargo t` (default
  members) and `cargo tf` (`--workspace -P full`). `cargo t -p <crate>`
  builds just that crate. Neither passes `--all-targets`: there are no
  examples or benches, and it only added harnesses for two `test = false`
  targets.
- `scripts/install-git-hooks.sh` installs an opt-in pre-push hook that runs
  the full tier, touching only its own marked block. Two independent things
  run the skipped tier: that hook, and CI, which uses `cargo test` and never
  reads the nextest config.

Slow tests fixed rather than tiered
- `project_observations` in the consolidator trimmed an over-budget
  projection one observation at a time, re-rendering the whole text and
  re-scoring every remaining candidate after each removal. Each score scans
  the body, so 256 observations of 4k chars cost ~65k body scans per prompt:
  14s in production consolidation, exactly as in the unit test. Scores and
  per-block sizes are now computed once and the prune subtracts; output is
  unchanged and pinned by the existing tests. 13.9s to 0.18s.
- Windows takes ~2s to refuse a loopback connect, so every hook test that
  posted to a closed port paid 2s per request. `dead_http_endpoint()` in the
  new `ai-memory-test-support` crate accepts and closes instead, with a
  fallback to the closed port where binding is denied. devin hook tests:
  4.2s to 0.15s each.
- The store unit fixture opened a file-backed SQLite with the default
  rollback journal and synchronous=FULL, so ~120 parallel fixtures fsynced
  every transaction. journal_mode=MEMORY + synchronous=OFF: 242s to 89s of
  test time, p90 1.6s to 0.5s.
- Windows-only tests resolve `powershell.exe` or `pwsh.exe` once per process
  and the auto-improve eval fixtures are `.ps1` scripts instead of cmd.exe
  batch files; a post-bind settle sleep is gone; the two unpinned
  multi-thread tokio tests pin `worker_threads = 4`. The four copies of the
  PowerShell resolver and the mcp suite's duplicated `post`/`get` helpers are
  now one each.

Not done, with the numbers in AGENTS.md: nextest vs in-process libtest is a
wash per crate and a rout for the workspace (20s vs 309s); the
`local-embeddings` default feature costs ~50s of cold build and ~27 MB per
binary but under a second per relink, so it stays a product default.

Measured: workspace loop ~380s to ~150s on macOS; on a 32-thread Windows box
the warm everyday run is 20s of test time across 2919 tests in 11 binaries,
and the rebuild after a core edit is 13s of cargo with lld plus the
first-run scans.
2026-09-05 08:27:08 -04:00

277 lines
9.5 KiB
Rust

//! Per-operator reinforcement, recorded alongside the shared counter.
//!
//! `pages.access_count` cannot distinguish "50 reads by one person" from "one
//! read by each of 50 people", although only the second says a page is
//! load-bearing for a team. `page_access` records the breakdown WITHOUT
//! replacing the scalar, so the retention formula, the hard-delete predicate
//! and every existing query keep reading exactly what they read before.
use ai_memory_core::{IdentityKey, NewPage, PagePath, Tier};
use ai_memory_store::{DecayParams, Store, retention_score, retention_score_with_breadth};
use rusqlite::{Connection, params};
/// The qualified TEXT the read path records operators under — built through
/// the contract (`IdentityKey::storage_key()`), never hand-written, so these
/// tests break if the storage encoding ever drifts from the API.
fn actor(name: &str) -> IdentityKey {
IdentityKey::User(name.into())
}
/// Default parameters must reproduce the historical score exactly, whatever the
/// breadth — otherwise adopting the table would silently move every eviction
/// decision on every existing database.
#[test]
fn breadth_is_identity_at_the_default_weight() {
let params = DecayParams::default();
let breadth_weight = 0.0;
for actors in [0, 1, 2, 10, 500] {
for (age, count, since) in [
(0.0, 0, None),
(10.0, 3, Some(2.0)),
(365.0, 100, Some(200.0)),
] {
assert_eq!(
retention_score_with_breadth(
&params,
age,
count,
since,
None,
actors,
breadth_weight,
),
retention_score(&params, age, count, since, None),
"default weight must be identity (actors={actors})"
);
}
}
}
/// Even with the weight turned up, 0 and 1 distinct actors score identically to
/// the old formula: a page nobody has read per-actor rows for (everything
/// written before this existed) and a page one person reads are unchanged. That
/// is what removes the eviction cliff and the need for any backfill.
#[test]
fn zero_and_one_actor_score_identically_even_when_weighted() {
let params = DecayParams::default();
let breadth_weight = 1.5;
let baseline = retention_score(&params, 30.0, 5, Some(3.0), None);
for actors in [0, 1] {
assert_eq!(
retention_score_with_breadth(&params, 30.0, 5, Some(3.0), None, actors, breadth_weight,),
baseline,
"actors={actors} must not change the score"
);
}
// More readers is worth strictly more, and monotonically so.
let two = retention_score_with_breadth(&params, 30.0, 5, Some(3.0), None, 2, breadth_weight);
let ten = retention_score_with_breadth(&params, 30.0, 5, Some(3.0), None, 10, breadth_weight);
assert!(two > baseline);
assert!(ten > two);
}
/// A page never accessed scores the same regardless of breadth: with no access
/// timestamp there is no access term to weight.
#[test]
fn never_accessed_pages_are_unaffected_by_breadth() {
let params = DecayParams::default();
assert_eq!(
retention_score_with_breadth(&params, 10.0, 0, None, None, 9, 2.0),
retention_score(&params, 10.0, 0, None, None)
);
}
#[test]
fn invalid_breadth_weights_fail_closed_to_the_historical_score() {
let params = DecayParams::default();
let baseline = retention_score(&params, 30.0, 5, Some(3.0), None);
for weight in [-1.0, f64::NAN, f64::INFINITY] {
assert_eq!(
retention_score_with_breadth(&params, 30.0, 5, Some(3.0), None, 50, weight),
baseline,
);
}
}
/// The breakdown is written in the same transaction as the scalar, so the two
/// cannot drift, and an unattributed read still bumps the scalar.
#[tokio::test]
async fn per_actor_rows_accumulate_without_replacing_the_scalar() {
let tmp = tempfile::tempdir().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default".to_string())
.await
.unwrap();
let proj = store
.writer
.get_or_create_project(ws, "app".to_string(), None)
.await
.unwrap();
let page = store
.writer
.upsert_page(NewPage {
workspace_id: ws,
project_id: proj,
path: PagePath::new("notes/x.md").unwrap(),
title: "x".into(),
body: "b".into(),
tier: Tier::Semantic,
frontmatter_json: serde_json::json!({}),
pinned: false,
links: Vec::new(),
author_id: None,
expires_at: None,
entities: Vec::new(),
})
.await
.unwrap();
store
.writer
.bump_access_for_actor(vec![page], Some(actor("alice")))
.await
.unwrap();
store
.writer
.bump_access_for_actor(vec![page], Some(actor("alice")))
.await
.unwrap();
store
.writer
.bump_access_for_actor(vec![page], Some(actor("bob")))
.await
.unwrap();
// An unattributed read: still counted in the shared scalar.
store.writer.bump_access(vec![page]).await.unwrap();
let error = store
.writer
.bump_access_for_actor(vec![page], Some(IdentityKey::User(" ".into())))
.await
.expect_err("a directly constructed blank identity must be rejected");
assert!(error.to_string().contains("normalized identity"));
let conn = Connection::open(store.db_path()).unwrap();
let scalar: i64 = conn
.query_row(
"SELECT access_count FROM pages WHERE id = ?1",
params![page.as_bytes()],
|r| r.get(0),
)
.unwrap();
assert_eq!(
scalar, 4,
"the historical counter counts valid reads, not rejected identities"
);
let distinct: i64 = conn
.query_row(
"SELECT COUNT(*) FROM page_access WHERE page_id = ?1",
params![page.as_bytes()],
|r| r.get(0),
)
.unwrap();
assert_eq!(distinct, 2, "two named operators, the anonymous read aside");
let alice_exists: bool = conn
.query_row(
"SELECT EXISTS(SELECT 1 FROM page_access WHERE page_id = ?1 AND actor = ?2)",
params![page.as_bytes(), actor("alice").storage_key()],
|r| r.get(0),
)
.unwrap();
assert!(alice_exists);
}
/// The bump is fired from a detached task AFTER the search responded, so a page
/// can be deleted in the interval. `page_access.page_id` REFERENCES `pages(id)`
/// with foreign keys ON, so an unguarded insert for that id aborts the
/// transaction and every OTHER page in the same result set silently loses its
/// once-per-window reinforcement — which then makes those pages likelier to be
/// evicted by the next sweep.
#[tokio::test]
async fn a_stale_page_id_does_not_cost_the_rest_of_the_batch_its_bump() {
let tmp = tempfile::tempdir().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default".to_string())
.await
.unwrap();
let proj = store
.writer
.get_or_create_project(ws, "app".to_string(), None)
.await
.unwrap();
let mut ids = Vec::new();
for path in ["notes/live-a.md", "notes/doomed.md", "notes/live-b.md"] {
ids.push(
store
.writer
.upsert_page(NewPage {
workspace_id: ws,
project_id: proj,
path: PagePath::new(path).unwrap(),
title: path.into(),
body: "b".into(),
tier: Tier::Episodic,
frontmatter_json: serde_json::json!({}),
pinned: false,
links: Vec::new(),
author_id: None,
expires_at: None,
entities: Vec::new(),
})
.await
.unwrap(),
);
}
// The page vanishes between the search and the bump.
store
.writer
.delete_page(ws, proj, PagePath::new("notes/doomed.md").unwrap(), None)
.await
.unwrap();
store
.writer
.bump_access_for_actor(ids.clone(), Some(actor("alice")))
.await
.unwrap();
let conn = Connection::open(store.db_path()).unwrap();
for (label, id) in [("live-a", ids[0]), ("live-b", ids[2])] {
let scalar: i64 = conn
.query_row(
"SELECT access_count FROM pages WHERE id = ?1",
params![id.as_bytes()],
|r| r.get(0),
)
.unwrap();
assert_eq!(scalar, 1, "{label} lost its bump to the stale sibling");
let per_actor: bool = conn
.query_row(
"SELECT EXISTS(SELECT 1 FROM page_access WHERE page_id = ?1 AND actor = ?2)",
params![id.as_bytes(), actor("alice").storage_key()],
|r| r.get(0),
)
.unwrap();
assert!(per_actor, "{label} lost its per-operator row");
}
// The unknown id stays a no-op, exactly as it was before per-actor rows.
let orphaned: i64 = conn
.query_row(
"SELECT COUNT(*) FROM page_access WHERE page_id = ?1",
params![ids[1].as_bytes()],
|r| r.get(0),
)
.unwrap();
assert_eq!(orphaned, 0);
}