Pre-release doc-staleness sweep for 2.5.0:
- ARCHITECTURE.md config reference: add the [maintenance] section
(enabled/forget_sweep/lint/embedding_backfill intervals +
reconcile_tombstones_deleted_pages, #929/#964) and release_base_url (#801).
- README support matrix: Hermes Agent Community -> Supported, aligning the
compact table with the authoritative docs/support-matrix.md (#933 hooks +
#942 skills are first-party).
- research-codebase-memory-mcp.md: correct the stale '19 tools' -> 23.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Named server profiles live in <data_dir>/servers.toml, with each token in
its own owner-only file under <data_dir>/auth-tokens/. A strict parser
refuses the whole registry on any unknown key, profile names are
validated so they can never become a path (or a Windows device), and
resolve() fails closed: unknown profile, missing token, outside roots,
or unrooted once several profiles exist.
`server add` keeps registered roots when --root is omitted and discards
a token whose profile moved to another URL, so rotating a token cannot
lift the roots restriction or pair one server's token with another.
`server list` never prints a token.
`uninstall` removes the stored profile tokens together with the hook
token, since the hooks it removes are their only readers, and keeps the
registry of URLs, so a later reinstall lists each profile as tokenless
rather than forgetting it.
- Adds HandoffAcceptStatus enum ('claimed' | 'consumed_by_hook' | 'none_pending')
returned alongside 'handoff' in memory_handoff_accept (#988, split from #920).
- Adds ReaderPool::handoff_claimed_by_live_session to verify if the caller's
live session already received the handoff via SessionStart hook.
- Reports 'consumed_by_hook' only when the caller forwards its session id
(e.g., Claude Code with --session-aware or OpenCode 2) and matches the
scope, live session, and owner filters.
- Reports 'none_pending' when no handoff was pending or for static clients
where the hook claim cannot be confirmed for that exact session.
- Updates documentation, routing skill, ARCHITECTURE.md, security boundaries (row 4g),
and CHANGELOG.md.
ai-memory run --yolo now, on an interactive TTY only (never in hook/CI/
detached paths):
- warns that --yolo runs every tool call unconfirmed, [Y/n] default-yes;
- offers to re-run inside ai-jail when it is installed, re-execing the
original argv under `ai-jail --network --agent-state --env <NAME>…`
(--network shares the host net namespace so the loopback ai-memory server
stays reachable; only already-set credential/config env is forwarded);
- skips both prompts when already inside ai-jail (Linux hostname ai-sandbox
/ macOS PS1 (jail) ; fails open to showing the warning; Windows never).
Opt-in Claude "true yolo" (--true-yolo / [config] claude_true_yolo,
AI_MEMORY_CLAUDE_TRUE_YOLO) silences the pauses --dangerously-skip-permissions
leaves: sets the CLAUDE_CODE_DISABLE_*_RM_* env vars and injects
--settings bypassPermissions. Claude-only, off by default, sandbox-first.
Detection and argv assembly are pure/OS-explicit (ai-memory-workstream::jail)
with unit tests for every branch; a CLI integration suite asserts the argv
against the real ai-jail via --dry-run (skips cleanly when ai-jail/bwrap are
absent). Design: docs/design-yolo-safety-ai-jail.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
memory_query gated the reserved _global preferences union on 'no named
workspace/project'. The routing doctrine tells static MCP clients to pass
workspace+project on every call, which set that gate false, so static clients
never received global_scope_hits despite the documented contract (#930). The
union now keys on single-project resolution (scopes empty); an explicit
multi-scopes set is the only opt-out, and global=true/as_of are unaffected. The
double-search guard now resolves the queried project (named or active) rather
than the active-project default. The union remains keyed strictly to the
reserved _global scope and never leaks another project's pages (adversarial
test + security-boundaries row 1b).
Closes#930
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Option A from #953 (proposed by @rntjr): [search.fts] stopwords config
key. Absent keeps the built-in English list byte-identical; [] disables
the filter; a non-empty list replaces the default outright. Case folds
with full Unicode case-folding (matching the unicode61 remove_diacritics
2 content tokenizer) but never strips diacritics, so an accented capital
still matches a configured lowercase entry while "e"/"é" stay distinct.
Threaded from Config::load through a new ReaderPool::set_fts_stopwords /
FtsStopwords, mirroring how [retrieval] already reaches ReaderPool.
Env override AI_MEMORY_SEARCH_FTS_STOPWORDS (CSV, matching
allowed_hosts/cors_allow_origins/trusted_proxy_cidrs) is applied as data
in Config::load rather than through the usual __-split figment Env
layer, since that layer merges raw values before any deserializer could
tell a present-but-blank var (must mean "unset") apart from a real
empty-list override.
Closes#953.
A multi-page batch writes each update at a path the model chose. When
that path named an existing pinned page, apply_batch replaced its body
and wrote the new version with pinned = 0, because neither the batch
loop nor the wiki write path looks at the page being replaced. Pinned
pages are documented as immutable to automation.
Skip such updates with a warning, next to the existing invariant-slot
skip. _slots/ pages are pinned automatically and keep their own
state/invariant regime, so they are not skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hbGY9oRywtM3Nz2NcXxUF
#660 stopped the indexer from superseding an OKF-conformed ledger on every
hook append, but the rows it already wrote stay and nothing removed them:
`compact` deletes nothing, `forget-sweep` only hard-deletes decay
tombstones (only `decay` writes `superseded_at`), and `reindex` loses the
DB-only state. One reported store holds 6,539 versions of 101 live pages,
95% of its rows.
Adds `ai-memory reclaim-ledger-versions` (`POST
/admin/reclaim-ledger-versions`), an online command that deletes only the
superseded versions of paths whose *content* opens with a hook log entry.
Dry run unless `--confirm`; `--drop-latest` also removes each ledger's live
row; `--compact` rebuilds the FTS index and VACUUMs to return the bytes.
The content gate is the load-bearing part and is now one definition:
`ai_memory_core::log_ledger` owns the predicates, `ai_memory_wiki::ledger`
keeps the file-reading half and delegates, and the store decides from a
`pages.body` with no filesystem access. A real page a human named
`log-2026-09.md` keeps its whole version chain. Only `is_latest=0 AND
superseded_at IS NULL` rows are eligible, so decay-owned rows stay with the
sweep that reasons about them.
Derived FTS/entity/vector/link rows go with the page through the existing
ON DELETE CASCADE. The FTS delete trigger is stood down for the bulk delete —
its DDL read back from sqlite_master and re-executed, so it cannot drift
from the schema — and pages_fts is rebuilt wholesale, which is what keeps
this from re-tokenizing tens of gigabytes of ledger body row by row.
Reading the gate needs a bounded prefix, not the whole body: GATE_PREFIX_BYTES
lives with the predicate, shared by the file reader and the store.
Refs #914
Sessions imported by `backfill` before it carried per-event timestamps all
have started_at/ended_at on the import day, which flattens every
time-ordered view of that history. This adds an admin repair for sessions
already imported; it is independent of the companion fix to `backfill`
itself (#919).
The repair runs online, through the server's single writer, because it does
not need direct disk access (invariant 9). POST /admin/repair-session-times
validates a caller-supplied batch of (session_id, started_at, ended_at)
candidates in one transaction, scoped to (workspace, project) like
purge-session. It only rewrites a row whose current started_at postdates
the candidate's own transcript end, the bug's signature, so a correctly
hook-captured session or a re-run is a no-op (`not_flattened` /
`unchanged`). Without confirm it is a real dry run: the write runs and
rolls back. Out-of-scope sessions are reported not-found and left
untouched; negative, inverted or future times are refused; an open session
never gets an end time; and a confirmed batch writes one audit_log row with
every repaired session's before/after times.
`ai-memory repair-backfill-timestamps` is the thin client. It reuses
backfill's session discovery and the workstream transcript reader, takes
each session's earliest and latest event time, resolves native ids with the
new `SessionId::from_native` (now shared with the hook router instead of a
private copy), chunks the batch at the server's cap, and prints each
repaired session's old and new times.
Every session/observation `ai-memory backfill` imported was stamped with
started_at/ended_at/created_at at import time, discarding each transcript
event's own timestamp even though the workstream adapters already captured
it (`NewWorkstreamEvent::occurred_at`) — a session imported today from a
transcript recorded weeks ago showed up as having happened today.
`NewSession` and `NewObservation` gain an optional `occurred_at`
(microseconds); the store falls back to `now()` when it is `None`, so live
hook capture is unaffected. This includes `admit_hook_session_event`'s own
session-row INSERT (the real path a live `/hook` session-start event takes,
separate from `begin_session_row`) — missing that one meant a backfilled
session's `ended_at` (original, past) could sit before its `started_at`
(import time).
The `/hook` body accepts an RFC 3339 `occurred_at`, read from the top level
of the body only (not the nested `payload`/`event`/`properties`/`info`/`path`
search other hook fields use, so a harness payload that happens to carry an
`occurred_at` key elsewhere in its own structure is never mistaken for this
field). It is client-controlled input arriving over the hook endpoint, so
`HookEnvelope::occurred_at_micros` bounds it (must be > 0 and no more than
five minutes ahead of server time) before it is trusted; anything else —
missing, unparsable, or out of bounds — resolves to `None` rather than
erroring, keeping hooks fire-and-forget. It is numeric metadata, not text, so
it never goes through the sanitizer.
Backfill validates each transcript event's own timestamp (an unparsable one
is treated as missing) and threads the resolved time through `map_event` and
into the session-start/session-end items: a missing timestamp inherits the
nearest preceding valid one, an event before the first valid timestamp
inherits that first one, and the session's start/end times are the earliest/
latest valid event time in the transcript rather than assuming it is already
time-sorted.
Because a backfilled session's `ended_at` can land well in the past, it can
sit below the auto-improve review watermark and the cross-session
experience-pass anchor (both keyed on `ended_at`), so a freshly imported
session may not get an automatic review pass until a newer session moves
those forward; an opt-in retention window measured from an observation's own
time can also make an old backfilled observation immediately prunable rather
than only after it ages in place; and the "most recently active project"
restart fallback, which looks at how recent observations are, may not pick a
project that was just backfilled. These are documented consequences, not
regressions introduced here — they follow directly from timestamps now being
honest.
Add a server-rendered page under /web that lists the pending
auto-improvement proposals of all projects. The page has a project
filter and a sort. It shows the rationale, the proposed body, and a
warning for proposals that write under _rules/ or replace a page.
The approve and reject buttons post from the browser to the existing
/admin/pending-writes/{id}/approve|reject handlers with the session
cookie and the CSRF header. Admission, audit, attribution, and the
single writer stay the same. ai-memory-web adds no write route.
The page uses the same Capability::Admin decision as /admin, and
serve passes the trusted-proxy setting to it. A non-root session gets
a 403 page. The HTML auth redirect does not send that session to the
change-password form.
Two store reads feed the page. One counts pending proposals per
project, so the total and the project filter include every project.
The other lists proposals, optionally for one project, with a cap of
500 rows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SwU4vj4tZLYVLR7vEKeh2W
Forward-merge the 2.4.x fix batch (#877,#891,#893,#888,#820,#885,#886,#884,#890,#894,#895) so main stays a subset of release/2.5.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
# Conflicts:
# CHANGELOG.md
# crates/ai-memory-cli/src/commands/run.rs
# crates/ai-memory-store/tests/suite/multi_session.rs
Two routine INFO lines made up most of the default server log:
- the wiki watcher's "reconciliation pass complete" summary fired every
RECONCILE_INTERVAL (30s) regardless of activity; drop it to debug. Its
failure signals stay loud (the per-page warn!, the watcher_degraded
error!, and the info! recovery transition).
- the external rmcp MCP SDK logs per-request lifecycle at info, which the
default filter did not cap. Prepend `rmcp=warn` to the default filter so
an operator can restore it via log_level (e.g. "info,rmcp=info") or
RUST_LOG, while `tracing_appender=warn` stays appended and thus
non-overridable — the invariant #15 feedback-loop guard.
Extract the filter string into a pure `default_filter` helper and unit-test
the ordering: the default suppresses both targets, log_level can restore
rmcp, and log_level cannot lower tracing_appender below warn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
A manual memory_consolidate MCP call wrote the page directly through the
consolidator and never touched session_consolidation_jobs, so a session
whose automatic SessionEnd job had reached the terminal `failed` state
(which claim_next never re-picks) kept showing `failed` even though the
operator had just consolidated it — a two-sources-of-truth inconsistency.
Add a writer method reconcile_session_consolidation_completed(session_id)
that flips a `failed`/`pending`/`superseded` row for the session to
`completed`. It deliberately never touches a `running` row: that is an
active worker lease, and stomping it would violate invariants #2/#16, so
the automatic path's claim-guarded complete/fail stays the only lease
transition. The memory_consolidate handler calls it best-effort after a
successful, non-dry consolidate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
`ai-memory run --env/--env-file` values reached the spawned harness and
native-session resolution but not first-launch auto-wire, so hooks and MCP
landed in the default config home, and a sentinel keyed only on agent and
version skipped a second account on the same version. Auto-wire now resolves
every install target from the launch environment, the Codex MCP entry follows
`CODEX_HOME`, and the sentinel also keys on the resolved hook and MCP paths.
Fixes found on the same paths: OMP profile and PI_CONFIG_DIR resolution,
blank relocation values, uninstall leaving sentinels behind, the Crush
context packet (default context files and the global crushrc), the Crush data
directory lookup, session discovery with concurrent launches in one checkout,
and Kiro v3 resume wiring. Spawned-server test fixtures now survive losing a
free port to another socket.
Refs #820
The zero-LLM contradiction detector in memory_lint used a hardcoded
absolute cosine band (0.4-0.75). On a single-language or single-domain
store the background similarity floor is elevated, so the fixed band
admits many same-domain-but-unrelated pairs as noise, up to the 25-finding
cap.
Add contradiction_band_min / contradiction_band_max config keys (env:
AI_MEMORY_CONTRADICTION_BAND_MIN / _MAX), defaulting to the historical
0.4 / 0.75 so existing installs see no behavior change. Validated at
config load (finite, 0.0 <= min < max <= 1.0). Threaded through
LintOptions and into the three call sites that build it (admin HTTP
lint, the memory_lint MCP tool, and the scheduled lint tick).
Closes#853.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
- Reformatted code in `upgrade.rs` for better readability, including consistent indentation and line breaks.
- Simplified the `refuse_oversized_content_length` function signature for clarity.
- Updated test cases for improved formatting and consistency in assertions.
- Adjusted the `ARCHITECTURE.md` documentation to reflect the new order of commands.
Adds `embedding_query_prefix` / `embedding_document_prefix` config keys
(env: AI_MEMORY_EMBEDDING_QUERY_PREFIX / AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX)
applied by the openai and openai-compat embedders: an optional string
prepended to query/document text before the existing truncation, so
truncation still bounds the whole input. Asymmetric embedding models
(Nemotron-3-Embed, base E5, Qwen3-Embedding) need a "query: " /
"passage: " instruction their publisher specifies; the OpenAI-compatible
/v1/embeddings wire format has no field for it. Empty by default, and not
trimmed so a publisher's trailing space is preserved.
Also fixes memory_query's embed_query helper (ai-memory-mcp/src/server.rs),
found while auditing every embed/embed_query/embed_document call site: it
called the generic Embedder::embed instead of Embedder::embed_query, so an
asymmetric embedder (google's task-typed embeddings, or these new prefixes)
embedded the search query on the document side.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
ai-memory serve leaked one file descriptor per dead hook/MCP peer.
A client whose connection dies without sending FIN (laptop sleep, a
VPN/Tailscale flap, an abrupt kill) leaves the accepted socket
ESTABLISHED forever, since the OS keepalive default is off. Over ~2-3
days of normal churn that exhausts the 1024-fd default and breaks the
healthcheck -- an unauthenticated availability/DoS. The rmcp
session-table half of this leak was already fixed in 2.4.0 by the
rmcp 2.x bump; this closes the remaining half-open-socket half.
Accepted sockets now get TCP keepalive via socket2, wired through
axum::serve::ListenerExt::tap_io (a hand-rolled axum::serve::Listener
newtype was tried first, but into_make_service_with_connect_info's
Connected<IncomingStream<'_, L>> bound is only implemented by axum for
its own TcpListener and, generically, for TapIo<L, F> -- never for an
arbitrary third-party L, and the orphan rule blocks implementing it
ourselves since neither Connected, SocketAddr, nor IncomingStream is
local to this crate). tap_io keeps the real peer SocketAddr flowing to
ConnectInfo while still touching every accepted stream.
New tcp_keepalive_secs config field (default 60s; AI_MEMORY_TCP_KEEPALIVE_SECS
env override; 0 disables keepalive entirely), read once through the
existing Config::load path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
`load_patchable_pages` hardcoded `_rules/` and `procedures/` as the only folders
whose page bodies reach the reviewer. Every other page arrives through
`render_recent_pages` as one line — path, title, kind, updated_at — with no
content, so the model cannot tell that what it is about to propose is already
written down.
For a project whose durable knowledge lives in `decisions/` or `gotchas/`, with
no `_rules/` pages at all, the filter returns an empty vector and the reviewer
sees no page body whatsoever. It then re-proposes invariants that already exist,
at high confidence, because the claim really is well evidenced by the session —
confidence tracks the evidence, and could never track "the wiki already says
this".
The workaround is to move normative content into `_rules/`. That is good hygiene
independently, but it makes folder choice silently decide model visibility, and
the wiki's own conventions encourage `decisions/` and `gotchas/`.
`patchable_page_prefixes` defaults to the historical pair, so every existing
config behaves exactly as before; a project can add the folders where its
knowledge actually lives instead of restructuring its wiki.
Two boundaries are deliberate and tested. An empty list sends no page bodies —
"no folders are patchable" must not mean "every folder is". An empty string in
the list is ignored, because `"".starts_with` matches every path and would turn
the filter off silently. Prefixes carry their trailing slash, so `decisions/`
does not also match `decisions-archive/`.
This is the narrowest of the three options in #834. It does not change the
retrieval strategy (embedding-nearest dedup context) or the recency ordering
that lets `sessions/` churn crowd the recent-page list; both are noted in the
docs as still open.
3382 workspace tests pass, clippy -D warnings and fmt clean. Mutation-tested:
adding `decisions/` to the default fails two tests, and letting an empty prefix
match fails another.
Closes#834
Add the LLM "dream" pass: cross-session rewrite/merge of cold clusters,
scheduled on idle with cancel-on-activity, surprisal-first ordering
(docs/design-memory-aging.md B2/B3/B4).
B2 — cross-session LLM merge. `dream::run_dream_pass` clusters the SAME
bounded cold set the forget sweep materialises (reuses A3's adaptive_eps /
dbscan; new pub(crate) `sweep::materialize_cold_set` shares the scoring,
invariant #2) and hands each multi-member cluster to the provider to rewrite
into ONE coherent page via JSON-schema structured output (invariant #7,
`DreamMergedPage`). Routes through the gated apply path
(`preflight_admission(Consolidate)` before the LLM, then `Wiki::apply_batch`)
with dry_run first. Never deletes a source (invariant #16): the survivor is
rewritten, every merged-away member is superseded with a merge-note stub
(reachable via git + supersession chain), and `page_evidence`
(reconsolidation + `b2_dream:<id>`) records the sources. No provider / no
embedder / flag off ⇒ a clean no-op; the zero-LLM A3 path is untouched
(invariant #13). No migration.
B3 — event + idle scheduling with cancel-on-activity. New `[dream]` config
and a scheduled job in serve.rs that runs the pass only after a configurable
idle window (read from the MCP tool router's shared `ActivityClock`, bumped on
every tool call) and cancels it the moment activity resumes (a cheap
`DreamCancel` polled between clusters; a watcher flips it on renewed activity).
Bounded to max_clusters_per_run per run (invariant #5), all writes via the
single-writer actor (invariant #2). Every run yields an observable
`DreamReport` — no silent window.
B4 — surprisal-first ordering. The work queue is ordered by each cluster's
embedding distance to the nearest existing (non-cold, latest) page —
most-novel first. Pure ordering, unit-tested.
Opt-in LLM, OFF by default, R2-gated before default-on. Tests use a fake
LlmProvider (zero live calls). CHANGELOG + ARCHITECTURE updated;
comparison.md / competitive-parity.md left for the closing docs pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Turn the dormant page_evidence substrate (V63) into a read-time, zero-LLM
belief-strength `confidence` that can shade retrieval ranking as ONE bounded
factor — exposed inertly by default, foldable into authority only behind a
default-OFF, R2-gated flag (design-memory-aging.md B1 /
design-hindsight-borrowings.md §3).
- New pure `ai_memory_store::belief` module (mirrors decay.rs): confidence is a
bounded [0.0, 0.95] function of distinct supporting sessions (breadth, not
raw count), recency of the newest sighting, and live `contradicts` count.
Anti-entrenchment baked in: distinct-session weighting, a residual cap so
non-session volume can't dominate, recency shading to a floor, and a hard
confidence cap so evidence alone never pins authority at the ceiling.
- New batched reader `page_belief_inputs` (one evidence-aggregate query + one
live-contradiction query, never per-hit; invariant #2). Fetched only when an
explained query needs it or the weight is on, so the hot default path is
untouched. Replaces the now-subsumed `page_evidence_counts`.
- SearchExplain gains `confidence` and `belief_factor` beside `evidence_count`;
all populated on the explained path (inert). memory_status gains
`evidence_rows`.
- Folded into PageAuthority inside the existing [0.55, 1.50] clamp — one more
bounded factor, no new multiplier tower — behind
`[retrieval] belief_authority_weight` (default 0.0 = OFF, byte-identical
ranking). A supersession always wins regardless of evidence: a superseded
version's stale evidence is never boosted, and confidence never gates whether
a write/correction takes (invariant #16).
R2: the authority factor ships OFF; enabling it is gated on a positive R2 delta
(retrieval triple / QA), not yet performed.
Tests: pure monotonicity/bound/cap/recency/contradiction unit tests; default-OFF
byte-identity proof; ranking-on boost; inert explain/status exposure; live
contradiction weakening; supersession-always-wins. Full gate green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Flag likely-conflicting pages through the existing memory_lint. Cold
semantic/procedural pages whose already-stored embeddings sit in the
0.4-0.75 cosine-similarity band ("same topic, not a near-duplicate" -- the
shape of a likely contradiction; >=0.75 is A3 dedup territory, <0.4
unrelated) get an advisory `contradiction` finding naming both pages with
newer-wins timestamp advice.
Zero generative LLM (invariant #13): reads only existing embeddings + cosine.
No embedder configured, or no embeddings for the (provider, model, dim)
triple, is a clean no-op -- never an error, never a provider call. Advisory
only (invariant #16): emits a finding, never deletes/edits/supersedes a page.
No persistence, no migration: `contradicts` edges live in the body-derived
`links` table (replace_links_in_tx rewrites it from the page body on every
write), so a programmatic edge would be silently wiped -- hence a lint-time
detector, not a stored edge. Bounded (invariant #2): one embeddings load
over the capped cold set, capped findings, deterministic ordering. No new
MCP tool (still 23).
Runs on the user-invoked memory_lint (MCP + admin) off the configured
embedder's triple; the automatic scheduled lint stays rule-based.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Two independent, opt-in, non-destructive "consolidation hygiene" features
(docs/design-memory-aging.md buckets A3 + A4). Both off by default, zero
generative LLM, and reuse existing tables (no new migration).
A3 — cold-cluster dedup. The forget-sweep can cluster near-duplicate cold
episodic pages by embedding (cosine DBSCAN, adaptive k-distance eps clamped
to a conservative ceiling, minPts=2) over the bounded cold set it already
materialises, and collapse each cluster to one survivor: the highest-retention
member, its body the extractive union of the cluster's keep-tokens (every
member's durable facts survive), with the other members superseded by a merge
note pointing at the survivor. No source is hard-deleted (invariant #16 — merged
members stay reachable via supersession + git, recoverable with restore-page);
merge provenance is recorded in page_evidence (reconsolidation kind, a3_merge:
source-id prefix — the closed vocabulary has no A3 kind and adding one would need
a CHECK-constraint migration). Reads only already-stored embeddings, so it is a
clean no-op with no embedder or no matching (provider,model,dim) triple
(invariant #13). New pure module cold_cluster.rs (DBSCAN + adaptive eps). Gated
by [decay] dedup_cold_clusters (default false), wired into the scheduled sweep;
results surface in SweepReport.merged.
A4 — entropy / boilerplate pre-filter. A pure, zero-LLM Shannon-entropy +
boilerplate gate (entropy_filter.rs) that skips low-information session pages
(near-empty, whitespace, single-character, highly-repetitive) from the
cross-session experience consolidation pass before they reach the prompt, eval
gate, or apply_batch. Advisory (skip, never delete; invariant #16), off by
default via [auto_improve.scheduler.experience_entropy_filter] with conservative
validated thresholds tuned so a terse-but-informative note is KEPT.
R2: both ship opt-in/off, so R2 is not a merge blocker; the recall
no-regression proof is the gate before any future default-on. No R2 number was
run for this change. No new MCP tool (still 23).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Add opt-in extractive tier-down (docs/design-memory-aging.md §A2): when
`[decay] compact_cold_episodic` is on, the forget-sweep COMPACTS a cold
episodic page instead of evicting it — keeping the L0 frontmatter
`abstract:`, an L1 first-paragraph summary, and an L2 regex-mined
keep-token set (paths, URLs, code spans, error codes, UPPER_SNAKE
constants, long identifiers), dropping the prose body. Tier-down beats
eviction because the durable facts survive.
Off by default, so an upgrade changes nothing. Fully zero-LLM (regex
only, #13). Reversible and non-destructive: the rewrite goes through the
wiki layer (write_page/apply_batch), so the full pre-compaction body
stays in git history and the supersession chain and is recoverable with
restore-page (#16). The compaction pass runs BEFORE the decay-eviction
pass and batches its writes (#2).
New V65 migration adds a nullable `pages.compacted_at` marker: additive
`ADD COLUMN` only, no backfill — populated lazily by the sweep, derived
at the single upsert choke point from a `compacted: true` frontmatter
mirror so the marker and body commit in one transaction (#3, #10). The
sweep and curator skip a marked page, so it is never re-compacted,
re-evicted, or re-reported as cold. Downgrade guard preserved
(set_abort_missing(true)).
Ships opt-in/off; the R2 recall no-regression proof is the gate before
any future default-on. No new MCP tool (still 23).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Replace the single global decay rate λ with an optional per-Tier
half-life lookup, config-driven, so an operator can keep episodic
history longer and working-tier scratch shorter (the
mcp-memory-service 365/180/90/30 shape) instead of one λ for every
tier.
The default is byte-identical to today: `DecayParams::tier_lambda`
is an empty `TierLambdas` (a Copy struct of four `Option<f64>`, not
a HashMap, so `DecayParams` stays `Copy`). `lambda_for(tier)` returns
the scalar `lambda` unchanged for any unset tier — no days↔λ
round-trip that could perturb an unconfigured store's scores — so an
upgrade changes no score and mass-evicts nothing on the first
post-upgrade forget-sweep. The load-bearing property test asserts
byte-for-byte (to_bits) equality with the historical scalar-λ formula
across a grid of inputs for every tier.
`retention_score`/`retention_score_with_breadth` now take the page's
tier; the two retention-scoring callers (sweep, curator) pass
`c.tier`. Config exposes an opt-in `[decay.half_life_days]` table in
DAYS (converted to λ = ln(2)/days), wired through `DecaySettings`,
`decay_params()`, and load-time validation (each override finite and
> 0). Pure math + config: no new column, no migration, no new MCP
tool (still 23).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Close the access-reinforcement gap so memory surfaced through more MCP read
paths resists decay. Previously only memory_query and memory_recent bumped
pages.access_count + last_accessed_at (the M8 decay reinforcement term);
memory_read_page, its include_related graph walk, and memory_explore
reinforced nothing.
Wire the existing sanctioned spawn_access_bump into all three:
- memory_read_page bumps the resolved page (via latest_page_id_by_ids), and,
when include_related is set, the walked neighbours too — fired only on the
success paths so an errored read reinforces nothing.
- the related walk now carries each node's PageId (serde-skipped, so the wire
payload is byte-identical) for the walk-bump.
- memory_explore resolves the surfaced snapshot pages (rules, slots, recent,
pinned, settled) to ids in one batched read (latest_page_ids_by_paths, no
per-page N+1) and bumps them.
Strictly additive: reuses the throttled (<=1 per (page,operator) per 60s),
FTS-exempt, fire-and-forget single-writer path, so invariants #2/#3/#13/#16
hold. Response payloads unchanged; no new MCP tool (still 23).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Rebase-via-merge of PR #774 (memory_status scope reporting) onto the
current release/2.4. Integrated the traced scope-resolution path
(resolve_read_args_traced / ScopeSource / ReadPointer) with the
error-classification and reasoning/pinned/answer features already on
release/2.4. Resolved conflicts in CHANGELOG.md (kept release/2.4's
Added/Changed lists and appended #774's memory_status entry, ref #757,
#774) and docs/ARCHITECTURE.md (kept #774's memory_status row plus
release/2.4's memory_briefing/memory_explore rows).
docs(changelog): reference #774 for memory_status scope reporting
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm