mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
1104 lines
83 KiB
Markdown
1104 lines
83 KiB
Markdown
# ai-memory - Architecture
|
||
|
||
> One canonical doc for "what is this thing and how is it shaped".
|
||
> Long-form research lives next to this file under [`docs/`](.); this
|
||
> page is the operational summary for someone reading the code.
|
||
|
||
## Purpose
|
||
|
||
ai-memory is a single Rust binary that gives the coding agents in the
|
||
[README Support Matrix](../README.md#support-matrix), plus other MCP-capable
|
||
clients, long-term memory shared across CLIs.
|
||
Quit one mid-task; open another in the same directory; continue. No
|
||
manual `write_note` ceremony, no copy-pasting summaries between
|
||
sessions.
|
||
|
||
The artifact you accrete is a **Karpathy-style LLM wiki**: a
|
||
git-versioned tree of markdown pages on disk that gets *compiled* over
|
||
time, appended-to. Pages are versioned in place via
|
||
supersession, semantic concepts compound, episodic logs decay. A
|
||
companion SQLite index gives FTS5 + lexical entity + link-neighbor retrieval,
|
||
with optional vectors; the markdown stays the source of truth.
|
||
|
||
## Data flow
|
||
|
||

|
||
|
||
Solid arrows are request, read, and write paths. Dashed arrows are
|
||
background reconciliation or provider-backed maintenance. The core invariant is
|
||
unchanged: the markdown wiki is the source of truth, and SQLite is the derived
|
||
index for search, sessions, observations, handoffs, audit, embeddings, and the
|
||
optional managed-workstream continuity ledger.
|
||
Auto-improvement sits on the provider-backed maintenance side: the server
|
||
schedules reviews for newly completed sessions in every project, records
|
||
validated proposals in the pending-writes audit trail, and auto-approves them
|
||
through the normal wiki write path by default. Scheduler ticks are
|
||
non-overlapping; long all-project review passes delay the next tick instead of
|
||
starting another copy. Scheduling and approval are separate. Admins can set
|
||
`[auto_improve.scheduler] enabled = false` to stop background review, or
|
||
`[auto_improve] require_approval = true` to leave scheduled and manual proposals
|
||
pending for review. Operators can additionally set `[auto_improve.eval]` to run
|
||
a project-supplied executable gate for selected proposal prefixes after LLM
|
||
validation and before staging/approval; it is disabled by default and never runs
|
||
from hook paths.
|
||
|
||
**Steady-state loop:**
|
||
|
||
1. Agent CLI emits a lifecycle hook (SessionStart, UserPromptSubmit,
|
||
PostToolUse, …). Shell-script hooks `curl` event JSON to `POST /hook`
|
||
with a short timeout. Native `ai-memory hook --event ...` commands spool
|
||
events locally with a stable per-entry idempotency key, do a short bounded
|
||
cleanup at session start, and hand
|
||
session-end delivery to a detached lock-aware `hook-drain` helper;
|
||
high-latency operators can raise the drain/handoff/background caps with
|
||
minute-based env vars.
|
||
Agent hot paths never block on the network; saturated servers return HTTP
|
||
429 instead of queueing unbounded work.
|
||
For supported native commands and generated OpenCode/OMP/Pi/OpenClaw
|
||
integrations, the nearest-marker capture policy runs first: a dropped
|
||
recognized file-tool event never enters spool, queue, transport, logs, or
|
||
storage. See [Capture exclusions](marker-file.md#capture-exclusions).
|
||
2. Server's hook router sanitises the payload (the only path from
|
||
untrusted text into the store), assigns an [`ObservationKind`], and
|
||
enqueues a `WriteCmd` to the writer actor. For native keyed events, the
|
||
project-scoped key and observation commit together. The key is marked
|
||
complete only after downstream processing: an incomplete replay resumes
|
||
wiki/handoff effects without another observation, while a completed replay
|
||
is acknowledged and skipped. A bounded per-project/key gate serializes an
|
||
overlapping retry with the original processor. Downstream effects remain
|
||
at-least-once until that completion marker, so a process crash during those
|
||
effects may repeat an already applied effect rather than silently lose the
|
||
rest. For an interrupted already-ended SessionEnd, the replay converges the
|
||
wiki commit, durable provider job, and pending key without appending another
|
||
observation. `log.md` gets an appended
|
||
`## [YYYY-MM-DDTHH:MM:SSZ] <event> | <title>` line.
|
||
3. On true `SessionEnd` events, the server synthesises a
|
||
`sessions/<id>.md` summary page (rule-based, no LLM) and opens a
|
||
`Handoff` row for the next agent. One SQLite transaction inserts that
|
||
automatic handoff, stamps the session ended, and records the covered
|
||
observation count, so recovery never sees only half of those DB effects. A
|
||
later SessionEnd re-runs the path only when that generation advances, so
|
||
resumed sessions are captured while duplicate delivery and clock skew
|
||
converge. Existing ended sessions are baselined at migration instead of
|
||
becoming historical catch-up work. Auto-commits the wiki. Clients
|
||
without a reliable true session-end hook need an explicit ending action:
|
||
`ai-memory finalize-session --agent antigravity-cli` for Antigravity CLI
|
||
(Codex has a native `SessionEnd` since CLI 0.145.0; `finalize-session
|
||
--agent codex` is only the fallback on older Codex).
|
||
The command selects the latest matching open session and enters the same
|
||
canonical SessionEnd path as a native hook. Generated session-page
|
||
frontmatter records `session_id` plus the immutable `sessions.agent_kind`
|
||
as `agent`; it describes the page's harness origin, not the later writer.
|
||
Manual page writes do not receive inferred agent metadata.
|
||
4. When `AI_MEMORY_LLM_PROVIDER` is set, `memory_consolidate` rewrites
|
||
that summary into a richer durable page or fans out into a
|
||
multi-page batch under `concepts/`, `decisions/`, `gotchas/`. Consolidation
|
||
prompts preserve the source material's dominant natural language and ask
|
||
the model to connect related pages with path-based wikilinks.
|
||
5. When an LLM provider is configured, the auto-improvement scheduler reviews
|
||
newly completed sessions across all projects outside hook latency. It records validated
|
||
`concepts/`, `decisions/`, `gotchas/`, `procedures/`, and `_rules/` proposals
|
||
in the pending-writes audit trail, then approves them through the wiki
|
||
mutation path by default. The scheduler initializes a per-project first-run
|
||
watermark so historical sessions are not processed automatically on upgrade,
|
||
then records per-session claims before LLM work so failed scheduled reviews do
|
||
not retry forever. Explicit CLI/admin/MCP auto-improve calls use the same
|
||
pipeline for targeted reruns or catch-up. With `[auto_improve]
|
||
require_approval = true`, scheduled and manual proposals remain pending until
|
||
explicit pending-writes approval. If `[auto_improve.eval] enabled = true`,
|
||
targeted proposals (default `_rules/` and `procedures/`) must pass the
|
||
configured executable JSON contract before they are staged; failures become
|
||
rejected candidates/rejection-buffer entries rather than wiki writes.
|
||
Every LLM prompt treats repository text, observations, wiki pages, and prior
|
||
proposals as untrusted data rather than instructions. The same explicit
|
||
trust boundary and delimiters precede automatically injected handoffs,
|
||
project briefs, and managed-workstream packets; current instructions and
|
||
checkout state remain authoritative.
|
||
6. `memory_query` answers via FTS5 + entity-match + link-neighbour RRF; when an
|
||
embedder is configured, vector cosine over `page_embeddings` joins the same
|
||
RRF. The entity index is derived from the canonical frontmatter `entities`
|
||
list, and an empty index contributes no candidates or score. Before final
|
||
truncation, a bounded authority multiplier adjusts relevance using canonical
|
||
page kind, tier, `pinned`, and explicit positive/negative frontmatter tags.
|
||
It favors maintained rules, decisions, procedures, and gotchas in close
|
||
contests while keeping episodic, historical, lint, and test evidence
|
||
searchable. No query-intent regex or hard exclusion participates. An
|
||
optional `AI_MEMORY_RERANKER=llm` pass sends a bounded query plus up to 30
|
||
bounded titles/snippets to the configured provider after project/scope
|
||
fusion; it is limited to one call per query and four calls in flight, and any
|
||
invalid, failed, timed-out, or saturated attempt preserves the local order.
|
||
Global and supplemental global-preference results do not take this path. If
|
||
compiled wiki pages miss entirely in default, explicit project, or explicit
|
||
`scopes` mode, bounded raw observation FTS returns fallback `raw_hits`;
|
||
`global=true` searches compiled wiki pages across projects only. Page hits
|
||
bump `access_count` + `last_accessed_at` - the M8 reinforcement term, which
|
||
`memory_feedback` complements with explicit per-page salience. Identified
|
||
operators also add one `page_access` row per page; an opt-in
|
||
`[decay] breadth_weight` can reward pages reinforced by several distinct
|
||
operators. That bump is throttled to at most once per page per minute, so a
|
||
burst of overlapping searches does not flood the writer actor with
|
||
redundant reinforcement writes. The same reinforcement fires from every
|
||
read path that surfaces a page, not just search: `memory_read_page` (and its
|
||
`include_related` walk, which reinforces the walked neighbours too) and
|
||
`memory_explore` (the pages it surfaces) bump the same counters through the
|
||
same throttled, FTS-exempt path, so a page a human opens directly or the
|
||
graph surfaces resists decay like a query hit (design-memory-aging.md C1).
|
||
7. The forget sweep runs on demand and on the server's `[maintenance]`
|
||
schedule: pages past their frontmatter `expires_at:` TTL are
|
||
hard-deleted through the wiki layer (file + rows, pin or not);
|
||
pages with `retention < cold_threshold` are evicted through the wiki layer,
|
||
which removes the authoritative file and leaves a decay tombstone;
|
||
tombstones older than `hard_delete_after_days` are purged with their full
|
||
version ancestry only within that sweep's resolved workspace/project,
|
||
together with entity-index rows orphaned by the purge. A newer page recreated
|
||
at the same path is preserved. Semantic / pinned / freshly-touched pages
|
||
survive. A fourth pass, disabled unless `[decay] observation_retention_days`
|
||
is positive, then deletes raw `observations` older than that age — but only
|
||
for sessions already consolidated into a summary page that is still live, so
|
||
raw capture is never removed while it is the last copy of that session's
|
||
work. It runs last so this run's own evictions and hard-deletes already
|
||
exclude their sessions, deletes in `observation_prune_batch` transactions so
|
||
a multi-million row prune cannot hold the write lock, and repairs
|
||
`sessions.ended_observation_count` downward in the same transaction. The
|
||
prune is irreversible in a specific sense: observations are the input to
|
||
consolidation, so a pruned session can never be re-consolidated — not with
|
||
a better model, a better prompt, or a fixed consolidator bug — and its
|
||
summary page becomes the only surviving account of that session. Freed
|
||
SQLite pages are reused, not returned to the OS: the `.db` file does not
|
||
shrink, the backup tarball does.
|
||
Scheduled sweep, rule-based lint, and opt-in embedding backfill ticks
|
||
enumerate every existing workspace/project scope before doing per-project
|
||
work, matching the auto-improvement scheduler's store-wide scope model. A
|
||
separate daily cleanup removes week-old project rows only when they contain
|
||
no pages, sessions, observations, handoffs, managed workstreams, or
|
||
auto-improvement data; managed continuity history therefore keeps its
|
||
project scope alive even when no lifecycle-hook session has been captured.
|
||
8. Backups: `ai-memory backup --to <tarball>` uses SQLite's online
|
||
backup API so the source stays writable; `ai-memory restore`
|
||
reverses. Or: `git push` the wiki dir + `rsync` the data dir.
|
||
|
||
**Optional managed-workstream loop:** `ai-memory run` opens a lease for the
|
||
current repository/worktree workstream, resolves an explicit harness or the
|
||
newest usable local/linked harness, creates or resumes that harness's native
|
||
session, and marks lifecycle calls with an invocation-scoped run id.
|
||
SessionStart injects an unseen bounded event range; Crush receives it through a
|
||
temporary supported global-context path because it lacks SessionStart. The host
|
||
imports the native transcript tail and a Git checkpoint when the child exits.
|
||
Every injected packet starts with a versioned origin marker. The Claude
|
||
transcript normalizer excludes a marked packet if Claude persists and reads it
|
||
back, preventing delivered history from recursively re-entering the ledger.
|
||
An explicitly pending handoff is delivered before the managed event range;
|
||
their single-use delivery claims share one writer transaction after the
|
||
complete startup response has been assembled. Manual handoffs take precedence;
|
||
otherwise the newest cwd-eligible automatic handoff is delivered, and that
|
||
same transaction expires older eligible automatic handoffs while preserving
|
||
manual and sibling-directory work. Insertion also expires prior open automatic
|
||
handoffs from the exact cwd, bounding repeated SessionEnds before any receiver
|
||
starts.
|
||
ai-memory opens native stores read-only. Raw sanitized JSONL segments are
|
||
immutable, while SQLite supplies monotonic sequences, FTS, native
|
||
source/delivery cursors, and idempotent retry state. A full-ledger
|
||
`workstream-search` path complements
|
||
the bounded startup packet. An interactive empty workstream may adopt a
|
||
checkout-matching native session once. Eligibility comes from authoritative
|
||
ledger/session state: after any harness establishes the workstream, a newly
|
||
joining harness starts fresh and receives portable history instead of adopting
|
||
unrelated old native history. Handled launcher failures cancel their lease;
|
||
normal reopen retries brief finalization conflicts, while an unclean process
|
||
death remains bounded by the renewable lease expiry. An explicit
|
||
`--force-unlock` recovery expires and replaces a selected active lease in the
|
||
same writer transaction, but only when its durable operator attribution equals
|
||
the new run's attribution; the informational `host:pid` lease label is never an
|
||
authorization key. The old run can no longer heartbeat or finish. See [Managed
|
||
cross-harness workstreams](managed-workstreams.md).
|
||
|
||
## Hook event vocabulary
|
||
|
||
The core observation vocabulary is a closed set of agent lifecycle
|
||
events. Hook bridges may accept client-specific aliases, but storage
|
||
normalises them to exactly one of these `ObservationKind` values:
|
||
|
||
| Stored kind | Semantics |
|
||
|---|---|
|
||
| `session-start` | Agent session began; cwd/model/session identity captured. |
|
||
| `user-prompt` | User submitted prompt text to the agent. |
|
||
| `pre-tool-use` | Agent is about to call a tool. |
|
||
| `post-tool-use` | Agent finished a tool call. |
|
||
| `pre-compact` | Agent is about to compact or compress its context. |
|
||
| `post-compaction` | Agent compacted context and supplied a post-facto summary or checkpoint. |
|
||
| `notification` | Agent emitted a notification-style event. |
|
||
| `stop` | Agent finished an interactive turn or stopped naturally. |
|
||
| `session-end` | Agent session ended; summary/handoff path may run. |
|
||
| `other` | Unknown or unsupported hook event. |
|
||
|
||
Antigravity CLI has no native SessionStart event. Its `PreInvocation` hook
|
||
fires before every model call, so the bridge maps only the documented
|
||
`invocationNum = 0` payload to `session-start`; later invocations are ignored
|
||
before spool or network side effects.
|
||
|
||
Unknown events do **not** expand the enum and, by default, leave no
|
||
source-event metadata in storage; they collapse to `other`. Third-party
|
||
integrations that need their own vocabulary can opt in by sending
|
||
`extension=<namespace>` on `/hook`. With a valid extension namespace,
|
||
ai-memory stores an explicit `source_event=<name>` when provided, or the
|
||
unknown `event` string when `source_event` is omitted. The stored pair is
|
||
nullable observation metadata; `kind` stays canonical. This is an
|
||
extension seam, not a runtime plugin system: external processors must use
|
||
the existing HTTP/MCP APIs and cannot bypass the sanitizer, hook
|
||
backpressure, or single-writer SQLite actor.
|
||
|
||
An external lifecycle producer can set `AI_MEMORY_CAPTURE_OWNER` on the harness
|
||
process to suppress its installed native capture while retaining supported
|
||
handoff/briefing delivery and MCP. The producer uses `extension`/`source_event`
|
||
for provenance and stable, namespaced `ingest_key` values for retries. See the
|
||
[external capture contract](external-lifecycle.md) for batching, identity and
|
||
the limits of this cooperative process-scoped mode.
|
||
|
||
Lifecycle bodies have content limits independent of the 10 MiB HTTP request
|
||
limit. User prompts and post-compaction summaries are capped UTF-8-safely at
|
||
16 KiB; notification and tool excerpts are capped at 2 KB. Native
|
||
`ai-memory hook` commands apply the event-specific cap before local spooling
|
||
and transport, and the server repeats it when parsing every request so direct
|
||
and older clients cannot bypass it. The typed sanitizer boundary then applies a
|
||
16 KiB backstop to every durable observation body after redaction. The
|
||
separately gated Claude Code assistant/Stop excerpt remains capped at 2 KB.
|
||
|
||
An optional RFC 3339 `occurred_at` on the hook body lets a client supply the
|
||
event's own original time — currently only `ai-memory backfill`, replaying a
|
||
transcript's per-event timestamps, so imported sessions/observations are
|
||
stamped with when they actually happened instead of import time. It is read
|
||
from the top level of the body only (unlike the nested `payload`/`event`/
|
||
`properties`/`info`/`path` search other hook fields use), so a real harness
|
||
payload that happens to carry an `occurred_at` key somewhere in its own
|
||
structure is never mistaken for this field. It is numeric metadata, not text,
|
||
so it never goes through the sanitizer (outside invariant #6's boundary); it
|
||
is still client-controlled input over `/hook`, so
|
||
`HookEnvelope::occurred_at_micros` bounds it (must be > 0 and no more than
|
||
five minutes ahead of server time) before trusting it. Anything else —
|
||
missing, unparsable, or out of bounds — resolves to `None`, which the store
|
||
treats as "now", never an error, keeping hooks fire-and-forget (invariant #5).
|
||
A backfilled session's `ended_at` can therefore land well in the past, which
|
||
can push it below the auto-improve watermark and the experience-pass anchor
|
||
(both keyed on `ended_at`), and a retention window measured from an
|
||
observation's own time can make an old backfilled observation immediately
|
||
prunable rather than only after it ages in place.
|
||
|
||
## Storage architecture
|
||
|
||
**Two layers, one source of truth.**
|
||
|
||
* `<data_dir>/wiki/` - markdown source of truth. Owned by a `git2`
|
||
repo so every consolidation pass + every session-end produces a
|
||
durable commit. Editable by hand in Obsidian / vim - the watcher
|
||
reconciles outside edits.
|
||
* `<data_dir>/db/memory.sqlite` - derived index. WAL mode. One
|
||
writer actor owns the writer `Connection`; reads go through a
|
||
cloneable read-only pool.
|
||
* `<data_dir>/raw/` - immutable sanitized managed-workstream JSONL segments.
|
||
Legacy raw fallback recall searches the durable `observations` table via
|
||
FTS5; lifecycle HookEnvelope JSON is not a complete transcript archive.
|
||
* `<data_dir>/logs/` - rolling daily `tracing` output.
|
||
* `<data_dir>/models/` - reserved for bundled embedding models
|
||
(M9.5+, when local `ort` lands).
|
||
* `<data_dir>/client-projects.json` - private, client-local checkout links for
|
||
`ai-memory show`, keyed by credential-free server identity plus workspace and
|
||
project. It is not part of the SQLite/wiki source of truth, and no server API
|
||
exposes host paths.
|
||
|
||
**Schema (current head):**
|
||
|
||
| Table | What |
|
||
|---|---|
|
||
| `workspaces`, `projects` | Top of the 3-tuple identity coordinate. `projects.identity` / `identity_source` (V70, #708) hold the repository identity a project routes by — an explicit marker `identity` or a normalised git remote, resolved client-side — unique per workspace when set; empty until a capture claims the project. See `docs/marker-file.md#repository-identity`. `projects.access_mode` (V68; `open` default / `restricted`) decides whether any authenticated user or only root, the creator (`projects.created_by`, V69) and grant holders (`project_grants`, V68) reach it, decided by `ai_memory_store::authorize_project` — `docs/users.md#per-project-access`. |
|
||
| `pages` | Versioned wiki pages with `is_latest` + `supersedes` chain. M8 columns: `last_accessed_at`, `access_count`, and decay-only tombstone marker `superseded_at`. M9 cols: `embedding_provider`, `embedding_model`, `embedding_dim`. V36: `expires_at` (frontmatter TTL). V37: `salience` (NULL = `salience_default`; derived from `page_feedback`). |
|
||
| `pages_fts` | FTS5 virtual table over `(title, body)`, auto-synced by triggers. |
|
||
| `sessions`, `observations` | Sanitized, bounded lifecycle-hook projections. `sessions.ended_observation_count` is the stable generation watermark for resumed-session re-end eligibility; wall clocks are not used for that decision. They are an operational audit trail, not a complete native transcript. |
|
||
| `session_consolidation_jobs` | Durable, observation-generation-idempotent queue for the *automatic* SessionEnd LLM consolidation worker. One bounded server worker leases jobs, retries provider failures with backoff, and recovers expired leases after restart. A manual `memory_consolidate` writes its page out-of-band and then reconciles this table (flipping a `failed`/`pending`/`superseded` row for the session to `completed`, never touching a live `running` lease), so the operator does not see a `failed` job for a session that is in fact consolidated. |
|
||
| `observations_fts` | FTS5 virtual table over raw observation `(title, body)`, used only as bounded fallback. |
|
||
| `workstreams`, `managed_runs`, `workstream_native_sessions` | Optional lease state plus per-harness native source and delivery cursors for `ai-memory run`. |
|
||
| `workstream_events`, `workstream_events_fts` | Append-only normalized visible transcript events and full-text search; immutable sanitized source batches also live under `raw/workstreams/`. |
|
||
| `links` | Wikilink / markdown cross-references. `to_page_id` (a global PageId) is nullable for unresolved forward links. `to_workspace` / `to_project` carry a cross-project scope (NULL = the source page's own project). |
|
||
| `handoffs` | Typed cross-agent handoff records (open / accepted / expired). |
|
||
| `page_embeddings` | Optional vector rows for latest pages, with `(provider, model, dim)` denormalised so hybrid search can ignore stale vectors after an embedding config change and report missing-embedding diagnostics. |
|
||
| `page_feedback` | Append-only `memory_feedback` signals (`helpful` / `not_helpful` / `stale` / `wrong`) keyed by page *version*, with an optional sanitized reason and `salience_after`. Source of truth for the derived `pages.salience`; the lint pass reads unresolved stale/wrong rows joined against `is_latest = 1`, so a rewrite retires the finding. |
|
||
| `page_access` | One row per latest page and qualified operator identity. Supplies the optional access-breadth retention term without changing the existing shared access counter. |
|
||
| `page_evidence` | V63 append-only record of what produced or reaffirmed each page version — consolidation cites the `session` it ran on, written in the page-upsert transaction and cascaded on purge. Feeds a read-time, zero-LLM **belief-strength `confidence`** (`ai_memory_store::belief`, B1 / `docs/design-hindsight-borrowings.md` P2): a bounded `[0.0, 0.95]` function of distinct supporting sessions (breadth, not raw count), recency of the newest sighting, and live `contradicts` count. Exposed as `confidence` + `evidence_count` in `memory_query(explain=true)`, as `evidence_rows` in `memory_status`, and used to order the opt-in `settled_first` briefing — all **ranking-inert**. It folds into `PageAuthority` only behind `[retrieval] belief_authority_weight` (**default `0.0`/OFF, R2-gated**), as one bounded factor inside the `[0.55, 1.50]` clamp; a supersession always wins regardless of evidence (a superseded version is never boosted). |
|
||
| `agent_messages` | V64 cross-project message inbox/queue (`docs/agent-messaging.md`). Directed, claim-once mail from a sender coordinate to a recipient coordinate; `pending`→`claimed` (popped exactly once, the handoff compare-and-set) or `pending`→`cancelled` (sender retracts). The one table that crosses per-project isolation, so reads are keyed by the recipient coordinate (inbox) or sender coordinate (outbox); `from_owner_user`/`claimed_by_user` are attribution only, never a read filter. Recipient inbox depth is capped. `ON DELETE CASCADE` on both coordinate pairs. |
|
||
| `pages.compacted_at` | V65 nullable A2 tier-down marker (`docs/design-memory-aging.md` §A2). Set when the forget-sweep extractively compacts a cold episodic page (opt-in `[decay] compact_cold_episodic`) instead of evicting it; derived at the single page-upsert choke point from a `compacted: true` frontmatter mirror, so the marker and compacted body land in one transaction. Additive `ADD COLUMN`, no backfill — populated lazily by the sweep. The sweep and curator skip a marked page so it is never re-compacted, re-evicted, or re-reported as cold. Reversible: the full pre-compaction body stays in git + the supersession chain. |
|
||
| `client_activity` | Server-wide MCP tool-call counters split into reads/writes and bucketed by UTC day. The MCP request choke point flushes buffered calls on a one-minute background interval; failed batches retry from bounded memory. Each day stores at most 128 sanitized client labels plus `other`, so an untrusted `clientInfo.name` cannot create traffic-proportional rows. |
|
||
| `auto_improve_proposals` | Staged learning and maintenance edits with immutable target snapshots and append-only decision events. Pending-target uniqueness is scoped by the qualified staging identity; unattributed proposals retain the historical shared bucket. |
|
||
| `entities`, `entity_page_links` | V38 noun index derived from canonical frontmatter. Names are normalized and unique per project; links target immutable page versions while retrieval filters to the latest version. Scope-pairing triggers prevent cross-project links. Powers the fourth RRF retrieval stream. |
|
||
| `audit_log` | Every mutation, addressable by `at DESC`. |
|
||
|
||
**Memory tiers (M8 policy):**
|
||
|
||
| Tier | Lifetime | Decay |
|
||
|---|---|---|
|
||
| Working | Current session only | Hard-drop on session end (kept in `observations` for forensics) |
|
||
| Episodic | 30d hot → 180d cold → evict (or tier-down, opt-in A2) | `salience · exp(−λΔt) + σ · log(1+access_count) · exp(−μ · days_since_access) · (1 + breadth_weight · ln(1 + max(distinct_actors−1, 0)))` |
|
||
| Semantic | Indefinite | None - only supersedeable via M7 LLM rewrite |
|
||
| Procedural | Indefinite | Frequency-decay if not re-observed |
|
||
|
||
`λ` is the scalar `[decay] lambda` by default, but can be set per tier via
|
||
`[decay.half_life_days]` (a half-life in days per tier, converted to
|
||
`λ = ln(2)/days`); an unset tier uses the scalar, so the default reproduces
|
||
today's single-λ scores exactly.
|
||
|
||
**Extractive tier-down (A2, opt-in).** With `[decay] compact_cold_episodic =
|
||
true`, the sweep runs a compaction pass *before* the decay-eviction pass: a cold
|
||
episodic page that is not already compacted is rewritten through the wiki layer
|
||
to keep its L0 `abstract:`, an L1 first-paragraph summary, and an L2 regex-mined
|
||
keep-token set (paths, URLs, code spans, error codes, identifiers), dropping the
|
||
prose — instead of being tombstoned. The rewrite supersedes the prior version,
|
||
so the full body stays reachable (git + supersession chain; `restore-page`
|
||
recovers it). The `V65` `pages.compacted_at` marker makes a compacted page
|
||
terminal for the decay pass (never re-compacted, never re-evicted, never
|
||
re-reported as cold). Zero-LLM, off by default; the R2 recall no-regression
|
||
proof gates any future default-on.
|
||
|
||
**LLM "dream" pass (B2/B3/B4, opt-in LLM, OFF by default, R2-gated before
|
||
default-on).** Where A3 collapses near-duplicate cold clusters *extractively*
|
||
(zero-LLM, keep-token union), the dream pass
|
||
(`ai-memory-consolidate::dream::run_dream_pass`) hands each cold cluster to the
|
||
configured provider to be rewritten into ONE coherent page — the prose-coherent
|
||
merge extraction cannot do. It reuses A3's clustering math (`adaptive_eps` /
|
||
`dbscan`) over the *same* bounded cold set the forget sweep materialises
|
||
(`sweep::materialize_cold_set`, invariant #2). It runs only when `[dream]
|
||
enabled` is set AND a provider AND an embedder are configured; a provider-less
|
||
store keeps the zero-LLM A3 path (invariant #13). It **never deletes a source**
|
||
(invariant #16): the highest-retention member is rewritten and every merged-away
|
||
member is *superseded* with a merge-note stub, so the full pre-merge body stays
|
||
reachable (git + supersession chain; `restore-page` recovers it), and
|
||
`page_evidence` (`reconsolidation` + `b2_dream:<id>`) records which members fed
|
||
each merge (the hallucinated-merge guard). The rewrite routes through the gated
|
||
apply path — `preflight_admission(Consolidate)` before the LLM, then
|
||
`Wiki::apply_batch` (single-writer actor, invariant #2) — with **`dry_run`
|
||
first** (returns the plan, calls neither the LLM nor the writer) and
|
||
**JSON-schema structured output only** (invariant #7). Scheduling (B3, in
|
||
`serve.rs`) starts a run only after `[dream] idle_window_secs` of no client
|
||
activity (read from the tool router's shared `ActivityClock`) and **cancels it
|
||
the moment activity resumes** — a cheap `DreamCancel` flag polled between
|
||
clusters — bounded to `max_clusters_per_run` clusters per run (invariant #5).
|
||
Work is ordered **surprisal-first** (B4): the most-novel clusters, farthest in
|
||
embedding space from the nearest existing (non-cold) page, first. Every run
|
||
returns an observable `DreamReport` so a bad run is never silent. No new
|
||
migration (reuses `page_evidence` + supersession); no new MCP tool.
|
||
|
||
Pinned pages (`pinned: true` in frontmatter) are exempt from all
|
||
decay paths. Pages under `_slots/` are pinned automatically and surfaced
|
||
in briefing/explore snapshots as tiny editable memory slots. Slot pages
|
||
may declare a write regime with `slot_kind: state` or
|
||
`slot_kind: invariant`; omitted means `state` for backwards
|
||
compatibility. Use `state` for mutable working context such as current
|
||
focus and pending items. Use `invariant` for high-resistance project
|
||
context, identity, rules, or user preferences; consolidation should not
|
||
rewrite an existing invariant slot unless new observations directly
|
||
contradict specific existing content.
|
||
|
||
Shared servers may opt into `[slots] per_user = true`. Engine and MCP slot
|
||
writes then use a bounded namespace derived from the authenticated
|
||
`IdentityKey`; session briefs and consolidation prompts include shared slots
|
||
plus the caller's namespace. Existing unnamespaced slots stay shared and the
|
||
default remains off. Exact wiki reads and searches are deliberately unchanged:
|
||
this boundary limits prompt injection, not page access.
|
||
|
||
## Cross-project links
|
||
|
||
Pages normally link within their own project (`[[decisions/0001.md]]`, or a
|
||
`label` pointing to `../gotchas/x.md`). A wikilink can also name another project
|
||
so that dependencies between projects become explicit edges in the graph:
|
||
|
||
* `[[project:path.md]]` — a sibling project in the same workspace.
|
||
* `[[workspace/project:path.md]]` — a project in another workspace.
|
||
|
||
The parser (`ai-memory-wiki::extract_links`) yields a `LinkTarget
|
||
{ workspace, project, path }`; the store resolves it against the named
|
||
project's latest page and records the scope in `links.to_workspace` /
|
||
`links.to_project` (NULL = the source's own project, the common case).
|
||
Resolution is deferred-safe: a link to a page that does not exist yet
|
||
stays `to_page_id = NULL` and is repointed by
|
||
`refresh_incoming_links_for_path` when that page later lands — across
|
||
projects, not only within one.
|
||
|
||
Because `to_page_id` is a global id and `ReaderPool::page_links` joins by
|
||
id without a project filter, a resolved cross-project link surfaces as a
|
||
backlink on its target for free; `RelatedPage` carries the source's
|
||
`workspace` / `project` so the dependency is labelled and navigable. This
|
||
is what turns the per-project wikis into one dependency graph (see also
|
||
the `memory_lint` dangling-ref check, the briefing dependents counts, and
|
||
the `/api/v1/graph` endpoint).
|
||
|
||
## Crate layout
|
||
|
||
```
|
||
crates/
|
||
├── ai-memory-core/ domain types, errors, ids. NO IO.
|
||
├── ai-memory-store/ SQLite + writer actor + reader pool + decay math.
|
||
├── ai-memory-wiki/ atomic markdown writes, file watcher, git.
|
||
├── ai-memory-mcp/ rmcp transport + tool router.
|
||
├── ai-memory-hooks/ payload schemas, sanitiser, /hook ingress.
|
||
├── ai-memory-llm/ provider auth boundary + LlmProvider / Embedder traits.
|
||
├── ai-memory-consolidate/ Karpathy ingest / lint / sweep / auto-improve pipeline.
|
||
├── ai-memory-workstream/ read-only native transcript + launch adapters.
|
||
└── ai-memory-cli/ `ai-memory` binary entry point + thin HTTP subcommands.
|
||
```
|
||
|
||
Each crate has a single responsibility and exposes a typed API. No
|
||
circular deps. Inter-crate boundaries enforce the cross-cutting
|
||
invariants below.
|
||
|
||
## MCP tool surface (23 tools)
|
||
|
||
| Tool | Hint | Purpose |
|
||
|---|---|---|
|
||
| `memory_query` | read-only | FTS5 + entity-match + graph RRF + optional vector RRF search, followed by bounded kind/tier/pinned/tag authority adjustment and raw fallback. Bumps access counters for page hits. Defaults to the current project; single-project calls (project implicit or named with `workspace`+`project`) also union the reserved `_global` preferences scope as `global_scope_hits`, and only an explicit multi-`scopes` set opts out (#930); `scopes` searches named sibling projects; `global=true` searches every project at once (each hit annotated with its workspace + project). With `AI_MEMORY_RERANKER=llm`, project/scopes candidate pools are fused before at most one final LLM relevance pass; query/title/snippet data is bounded and JSON-encoded, and any timeout, provider error, invalid/incomplete score set, or four-call concurrency saturation preserves the adjusted order. The distinct `global=true` FTS-only ranker and supplemental global-preference hits are not reranked. `explain=true` attaches per-hit `score_details` (per-stream ranks, matched entities, raw FTS/cosine/entity inverse-frequency scores, RRF contributions, graph provenance including the typed edge kind (`causes`/`fixes`/`contradicts`) a neighbour was reached by, the page's evidence count, authority multiplier, and optional rerank score) to project/scopes hits plus a top-level `streams_active` list. The global FTS-only ranker reports its active stream without per-hit details. `include_expired=true` also returns TTL-expired pages. `include_superseded=true` also returns superseded (non-latest) page versions across the FTS/entity/vector/graph streams, each hit labelled `superseded: true` (the current version is never marked); default-off is byte-identical to the latest-only behaviour, and `global=true` / `as_of` are unaffected. `pin_first=true` prepends the project's bounded pinned latest pages (`ReaderPool::list_pinned_pages`, cap 10) ahead of the fused hits, deduped by page id (a pinned page that also matches appears once, marked `pinned: true`) and re-truncated to the requested limit; it applies to single-project searches (default or `workspace`+`project`), is ignored on `scopes`/`global`/`as_of`, and default-off is byte-identical. `answer=true` (opt-in, off by default) additionally synthesizes a cited natural-language answer over the top hits via the configured LLM provider (`complete_structured`, JSON-schema `{ answer, citations }`), attached as `answer: { text, citations }`; with no provider configured it returns the hits plus an `answer_unavailable` note instead of erroring, and with `answer` unset/`false` no provider is accessed and the response is byte-identical (invariant #13). It applies to the normal single-project/`scopes` path; `global`/`as_of` ignore it. Answer quality is not yet eval-validated. An optional `reasoning` tier (`minimal` (default) / `low` / `medium` / `high` / `max`) tunes the synthesis effort: `ChatRequest` carries no per-request reasoning field (the provider-level `reasoning_effort` is fixed at construction from config), so the tier maps to a per-tier max-token budget scaled off the path's base (answer base 2 000; `minimal` = 1x = byte-identical, `low` 1.5x, `medium` 2x, `high` 3x, `max` 4x). The tier is inert unless the `answer` LLM path runs (invariant #13); an unknown value is rejected by the schema (invariant #7). |
|
||
| `memory_recent` | read-only | Most-recently-updated `is_latest=1` pages. |
|
||
| `memory_read_page` | read-only | Fetch the FULL body of a single wiki page by `path` or by top FTS5 hit for a `query`; optional `workspace` + `project` targets a named sibling workspace/project. Use when an agent needs more than the 24-word snippets from `memory_query`. `include_related=true` also walks the link graph outward from the page (bounded BFS reusing the `page_links` primitive per node: default 1 hop, hard cap 3, global visited set for dedup/cycle-safety, total-node cap 50, cross-project aware) and returns a `related` array of reachable pages, each labelled with its `depth` (hop distance) and `direction` (`link`/`backlink`); default-off is byte-identical (no `related` field). |
|
||
| `memory_read_session_observations` | read-only | Page through ONE session's raw hook observations (`ObservationRecord` with full sanitized body, capped per row by `body_max_chars`), restricted to the rows that landed in the resolved scope and to sessions the caller may see; `total` and `elided_other_scope` report the in-scope count and the rows the session left in another project. `session_id` omitted reads the latest completed visible session. |
|
||
| `memory_status` | read-only | Counts, paths, version, plus the `scope` that answered: `workspace`, `project`, and `resolved_by` (`explicit`, `session`, `shared_slot`, `startup_seed`, `default`, `default_after_mismatch`). Unscoped reads resolved by `startup_seed` or `default_after_mismatch` also log a server warning. |
|
||
| `memory_briefing` | read-only | Structured counts/activity/rules/slots/recent snapshot. Project-scoped snapshots also carry a bounded `pinned` list (up to 10) of the project's pinned latest pages (`pinned = 1`, newest first) as standing SessionStart hot-context — distinct from `slots`, which is keyed by the `_slots/` path prefix; empty and omitted from JSON when the project has no pins, so the default shape is unchanged. Opt-in `settled_first: true` leads with up to 8 of the project's highest-standing `rule`/`decision` pages, ordered by evidence count then recency; off by default. |
|
||
| `memory_explore` | read-only | LLM prose digest over the briefing snapshot, degrading to JSON without a provider. An optional `reasoning` tier (`minimal` (default) / `low` / `medium` / `high` / `max`) scales the digest's max-token budget off its base (16 000; same 1x/1.5x/2x/3x/4x mapping as `memory_query`); inert on the no-provider briefing-only path, and `minimal` is byte-identical. |
|
||
| `memory_handoff_begin` | destructive | Open an owner-scoped handoff for the next agent; `shared=true` deliberately publishes it to the project. Optional `workspace` + `project` targets a named sibling workspace/project. |
|
||
| `memory_handoff_list` | read-only | List open own/shared handoffs with inspectable body and identity fields; does not claim or expire. Root-only `any_owner=true` recovers across operators. Optional `workspace` + `project` targets a named sibling workspace/project. |
|
||
| `memory_handoff_accept` | destructive | Fetch + ack an open own/shared handoff. Pass `handoff_id` from `memory_handoff_list` to claim that exact row; omitting it still claims the latest eligible open handoff (automatic handoffs are cwd-matched). Root-only `any_owner=true` recovers across operators. Optional `workspace` + `project` targets a named sibling workspace/project. Returns `handoff` plus `status`: `claimed` (this call won it), `consumed_by_hook` (the calling session's own SessionStart claimed it, so it is already in context; needs the session id the hook claimed under, i.e. a session-aware client), or `none_pending`. |
|
||
| `memory_handoff_cancel` | destructive | Mark an exact visible open handoff id expired when it was created by mistake; root-only `any_owner=true` recovers across operators. |
|
||
| `memory_message_send` | destructive | Send a directed cross-project message into another project's inbox (V64). Requires `to_workspace` + `to_project`; the recipient must already exist (fail-closed, never created). Body is secret-scrubbed and size-capped. The one tool that crosses project isolation on purpose. |
|
||
| `memory_message_list` | read-only | List pending mail for this project — `box="inbox"` (poppable, default) or `box="outbox"` (cancellable). Bodies are untrusted cross-project input. |
|
||
| `memory_message_pop` | destructive | Claim ONE inbox message exactly once (oldest, or a specific `message_id`); returns it fenced as untrusted input with sender provenance, or `null` when empty. |
|
||
| `memory_message_cancel` | destructive | Retract a pending sent message by `message_id`, or clear the whole outbox when omitted. Scoped to the sender project. |
|
||
|
||
`memory_handoff_list` is the inspect-without-claim path for clients that cannot inject SessionStart stdout. `memory_handoff_cancel` needs an exact id. `ai-memory handoffs` lists the open
|
||
handoffs for a project, oldest first, with their ids — read-only, and
|
||
content-free (identity, provenance and age, never the summary body). Automatic
|
||
expiry deliberately spares manual and sibling-directory handoffs, so a
|
||
long-lived entry appearing there is that policy working rather than a fault.
|
||
|
||
| `memory_consolidate` | destructive | LLM-driven page rewrite. `multi_page=true` for atomic fan-out; an update whose path names an existing pinned page is skipped (`_slots/` excepted). Omitting `session_id` (or sending a blank one) consolidates the latest completed session in the resolved project; the same omission on a project with none fails as `no completed session in <scope>`. Consolidation prompts append the target project's active reserved `_prompts/consolidation.md` body as sanitized, 2,000-character-capped, JSON-encoded, untrusted advisory preferences; TTL-expired pages are ignored and a per-call `instructions` argument overrides the page for one call. Both system prompts keep schema, evidence, disclosure, tool-use, and output rules authoritative. |
|
||
| `memory_feedback` | write | Record a quality signal for one page by exact `path`: `helpful`/`not_helpful` step `pages.salience` for sweep-eligible episodic pages, while `stale`/`wrong` floor salience and surface any current page as a `feedback_flagged` lint finding. Never deletes; the path resolves to the current version in the transaction, so a later rewrite clears it. Retrieved content never authorizes feedback by itself. |
|
||
| `memory_auto_improve` | write | Manually review a completed session and apply or stage validated wiki edits through the auto-improvement approval path. Without a session ID, selects the newest completed session with no persisted auto-improvement run so repeated calls advance through preflight skips; an explicit ID remains rerunnable. The server also schedules review for new sessions; `[auto_improve] require_approval = true` leaves proposals pending for manual review. |
|
||
| `memory_write_page` | destructive | Write durable wiki knowledge when the user explicitly asks to remember/annotate it. `scope: "global"` writes into the reserved `_global` preferences scope; optional `expires_at` sets an RFC3339 or date-only TTL. |
|
||
| `memory_delete_page` | destructive | Delete a single page by exact `path`. Fires the admission chain (op=delete); idempotent. |
|
||
| `memory_forget_sweep` | destructive | Retention pass: evict cold pages through the wiki layer, purge aged tombstone ancestry, and hard-delete TTL-expired pages. `dry_run=true` for preview. |
|
||
| `memory_lint` | destructive | Rule-based + LLM contradiction findings → `wiki/_lint/`. Also runs a **zero-LLM contradiction detector** (design-memory-aging.md A5): cold semantic/procedural pages whose already-stored embeddings sit in the `contradiction_band_min`–`contradiction_band_max` cosine-similarity band (default 0.4–0.75; "same topic, not a near-duplicate" — at/above the max is A3 dedup, below the min unrelated) get an advisory `contradiction` finding with newer-wins timestamp advice. Bounded (one embeddings load over the capped cold set, capped findings, deterministic); a clean no-op with no embedder configured; advisory-only — never deletes/edits/supersedes a page and persists no edge (invariants #13, #16, #2), so no migration. On a single-language or single-domain store, background similarity between unrelated pages already sits well above the default floor, so the band measures domain proximity more than conflict and produces noisy findings — raise `contradiction_band_min` (`config.toml` or `AI_MEMORY_CONTRADICTION_BAND_MIN`) for such a store. |
|
||
| `memory_install_self_routing` | read-only | Return the canonical slim routing snippet plus managed Agent Skill payloads and target hints for CLAUDE.md / AGENTS.md installs. |
|
||
|
||
`memory_briefing`, `memory_explore`, `memory_write_page`,
|
||
`memory_install_self_routing`, `memory_read_page`,
|
||
`memory_read_session_observations`, `memory_delete_page`,
|
||
`memory_handoff_cancel`, `memory_auto_improve`, and `memory_feedback`
|
||
post-date the original "narrow on purpose" cut (§10 of
|
||
`design-decisions.md`): briefing/explore separate the structured vs.
|
||
prose halves of "what's going on", `memory_write_page` covers explicit
|
||
durable annotations without abusing single-use handoffs,
|
||
`memory_install_self_routing` exists for the meta case where the agent
|
||
must re-write its own routing rules into a project's `CLAUDE.md` /
|
||
`AGENTS.md` and install the companion managed Agent Skills into
|
||
`.claude/skills` or `.agents/skills`, `memory_read_page` complements
|
||
`memory_query` for the "I need the full page, not a snippet" case
|
||
(e.g. opening a decision page end-to-end),
|
||
`memory_read_session_observations` opens the raw evidence behind a compiled
|
||
page or a raw hit (one session, in scope, paged and body-capped) so an agent
|
||
can audit what the hooks actually captured, `memory_auto_improve` exposes a
|
||
safe default-on learning review through the same approval/write path as
|
||
pending writes, and `memory_delete_page` is the exact-path destructive pair
|
||
needed by admission-aware mirrors. `memory_handoff_cancel` is the safety valve
|
||
for mistaken handoff creation. `memory_feedback` implements the
|
||
"finer-grained reinforcement beyond access counts" P2 item from
|
||
`prior-art-implementation-findings.md`: it cannot ride on a read tool
|
||
without conflating read and write semantics, and the access counter it
|
||
supplements cannot tell "this page answered the question" from "this page
|
||
wasted a read". The narrow-surface discipline still holds —
|
||
every new tool has to earn its slot — but the count is now 23, not 10.
|
||
|
||
The managed Agent Skills are a narrow prompt-packaging exception to the
|
||
otherwise wiki-centered architecture. They are static `SKILL.md` files that
|
||
teach agents when to call ai-memory MCP tools; they are not durable wiki pages,
|
||
not auto-improvement output, and not a runtime skill router inside ai-memory.
|
||
|
||
MCP parameter aliases are intentionally sparse: `memory_query.query` accepts
|
||
`q|search`, and limit fields accept `n` / `top_k` where shipped. Project and
|
||
cwd parameters use their canonical names.
|
||
|
||
Claude Code's optional session-aware MCP registration is a transport adapter,
|
||
not a second tool implementation. `ai-memory mcp-bridge` serves the upstream
|
||
tool catalogue over local stdio, delegates tool calls to the configured HTTP
|
||
server through rmcp's client transport, and injects the inherited
|
||
`CLAUDE_CODE_SESSION_ID` as `X-Memory-Actor-Session-Id`. The server therefore
|
||
keeps the same auth, scope resolver, and tool handlers as direct HTTP clients.
|
||
The adapter fails closed without a Claude session id and is installed only by
|
||
the explicit `install-mcp --client claude-code --session-aware` option.
|
||
|
||
## HTTP authentication classes
|
||
|
||
The process separates four active credential classes from one transitional
|
||
browser compatibility path:
|
||
|
||
| Class | Wire | Authorizes |
|
||
|---|---|---|
|
||
| Human password | `POST /auth/login` body | Session issuance only |
|
||
| Web session | `ai_memory_session` cookie + CSRF | `/auth/me`, `/admin/*`, `/api/v1/*` by `AuthLevel`; never `/mcp` or hooks |
|
||
| Recovery | `POST /auth/recovery` body | Root password reset; no session |
|
||
| API key | `Authorization: Bearer` (`aim_`, root `AI_MEMORY_AUTH_TOKEN`, or external `amk_`) | Machine APIs; never a web session |
|
||
| Deprecated browser compatibility | HTTP Basic root bearer, then HttpOnly `ai_memory_auth` cookie | GET-only browser routes until any human password or completed bootstrap exists; never machine routes |
|
||
|
||
The deprecated Basic/cookie path stops immediately when human auth becomes
|
||
active; restart is not required. `/web` SPA HTML is public static; the builtin
|
||
wiki browser and JSON APIs stay behind the route class above.
|
||
|
||
## CLI subcommand surface
|
||
|
||
```
|
||
init status run
|
||
show continue resume
|
||
workstreams rename-workstream workstream-search
|
||
audit-contamination search read-page
|
||
write-page delete-page serve
|
||
reset backup restore
|
||
reindex install-hooks hook
|
||
install-mcp commit checkpoints
|
||
restore-page llm-test forget-sweep
|
||
lint curator auto-improve-report
|
||
auto-improve finalize-session pending-writes
|
||
embed generate-auth-token setup-agent
|
||
bootstrap install-instructions install-skills
|
||
reorg purge-project rename-project
|
||
move-project move-session uninstall
|
||
upgrade auth user
|
||
completions handoffs purge-session
|
||
compact api-key export-okf
|
||
message doctor backfill
|
||
project reclaim-ledger-versions repair-backfill-timestamps
|
||
server
|
||
```
|
||
|
||
Run `ai-memory --help` for the full tree.
|
||
|
||
`auto-improve-report` is read-only by default; `--stage` creates one pending
|
||
telemetry report page for audit/approval without staging learning-memory edits.
|
||
|
||
`reclaim-ledger-versions` drops the superseded versions of the raw hook event
|
||
ledger that the pre-2.1.1 indexer left behind (#660). It is a dry run unless
|
||
`--confirm` is passed. A path is only a candidate when its *content* opens
|
||
with a hook log entry — the same
|
||
`ai_memory_core::log_ledger::body_opens_with_log_ledger` gate the indexer
|
||
(#660), the OKF conformance migration (#669) and the bundle export (#748) use
|
||
— so a real page named `log-2026-09.md` keeps its whole version chain. Only
|
||
`is_latest=0 AND superseded_at IS NULL` rows are eligible, so rows a decay
|
||
tombstone owns stay with `forget-sweep`. Derived FTS/entity/vector/link rows
|
||
go with the page through the existing `ON DELETE CASCADE`s, and the FTS delete
|
||
trigger is stood down for the bulk delete (its DDL is read back from
|
||
`sqlite_master` and re-executed) so the cleanup does not re-tokenize tens of
|
||
gigabytes of ledger body row by row; `pages_fts` is then rebuilt wholesale.
|
||
`--compact` additionally `VACUUM`s to return the bytes, at the cost `compact`
|
||
documents.
|
||
|
||
## Cross-cutting invariants
|
||
|
||
Carved in M0/M1; every milestone has to respect them. Each comes from
|
||
a documented prior-art bug; cite the source when reviewing changes
|
||
that touch the relevant area.
|
||
|
||
1. **One config-read path.** `Config::load()` called once at startup.
|
||
No `std::env::var` outside it. (agentmemory #456 / #469.)
|
||
2. **Single-writer SQLite actor.** All writes go through one `mpsc`
|
||
channel to one dedicated OS thread. (cognee #2717.)
|
||
3. **Indexes commit in the same transaction as the data.** No
|
||
background-task-indexing-after-return. (basic-memory #763 / #578.)
|
||
4. **Typed 3-tuple identity** (`workspace_id`, `project_id`, path)
|
||
in every domain row from day one. (basic-memory #783 / #834.)
|
||
5. **Hooks are fire-and-forget.** Hook scripts hard-timeout at
|
||
≤200 ms; server returns 202 immediately or 429 when saturated.
|
||
(agentmemory #221 / #143.)
|
||
6. **Privacy strip is a typed boundary.** `Sanitized<NewObservation>`
|
||
has no other constructor than `sanitize()`. (design-decisions §14.)
|
||
The opt-in assistant/Stop excerpt (#196) enters through this same
|
||
boundary: the client sanitizes it before it reaches the wire, and the
|
||
server re-scrubs it here with its configured patterns before the write.
|
||
7. **JSON-schema structured outputs only.** Native provider JSON
|
||
modes; no XML, no Instructor wrapping. (agentmemory #492 / #539,
|
||
cognee #2840.)
|
||
8. **`{provider, model, dim}` denormalised next to every embedding.**
|
||
Warn and ignore stale vectors on mismatch until re-embedding completes.
|
||
(agentmemory #469.)
|
||
9. **Live-process check before direct-disk lifecycle ops.** `ai-memory reset`,
|
||
`restore`, `reindex`, and `uninstall --purge-data` consult `sysinfo`; the
|
||
uninstall guard is conditional on `--purge-data`. `backup` is a thin HTTP
|
||
client instead: the server snapshots SQLite with its online backup API while
|
||
the writer remains live. (basic-memory #765.)
|
||
10. **Atomic file writes** (tmp + rename + fsync). Watcher ignores
|
||
own writes by filename prefix.
|
||
11. **Absolute canonical data dir** default; logged loudly on
|
||
startup. (agentmemory #303.)
|
||
12. **No global singletons / `lazy_static` configs.** All deps
|
||
explicit. (cognee #2228.)
|
||
13. **Zero-LLM default path.** LLM has opt-in via env. The
|
||
system works without any provider configured.
|
||
14. **Provider auth resolves before provider construction.** Native
|
||
provider clients consume typed `ProviderAuth` material; they never
|
||
read env vars directly. Token-backed providers receive explicit
|
||
auth-file paths / env-derived token material through that boundary,
|
||
then own provider-specific refresh and persistence.
|
||
15. **Tracing subscribers explicitly filter their own module.**
|
||
No feedback loops. (agentmemory #519.)
|
||
|
||
## Configuration (`config.toml`)
|
||
|
||
Lives at `<data_dir>/config.toml`. All values overridable by env vars
|
||
prefixed `AI_MEMORY_*`.
|
||
|
||
```toml
|
||
bind = "127.0.0.1:49374"
|
||
log_level = "info" # default filter also pins `rmcp=warn` (the MCP SDK's
|
||
# per-request info logs) and drops the 30s reconcile
|
||
# summary to debug (#894). Restore either via log_level
|
||
# (e.g. "info,rmcp=info", "debug") or RUST_LOG;
|
||
# `tracing_appender=warn` stays forced (feedback-loop guard)
|
||
tcp_keepalive_secs = 60 # idle time before TCP keepalive probes an accepted `serve`
|
||
# connection; reaps sockets left half-open by a dead peer
|
||
# (laptop sleep, VPN flap) that would otherwise leak fds
|
||
# until EMFILE (#792). 0 disables keepalive. Env:
|
||
# AI_MEMORY_TCP_KEEPALIVE_SECS
|
||
contradiction_band_min = 0.4 # `memory_lint`'s A5 zero-LLM contradiction band
|
||
contradiction_band_max = 0.75 # (lower/upper cosine-similarity edge). The band is a
|
||
# fixed absolute cosine value, but a single-language or
|
||
# single-domain store's background similarity sits well
|
||
# above the general-purpose default, so the default band
|
||
# ends up measuring domain proximity rather than conflict
|
||
# and produces noisy findings — raise `contradiction_band_min`
|
||
# for such a store. Must satisfy 0.0 <= min < max <= 1.0.
|
||
# Env: AI_MEMORY_CONTRADICTION_BAND_MIN /
|
||
# AI_MEMORY_CONTRADICTION_BAND_MAX
|
||
|
||
# Capture / launch UX (all default-on where noted). Each has an AI_MEMORY_* env
|
||
# override (AI_MEMORY_CAPTURE_ASSISTANT / AI_MEMORY_BACKFILL_ON_START /
|
||
# AI_MEMORY_RUN_AUTOWIRE / AI_MEMORY_CLAUDE_TRUE_YOLO).
|
||
capture_assistant = false # server-side opt-in: honor a Claude Code / Codex
|
||
# client's sanitized assistant-final-message marker
|
||
# on Stop (#196). Client half is baked separately by
|
||
# `install-hooks --capture-assistant`.
|
||
backfill_on_start = true # on first SessionStart in a brand-new (empty) project,
|
||
# import that project's existing local harness history
|
||
# once so hooks-mid-project isn't amnesiac. Only ever
|
||
# bootstraps an empty project; hard-capped. `ai-memory
|
||
# backfill` runs it by hand.
|
||
run_autowire = true # `ai-memory run <harness>` auto-installs that harness's
|
||
# hooks + MCP on first launch if missing (idempotent,
|
||
# one-time per harness+version+install location).
|
||
# Also `--no-autowire`.
|
||
claude_true_yolo = false # opt-in: on a Claude `ai-memory run --yolo`, also
|
||
# silence the residual `--dangerously-skip-permissions`
|
||
# prompts (rm timeout/confirmation, PowerShell rm deny)
|
||
# and force `bypassPermissions` via `--settings`.
|
||
# Claude-only, no-op for every other harness. Also
|
||
# `--true-yolo`. See
|
||
# docs/design-yolo-safety-ai-jail.md.
|
||
release_base_url = "" # override the GitHub releases base URL that `ai-memory
|
||
# upgrade` checks and downloads from (#801). Empty =
|
||
# https://github.com/akitaonrails/ai-memory/releases.
|
||
# For hermetic tests / mirrors, not day-to-day installs.
|
||
# Env: AI_MEMORY_RELEASE_BASE_URL.
|
||
|
||
[maintenance] # scheduled server jobs (run outside hook latency)
|
||
enabled = true # master switch for the scheduled jobs below
|
||
forget_sweep_interval_secs = 86400 # retention forget sweep; 0 disables. Cadence persists
|
||
# across restarts; overdue work starts after a bounded delay
|
||
lint_interval_secs = 86400 # rule-based wiki lint; 0 disables (same persistence)
|
||
embedding_backfill_interval_secs = 0 # embedding backfill; 0 = off (may call a paid provider)
|
||
reconcile_tombstones_deleted_pages = false
|
||
# opt-in (#929/#964): the 30s reconcile pass soft-tombstones
|
||
# (is_latest=0 + superseded_at — never a filesystem write) an
|
||
# OKF-imported content page whose file has been missing on
|
||
# two consecutive passes, behind a circuit breaker and with
|
||
# session pages excluded. OFF = byte-identical to pre-2.5
|
||
# behavior (deletions still need `ai-memory delete-page`).
|
||
# Docs: docs/okf.md, docs/install.md.
|
||
|
||
[decay] # M8 retention params
|
||
lambda = 0.02 # ↓ to forget less aggressively (fallback λ)
|
||
sigma = 0.6 # ↑ to reward query-hits more
|
||
mu = 0.04 # ↑ if recent hits should count more
|
||
cold_threshold = 0.20 # below this → remove file + retain tombstone
|
||
hard_delete_after_days = 180
|
||
breadth_weight = 0.0 # opt-in reward for distinct operators
|
||
observation_retention_days = 0 # 0 = never prune raw observations
|
||
observation_prune_batch = 5000 # rows per prune transaction
|
||
compact_cold_episodic = false # A2 opt-in: tier-down (compact) a cold
|
||
# episodic page instead of evicting it —
|
||
# keep abstract+summary+keep-tokens, drop
|
||
# prose. Reversible (git + supersession),
|
||
# zero-LLM. false = today's evict behaviour.
|
||
dedup_cold_clusters = false # A3 opt-in: cluster near-duplicate cold
|
||
# episodic pages by embedding (cosine DBSCAN,
|
||
# adaptive eps) and collapse each cluster to
|
||
# one survivor (union of keep-tokens), others
|
||
# superseded with a merge note. Reversible
|
||
# (git + supersession), zero generative LLM,
|
||
# no-op with no embedder. Merge provenance in
|
||
# page_evidence. false = no clustering.
|
||
# dedup_min_pts = 2 # DBSCAN density floor (0 ⇒ default 2)
|
||
# dedup_max_eps = 0.15 # conservative eps ceiling (cosine distance;
|
||
# 0 ⇒ default). Lower = merges less.
|
||
|
||
[decay.half_life_days] # opt-in per-tier retention curves (all keys
|
||
# optional). Half-life in DAYS; converted to
|
||
# λ = ln(2)/days. An omitted key falls back to
|
||
# the scalar `lambda` above, so the default
|
||
# (no keys) is byte-identical to today — no
|
||
# score change or mass-eviction on upgrade.
|
||
# working = 7 # e.g. keep scratch short…
|
||
# episodic = 365 # …and session history long
|
||
# semantic = 180
|
||
# procedural = 90
|
||
|
||
[slots] # optional shared-server injection boundary
|
||
per_user = false # shared + own slots in agent context
|
||
|
||
[consolidation] # LLM consolidation prompt sizing
|
||
max_input_tokens = 100000 # approximate whole-input target; min 6000
|
||
# a flat chars-per-token heuristic, so it
|
||
# UNDER-budgets denser corpora: pt-BR prose
|
||
# and source code tokenize at fewer chars per
|
||
# token than English and can overshoot the
|
||
# provider's real limit by ~40% — lower this
|
||
# (or input_token_safety_margin) for such a corpus
|
||
max_output_tokens = 32000 # provider generation limit; min 1000
|
||
# their sum must fit the model context window;
|
||
# leave headroom for tokenizer variance
|
||
input_token_safety_margin = 0.8 # scales the char budget, (0.0, 1.0]; the
|
||
# 0.8 default buys pt-BR/code headroom
|
||
|
||
[auto_improve] # default-available learning reviewer
|
||
require_approval = false # true leaves proposals pending for review
|
||
min_observations = 8
|
||
min_session_duration_secs = 120
|
||
min_confidence = 0.75
|
||
max_input_tokens = 24000
|
||
max_proposals_per_run = 5
|
||
max_patchable_pages = 8
|
||
patchable_page_prefixes = ["_rules/", "procedures/"]
|
||
max_patchable_body_chars = 8000
|
||
max_edits_per_proposal = 5
|
||
max_edit_content_chars = 4000
|
||
max_changed_chars_per_proposal = 12000
|
||
max_patch_edits_per_run = 8
|
||
max_rejection_context = 50
|
||
rejection_context_days = 180
|
||
max_final_body_chars = 32000
|
||
max_rule_page_tokens = 2000
|
||
max_procedure_page_tokens = 2000
|
||
include_raw_fallback = false
|
||
proposal_actor = "auto_improve"
|
||
pending_path = "_pending/auto-improve"
|
||
|
||
[auto_improve.scheduler] # background review; separate from approval
|
||
enabled = true
|
||
interval_secs = 3600
|
||
max_sessions_per_tick = 1 # per project; scheduler ticks do not overlap
|
||
min_session_age_secs = 600
|
||
experience_every_sessions = 0 # 0 disables the cross-session experience pass
|
||
experience_sessions = 10 # session summaries one experience pass reads
|
||
|
||
[auto_improve.scheduler.experience_entropy_filter] # A4 opt-in; off by default
|
||
enabled = false # true: skip low-information session pages from
|
||
# the experience consolidation pass BEFORE the
|
||
# prompt/eval-gate/apply_batch. Advisory (skip,
|
||
# never delete); zero-LLM. false = no filtering.
|
||
# min_chars = 16 # near-empty floor (non-whitespace chars)
|
||
# min_entropy_bits_per_char = 2.0
|
||
# max_repetition_ratio = 0.7 # 1 - distinct/total tokens above this ⇒ skip
|
||
# repetition_min_tokens = 6 # repetition check applies only above this
|
||
|
||
[retrieval] # opt-in ranking signals; all off by default
|
||
query_intent = false # lexical session-recall routing: queries phrased as
|
||
# "上次 / …的会话 / last time / yesterday" hand session
|
||
# pages back their default kind/tier authority penalty
|
||
session_recall_bonus = 0.25 # extra authority on top of the cancelled penalty;
|
||
# lower it (e.g. 0.15) if rank drift on
|
||
# "之前/上次"-prefixed fact queries matters more
|
||
abstract_vectors = false # fifth RRF stream over page_abstract_embeddings
|
||
# (L0: each page's frontmatter `abstract:` line, embedded
|
||
# by the same backfill as the body)
|
||
belief_authority_weight = 0.0 # fold read-time belief-strength confidence (P2) into page
|
||
# authority as ONE bounded factor inside the [0.55, 1.50]
|
||
# clamp. 0.0 = OFF (default): ranking is byte-identical and
|
||
# no belief query runs. confidence/evidence_count are still
|
||
# exposed in explain regardless (inert). DEFAULT OFF,
|
||
# R2-gated: do not default on without a positive R2 delta.
|
||
|
||
[search.fts] # FTS5 query-preparation tuning (contrast with [retrieval]:
|
||
# this is not a ranking signal). Omit the whole section for
|
||
# byte-identical behaviour on every existing install.
|
||
# stopwords = [] # words dropped from a bare natural-language FTS query
|
||
# before the OR-join. Three states:
|
||
# - key absent (the default): the built-in English list
|
||
# (a/an/and/the/…, ~60 entries) — unchanged behaviour.
|
||
# - stopwords = []: disables the filter entirely, English
|
||
# included.
|
||
# - a non-empty list: REPLACES the default outright with
|
||
# exactly those words (does not extend the English list).
|
||
# Entries are folded to lowercase with Unicode case folding
|
||
# (not ASCII-only — a sentence-initial "É" or all-caps "NÃO"
|
||
# still matches a lowercase entry), and a query token is
|
||
# compared the same way, but diacritics are never stripped:
|
||
# "e" and "é" stay distinct. What matters here is how a query
|
||
# is actually TYPED, not how wiki content is spelled — content
|
||
# matches through the FTS index's own diacritic-folding
|
||
# tokenizer regardless, but this filter only ever sees the
|
||
# literal characters someone typed. So a list for an accented
|
||
# language should include every spelling a user or agent might
|
||
# type, e.g. Portuguese BOTH "e" and "é", BOTH "nao" and "não".
|
||
# The filter also matches whitespace-split raw tokens before
|
||
# any punctuation handling, so an entry never matches a token
|
||
# with attached punctuation ("de," / "que?") — the same
|
||
# limitation English stopwords have always had. Explicit FTS5
|
||
# syntax (quoted phrases, OR/AND/NOT/NEAR, parens) always
|
||
# bypasses this filter, exactly as it does with the built-in
|
||
# list; a bare query made ONLY of configured stopwords keeps
|
||
# them all rather than returning nothing. At most 2000 entries
|
||
# of at most 64 characters each, no internal whitespace
|
||
# (entries are matched against single whitespace-split
|
||
# tokens); anything past that fails startup.
|
||
#
|
||
# Env override: AI_MEMORY_SEARCH_FTS_STOPWORDS as a
|
||
# comma-separated string — the same convention
|
||
# allowed_hosts/cors_allow_origins/auth.trusted_proxy_cidrs
|
||
# use for a Vec<String> — but read once as data in
|
||
# Config::load rather than through the usual `__`-split env
|
||
# layer, because a present-but-blank env var must mean
|
||
# "unset" (leave config.toml's value alone), never an
|
||
# accidental "disable filtering"; the automatic layer merges
|
||
# raw values before any such distinction could be made.
|
||
#
|
||
# Non-English example — a Portuguese-majority wiki, so its
|
||
# own function words (not English's) get filtered, spelling
|
||
# out both accented and unaccented forms someone might type:
|
||
# stopwords = [
|
||
# "a", "o", "as", "os", "de", "da", "do", "das", "dos",
|
||
# "em", "um", "uma", "uns", "umas", "com", "para", "por",
|
||
# "que", "se", "no", "na", "nos", "nas", "e", "ou",
|
||
# "nao", "não", "voce", "você",
|
||
# ]
|
||
|
||
[dream] # B2/B3/B4 opt-in LLM "dream" pass. OFF by default,
|
||
# R2-gated before it may default on. Never deletes a source.
|
||
enabled = false # true starts the scheduled pass — but ONLY if a provider AND
|
||
# an embedder are also configured. A provider-less store keeps
|
||
# the zero-LLM A3 path untouched (invariant #13).
|
||
interval_secs = 3600 # how often the scheduler CONSIDERS a run (0 ⇒ 3600)
|
||
idle_window_secs = 300 # operator must be quiet this long before a run starts; returning
|
||
# activity CANCELS an in-flight run at the next cluster boundary
|
||
# (B3). 0 ⇒ default 300.
|
||
# min_pts = 2 # DBSCAN density floor (0 ⇒ default 2)
|
||
# max_eps = 0.15 # conservative eps ceiling (cosine distance; 0 ⇒ default)
|
||
# max_clusters_per_run = 8 # bounded fan-out per run (invariant #5; 0 ⇒ default 8)
|
||
# min_cold_pages = 2 # events-accrued gate: skip a run below this many cold pages
|
||
```
|
||
|
||
**LLM provider env** (opt-in):
|
||
```
|
||
AI_MEMORY_LLM_PROVIDER anthropic | anthropic-oauth | openai | openai-oauth | codex | copilot |
|
||
gemini | openai-compat | opencode
|
||
AI_MEMORY_LLM_MODEL optional when the provider has a default; e.g. claude-haiku-4-5, gpt-5.4-mini
|
||
ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY / LLM_API_KEY
|
||
AI_MEMORY_LLM_BASE_URL required for openai-compat (Ollama, vLLM); optional override for
|
||
opencode (defaults to the Go endpoint, set
|
||
https://opencode.ai/zen/v1 for Zen's catalogue).
|
||
Applies to any provider: naming ai-memory is how an
|
||
operator says a vendor endpoint is proxied on purpose
|
||
LLM_BASE_URL the unprefixed cross-tool convention, accepted for
|
||
openai-compat and opencode only. Providers with a
|
||
fixed vendor endpoint (anthropic, openai, gemini,
|
||
the OAuth backends, copilot) ignore it and log why —
|
||
an operator's leftover Ollama URL must not silently
|
||
rewrite every Gemini request into a 404
|
||
AI_MEMORY_LLM_COMPAT_STRICT true by default; false disables response_format=json_schema
|
||
AI_MEMORY_LLM_TIMEOUT_SECS per-request timeout for chat providers; 300 by default
|
||
AI_MEMORY_LLM_REASONING_EFFORT optional reasoning/thinking effort
|
||
(none|minimal|low|medium|high|xhigh|max|ultra|persistent)
|
||
mapped per provider: OpenAI `reasoning_effort`,
|
||
OpenRouter `reasoning.effort`, xAI Grok
|
||
`reasoning_effort`, Anthropic `output_config.effort`,
|
||
Codex `reasoning.effort`. Gemini and Copilot
|
||
ignore the key. Host-unsupported values are
|
||
clamped to each provider's published enum.
|
||
AI_MEMORY_LLM_HEADERS optional extra HTTP headers on every chat request, as
|
||
comma-separated `Name=Value` (or `Name: Value`) entries;
|
||
e.g. `x-opencode-session=prod-01,x-opencode-client=ai-memory`.
|
||
For gateways that require a caller-identifying header.
|
||
Headers ai-memory sets itself (authorization,
|
||
content-type, x-api-key, x-goog-api-key,
|
||
anthropic-version, anthropic-beta, openai-beta,
|
||
host, content-length) are refused at startup.
|
||
Values are never logged. A header value cannot
|
||
contain a comma through the env var — use
|
||
`llm_headers = [...]` in config.toml for that.
|
||
AI_MEMORY_RERANKER optional `llm`; reranks project/scopes query candidates
|
||
COPILOT_GITHUB_TOKEN optional GitHub token for copilot
|
||
AI_MEMORY_CODEX_EXECUTABLE optional Codex executable; defaults to codex on PATH
|
||
GITHUB_COPILOT_API_TOKEN optional pre-minted Copilot API token
|
||
COPILOT_API_URL optional Copilot API base URL override
|
||
```
|
||
|
||
**Ordered LLM fallback chain** (opt-in, TOML only — see
|
||
`docs/llm-provider-fallback.md`, #648): the primary provider above always
|
||
runs first; `[[llm_fallbacks]]` entries run only after a transient failure
|
||
(429, 5xx, timeout, connection error), in declaration order, with the
|
||
original request/schema/operation id preserved on every attempt. A
|
||
deterministic failure (4xx other than 429, an unsupported schema, a
|
||
malformed response) still stops on the first candidate — no chain-wide
|
||
retry loop.
|
||
|
||
```toml
|
||
llm_provider = "opencode"
|
||
llm_model = "mimo-v2.5-free"
|
||
|
||
[[llm_fallbacks]]
|
||
provider = "openai-compat" # same wire names as llm_provider
|
||
model = "poolside/laguna-s-2.1-free"
|
||
base_url = "http://127.0.0.1:49375/v1" # required for openai-compat, as above
|
||
api_key_env = "AI_MEMORY_LOCAL_ROUTER_TOKEN" # env var *name*; the key itself
|
||
# never lives in config.toml
|
||
|
||
[[llm_fallbacks]]
|
||
provider = "gemini"
|
||
model = "gemini-3.5-flash"
|
||
api_key_env = "GEMINI_API_KEY"
|
||
```
|
||
|
||
`api_key_env` is required for any provider that needs an API key
|
||
(`anthropic`, `openai`, `gemini`, `opencode`; optional for `openai-compat`,
|
||
which may run keyless) — it is never inherited from the primary's own
|
||
fixed env var, so a fallback cannot look configured while actually
|
||
resolving no credential. It is optional only for a provider with a native
|
||
credential source (`openai-oauth`, `copilot`, `anthropic-oauth`), which
|
||
shares the primary's process-wide token material. `Config::load` validates
|
||
every profile and resolves its credential once, at startup: a
|
||
missing/empty provider or model, an unknown provider, or a missing
|
||
credential fails startup rather than leaving a latent fallback that only
|
||
fails once the primary is already down. Each candidate carries its own
|
||
30s in-memory circuit (`ai_memory_llm::fallback::CIRCUIT_COOLDOWN`): a
|
||
transient failure opens it, a success closes it, and a restart clears all
|
||
circuit state — there is no durable circuit or forced chain-wide deadline.
|
||
|
||
`GET /admin/status` (`ai-memory status`) reports an `llm_candidates` list
|
||
alongside the existing `llm`/`embedding` roles: each candidate's
|
||
provider/model label, whether it answered the most recently completed
|
||
call, its last success/error timestamp, a redacted error class + HTTP
|
||
status (never a response body or credential), and its circuit-open-until
|
||
timestamp. It is empty for a plain single-provider setup; the top-level
|
||
`llm` role fields are unchanged.
|
||
|
||
Every chat request carries `User-Agent: ai-memory/<version>`
|
||
(`ai_memory_llm::DEFAULT_USER_AGENT`, layered in `build_provider`). `reqwest`
|
||
sends no user agent unless configured, so provider requests used to arrive
|
||
anonymous — which gateways that require callers to identify themselves report
|
||
as an unknown client. `AI_MEMORY_LLM_HEADERS=user-agent=...` overrides it. The
|
||
Copilot provider keeps `GitHubCopilotChat/<version>` instead, the
|
||
editor-plugin agent GitHub's Copilot API expects.
|
||
|
||
`openai-oauth` uses `auth login openai-oauth` and stores the ChatGPT/Codex
|
||
refresh token in `<data_dir>/auth.json`; it is separate from MCP/server bearer
|
||
auth and from OpenAI Platform API keys.
|
||
|
||
`codex` reads only the access token and account id from the Codex CLI-owned
|
||
`auth.json`, resolved from `CODEX_HOME` or the platform home. It never persists
|
||
Codex credentials. A single 401 recovery is serialized and delegated to
|
||
`codex app-server --stdio`, with bounded JSONL/stdout/stderr and a 30-second
|
||
maximum recovery timeout.
|
||
|
||
`copilot` uses `auth login copilot` or `COPILOT_GITHUB_TOKEN`, exchanges the
|
||
GitHub token through `/copilot_internal/v2/token`, and calls Copilot Chat with
|
||
the `vscode-chat` integration headers. The raw GitHub token is not sent to the
|
||
Copilot chat endpoint.
|
||
|
||
**Embedder env** (opt-in):
|
||
```
|
||
AI_MEMORY_EMBEDDING_PROVIDER openai | voyage | google | gemini | openai-compat | copilot
|
||
AI_MEMORY_EMBEDDING_MODEL e.g. text-embedding-3-small, gemini-embedding-001
|
||
AI_MEMORY_EMBEDDING_BASE_URL optional override; required for openai-compat
|
||
AI_MEMORY_EMBEDDING_DIM 1536 (OpenAI, Copilot), 1024 (Voyage), 768 (Google);
|
||
required explicitly for openai-compat
|
||
AI_MEMORY_EMBEDDING_QUERY_PREFIX optional; prepended to query text before
|
||
embedding (openai / openai-compat only)
|
||
AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX optional; prepended to document text
|
||
before embedding (openai / openai-compat
|
||
only); e.g. "query: " / "passage: " for
|
||
Nemotron-3-Embed / base E5
|
||
OPENAI_API_KEY / VOYAGE_API_KEY / GEMINI_API_KEY / GOOGLE_API_KEY
|
||
LLM_API_KEY accepted for openai with a custom base URL and as
|
||
optional bearer auth for openai-compat
|
||
EMBEDDING_API_KEY optional embedding-only key; checked before
|
||
OPENAI_API_KEY and LLM_API_KEY for the openai and
|
||
openai-compat embedders
|
||
```
|
||
|
||
`EMBEDDING_API_KEY` credentials the embedding role alone, so the embedder can
|
||
target a different provider than the chat model — `openai` on `api.openai.com`
|
||
for the LLM, a cheaper or self-hosted OpenAI-compatible endpoint for vectors.
|
||
Without it the `openai` embedder takes `OPENAI_API_KEY`, then `LLM_API_KEY`
|
||
when a custom embedding base URL is set, exactly as before. `voyage` and
|
||
`google`/`gemini` keep reading only their own `VOYAGE_API_KEY` and
|
||
`GEMINI_API_KEY`/`GOOGLE_API_KEY`.
|
||
|
||
`openai-compat` also requires an explicit model because self-hosted engines have
|
||
no safe shared model or dimensionality default. It sends no authorization header
|
||
when both `EMBEDDING_API_KEY` and `LLM_API_KEY` are absent and stores vectors
|
||
under the distinct `provider="openai-compat"` identity.
|
||
|
||
`copilot` takes no API key at all: it resolves the same `CopilotAuth` as the
|
||
`copilot` LLM provider (`auth login copilot`, `COPILOT_GITHUB_TOKEN`, or
|
||
`GITHUB_COPILOT_API_TOKEN`) and shares its GitHub-token -> short-lived
|
||
Copilot-API-token exchange (`CopilotAuthState` in `ai-memory-llm::copilot`) —
|
||
no separate exchange path. It defaults to `text-embedding-3-small` / dim 1536
|
||
and calls Copilot's `/embeddings` endpoint with the same `vscode-chat`
|
||
integration headers as chat. That endpoint's wire shape follows the
|
||
OpenAI-compatible contract Copilot documents for chat, not a published
|
||
embeddings spec, and is not covered by a live test against Copilot here.
|
||
|
||
## Future work
|
||
|
||
* **M9.5 - local embeddings via `ort`.** Bundle `bge-small-en-v1.5`
|
||
for an API-key-free homelab path. ~200 MB image bloat; trait is
|
||
ready, just needs the `OrtBgeSmallEmbedder` impl + tokenizer wiring.
|
||
* **`sqlite-vec` integration.** Brute-force cosine works fine to a few
|
||
thousand pages; past that, the `sqlite-vec` extension is the next
|
||
step. See [`docs/vector-backend-policy.md`](vector-backend-policy.md)
|
||
for the criteria that should justify adding it.
|
||
* **Scheduled consolidation queue.** Forget sweep, lint, and auto-improvement
|
||
already run on server-side schedules; a future queue can compile session
|
||
summaries outside hook latency.
|
||
* **Richer curator actions.** The shipped curator stages only one report page;
|
||
future work can add individual merge/supersession/link-fix proposals while
|
||
keeping deletes and semantic rewrites review-gated.
|
||
* **Richer read surfaces for the web UI.** The multi-workspace read-only
|
||
wiki browser shipped in `ai-memory-web` (`/web` — project list, page
|
||
tree, page view, search, and the root-only `/web/pending` triage page,
|
||
whose approve and reject buttons post to the existing
|
||
`/admin/pending-writes/*` routes). It stays read-only by design: the wiki is a
|
||
machine-authored record, and a browser edit surface would break the
|
||
invariant the whole store rests on (#482). Better *reading* — richer
|
||
navigation, diff/history views, graph exploration — is open. See
|
||
[`docs/frontend-api.md`](frontend-api.md#10-known-gaps-and-deliberate-non-goals).
|
||
* **Real LongMemEval-S harness.** The recall-eval framework exists
|
||
([`crates/ai-memory-consolidate/tests/recall_eval.rs`](../crates/ai-memory-consolidate/tests/recall_eval.rs));
|
||
porting LongMemEval-S itself requires the dataset.
|
||
|
||
## Reading order
|
||
|
||
* This file - operational summary, you are here.
|
||
* [`docs/design-decisions.md`](design-decisions.md) - the full v1 spec.
|
||
* [`docs/research-karpathy-llm-wiki.md`](research-karpathy-llm-wiki.md)
|
||
- what "Karpathy-faithful" means.
|
||
* [`docs/research-agentmemory.md`](research-agentmemory.md),
|
||
[`research-basic-memory.md`](research-basic-memory.md),
|
||
[`research-cognee.md`](research-cognee.md),
|
||
[`research-ecc.md`](research-ecc.md),
|
||
[`research-codebase-memory-mcp.md`](research-codebase-memory-mcp.md) -
|
||
prior art studied.
|
||
* [`docs/auto-improvement-loop.md`](auto-improvement-loop.md) -
|
||
Hermes Agent-inspired learning-loop research and safety boundaries.
|
||
* [`docs/issues-*.md`](.) - concrete failure modes we've designed to
|
||
avoid.
|
||
* [`CLAUDE.md`](../CLAUDE.md) - per-session operating rules pinned
|
||
into Claude Code conversations.
|