mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
M0: bootstrap workspace, CI, config loader, and init command
Sets up the foundation: an 8-crate Cargo workspace, GitHub Actions CI (fmt + clippy -D warnings + test + deny + audit), the typed identity 3-tuple in ai-memory-core, the figment-based single-read config loader, and the `ai-memory init` / `status` subcommands. Design and research that drove these choices live in docs/; CLAUDE.md holds the per-session operating rules. See design-decisions.md §14 for the cross-cutting invariants (single config read path, typed identity, atomic writes, no global singletons, …) that every milestone must respect. Verification: - cargo build --workspace clean - cargo clippy --workspace --all-targets -- -D warnings clean - cargo fmt --all -- --check clean - cargo test --workspace: 12 passed - ai-memory --version, init, status --json all working - AI_MEMORY_DATA_DIR override honoured - Second init leaves config untouched (idempotent) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
name: ci
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
RUSTFLAGS: "-D warnings"
|
||||
|
||||
jobs:
|
||||
fmt:
|
||||
name: rustfmt
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
with:
|
||||
components: rustfmt
|
||||
- run: cargo fmt --all -- --check
|
||||
|
||||
clippy:
|
||||
name: clippy
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
with:
|
||||
components: clippy
|
||||
- uses: Swatinem/rust-cache@v2
|
||||
- run: cargo clippy --workspace --all-targets -- -D warnings
|
||||
|
||||
test:
|
||||
name: test
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: dtolnay/rust-toolchain@stable
|
||||
- uses: Swatinem/rust-cache@v2
|
||||
- run: cargo test --workspace --all-targets
|
||||
|
||||
deny:
|
||||
name: cargo-deny
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: EmbarkStudios/cargo-deny-action@v2
|
||||
with:
|
||||
log-level: warn
|
||||
command: check
|
||||
arguments: --all-features
|
||||
|
||||
audit:
|
||||
name: cargo-audit
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: rustsec/audit-check@v2
|
||||
with:
|
||||
token: ${{ secrets.GITHUB_TOKEN }}
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
# Rust build artefacts.
|
||||
/target/
|
||||
**/*.rs.bk
|
||||
**/*.rs.orig
|
||||
|
||||
# Local environment overrides.
|
||||
.env
|
||||
.env.local
|
||||
/.local-data/
|
||||
|
||||
# Editor / OS noise.
|
||||
.DS_Store
|
||||
*.tmp
|
||||
*.swp
|
||||
*.swo
|
||||
.idea/
|
||||
.vscode/
|
||||
|
||||
# Local debugging.
|
||||
*.log
|
||||
flamegraph.svg
|
||||
perf.data*
|
||||
|
||||
# Sandbox marker (not project content).
|
||||
.ai-jail
|
||||
@@ -0,0 +1,146 @@
|
||||
# CLAUDE.md — ai-memory project directives
|
||||
|
||||
> Read this every session before touching code. The long-form research and
|
||||
> design specs live under [`docs/`](docs/); this file is the operating rules.
|
||||
|
||||
## What this project is
|
||||
|
||||
A self-contained Rust binary providing long-term memory for AI coding agents
|
||||
(Claude Code, OpenAI Codex, OpenCode) over the Model Context Protocol.
|
||||
Storage = markdown-in-git wiki (source of truth) + SQLite (derived index).
|
||||
Capture = automatic via agent lifecycle hooks, never manual `write_note`.
|
||||
Consolidation = Karpathy "LLM Wiki" pattern with versioned supersession.
|
||||
|
||||
Read [`docs/design-decisions.md`](docs/design-decisions.md) for the full spec.
|
||||
Read [`docs/research-karpathy-llm-wiki.md`](docs/research-karpathy-llm-wiki.md)
|
||||
for what "Karpathy-faithful" means in practice.
|
||||
|
||||
## Stack (do not deviate without updating `docs/design-decisions.md` §4)
|
||||
|
||||
- **Runtime:** Rust 1.95 (pinned in `rust-toolchain.toml`), edition 2024,
|
||||
resolver 3, async via `tokio`.
|
||||
- **MCP:** `rmcp` (official `modelcontextprotocol/rust-sdk`).
|
||||
- **Store:** `rusqlite` + `refinery` migrations, FTS5 in v1, `sqlite-vec` in
|
||||
v0.2. **One file**, one writer actor, one read pool.
|
||||
- **Wiki:** markdown on disk, `notify-debouncer-full` watcher with heartbeat
|
||||
+ reconciliation, `git2` for versioning.
|
||||
- **HTTP:** `axum` for hook ingress + MCP HTTP/SSE.
|
||||
- **LLM:** typed clients per provider (Anthropic, OpenAI, OpenAI-compat) via
|
||||
`reqwest`. **Never** a generic gateway like LiteLLM (cognee #2840 lesson).
|
||||
- **Config:** `figment`, one read at startup, passed by `&Arc<Config>`.
|
||||
- **Logging:** `tracing` with module filters; never let the appender's own
|
||||
module log at INFO+ (agentmemory #519 lesson).
|
||||
|
||||
## Repository layout
|
||||
|
||||
```
|
||||
crates/
|
||||
ai-memory-core/ # domain types, errors. NO IO.
|
||||
ai-memory-store/ # SQLite, single-writer actor, migrations.
|
||||
ai-memory-wiki/ # markdown read/write, watcher, git.
|
||||
ai-memory-mcp/ # rmcp transport + tool router.
|
||||
ai-memory-hooks/ # hook payload schemas + HTTP ingress.
|
||||
ai-memory-llm/ # LlmProvider trait + 3 impls.
|
||||
ai-memory-consolidate/ # Karpathy ingest/query/lint pipeline.
|
||||
ai-memory-cli/ # `ai-memory` binary entry point.
|
||||
hooks/ # vendored hook scripts per agent.
|
||||
docker/ # Dockerfile + compose.
|
||||
docs/ # research + design (DO NOT delete).
|
||||
tests/ # workspace integration tests.
|
||||
```
|
||||
|
||||
## Workflow rules
|
||||
|
||||
1. **Milestone by milestone.** Do not start M(n+1) until every "Done when"
|
||||
bullet in M(n) passes. See [`docs/design-decisions.md`](docs/design-decisions.md)
|
||||
for the milestone list. No mixing.
|
||||
2. **No dead code, no half-built features.** If a feature is not finished,
|
||||
it does not land. If you must stub something, document it as `M(n) TODO`
|
||||
in the relevant module's doc-comment with the milestone number.
|
||||
3. **Tests before claiming done.** Every milestone requires:
|
||||
- `cargo fmt --all -- --check` (no diffs)
|
||||
- `cargo clippy --workspace --all-targets -- -D warnings` (no warnings)
|
||||
- `cargo test --workspace` (all green)
|
||||
- Manual exercise of the new feature against a real agent CLI when applicable.
|
||||
4. **Document the why in code, not the what.** No comments restating the line
|
||||
above; only comments explaining a constraint, an incident, or a non-obvious
|
||||
invariant.
|
||||
5. **Add a unit test before the implementation, not after.** Especially for
|
||||
parsers, ID derivation, and any retention/decay math.
|
||||
6. **Don't refactor outside the milestone.** Touch only what the current
|
||||
milestone requires; resist scope creep.
|
||||
|
||||
## Cross-cutting invariants (carved in, never violated)
|
||||
|
||||
These come straight from issue-tracker research on agentmemory, basic-memory
|
||||
and cognee — every one of them is in `docs/design-decisions.md` §14 with
|
||||
issue citations. **Treat any code review that violates one of these as a
|
||||
blocking issue:**
|
||||
|
||||
1. One config-read path. No `std::env::var` outside `Config::load()`.
|
||||
2. Single-writer SQLite actor. All writes through one `mpsc` channel.
|
||||
3. Indexes commit in the same transaction as the data. No
|
||||
background-task-indexing-after-return.
|
||||
4. Typed `(WorkspaceId, ProjectId, PagePath)` identity in every layer.
|
||||
5. Hooks are fire-and-forget. Sub-second timeouts. Return 202 immediately.
|
||||
6. Privacy strip is a typed boundary (`RawHookPayload → Sanitized<Observation>`).
|
||||
7. JSON-schema structured outputs only. No XML, no `instructor`-style wrapping.
|
||||
8. `{provider, model, dim}` stored next to every embedding. Refuse on mismatch.
|
||||
9. Live-process check (`sysinfo`) before any destructive op.
|
||||
10. Atomic file writes (tmp + rename + fsync). Watcher ignores own writes.
|
||||
11. Default data dir is an absolute canonical platform path.
|
||||
Logged loudly on startup.
|
||||
12. No global singletons / `lazy_static` configs.
|
||||
13. Zero-LLM default path. LLM features opt-in via env.
|
||||
14. Tracing subscribers explicitly filter their own module.
|
||||
|
||||
## Mistakes documented in the research — do NOT repeat
|
||||
|
||||
- [`docs/issues-agentmemory.md`](docs/issues-agentmemory.md): install/ops
|
||||
landmines (iii-engine sidecar, distroless volumes, cwd-relative paths).
|
||||
- [`docs/issues-basic-memory.md`](docs/issues-basic-memory.md): file watcher
|
||||
pain, manual-capture friction, multi-workspace retrofit.
|
||||
- [`docs/issues-cognee.md`](docs/issues-cognee.md): LiteLLM/instructor wire
|
||||
drift, multi-store sync bugs, dependency landmines.
|
||||
|
||||
When in doubt about a design decision, search those files for the keyword.
|
||||
|
||||
## Quick commands
|
||||
|
||||
```bash
|
||||
# Build everything.
|
||||
cargo build --workspace
|
||||
|
||||
# Lint + format + test (run before every commit).
|
||||
cargo fmt --all -- --check
|
||||
cargo clippy --workspace --all-targets -- -D warnings
|
||||
cargo test --workspace
|
||||
|
||||
# Auto-format.
|
||||
cargo fmt --all
|
||||
|
||||
# Exercise the binary.
|
||||
./target/debug/ai-memory --version
|
||||
./target/debug/ai-memory init
|
||||
AI_MEMORY_DATA_DIR=/tmp/x ./target/debug/ai-memory init
|
||||
./target/debug/ai-memory status --json
|
||||
|
||||
# CI parity (requires cargo-deny + cargo-audit installed).
|
||||
cargo install cargo-deny cargo-audit
|
||||
cargo deny check
|
||||
cargo audit
|
||||
```
|
||||
|
||||
## What this project is NOT (v1 non-goals)
|
||||
|
||||
See [`docs/design-decisions.md`](docs/design-decisions.md) §13 for the full
|
||||
list. Highlights: no multi-tenant auth, no web UI, no Postgres backend in v1,
|
||||
no alternative vector backends, no remote sync (use `git remote` on the wiki
|
||||
dir), no multimodal.
|
||||
|
||||
## Plan & status
|
||||
|
||||
The current execution plan is at
|
||||
[`/home/akitaonrails/.claude/plans/cuddly-moseying-karp.md`](/home/akitaonrails/.claude/plans/cuddly-moseying-karp.md)
|
||||
(local to the maintainer's `~/.claude/`).
|
||||
Live progress is tracked via the TaskList tool inside Claude Code sessions.
|
||||
Generated
+1479
File diff suppressed because it is too large
Load Diff
+79
@@ -0,0 +1,79 @@
|
||||
[workspace]
|
||||
resolver = "3"
|
||||
members = [
|
||||
"crates/ai-memory-core",
|
||||
"crates/ai-memory-store",
|
||||
"crates/ai-memory-wiki",
|
||||
"crates/ai-memory-mcp",
|
||||
"crates/ai-memory-hooks",
|
||||
"crates/ai-memory-llm",
|
||||
"crates/ai-memory-consolidate",
|
||||
"crates/ai-memory-cli",
|
||||
]
|
||||
|
||||
[workspace.package]
|
||||
version = "0.1.0"
|
||||
edition = "2024"
|
||||
rust-version = "1.95"
|
||||
license = "MIT OR Apache-2.0"
|
||||
repository = "https://github.com/akitaonrails/ai-memory"
|
||||
authors = ["Fabio Akita <boss@akitaonrails.com>"]
|
||||
|
||||
[workspace.dependencies]
|
||||
# Inter-crate dependencies.
|
||||
ai-memory-core = { path = "crates/ai-memory-core", version = "0.1.0" }
|
||||
ai-memory-store = { path = "crates/ai-memory-store", version = "0.1.0" }
|
||||
ai-memory-wiki = { path = "crates/ai-memory-wiki", version = "0.1.0" }
|
||||
ai-memory-mcp = { path = "crates/ai-memory-mcp", version = "0.1.0" }
|
||||
ai-memory-hooks = { path = "crates/ai-memory-hooks", version = "0.1.0" }
|
||||
ai-memory-llm = { path = "crates/ai-memory-llm", version = "0.1.0" }
|
||||
ai-memory-consolidate = { path = "crates/ai-memory-consolidate", version = "0.1.0" }
|
||||
|
||||
# Error handling.
|
||||
anyhow = "1"
|
||||
thiserror = "2"
|
||||
|
||||
# Serialization.
|
||||
serde = { version = "1", features = ["derive"] }
|
||||
serde_json = "1"
|
||||
serde_yaml = "0.9"
|
||||
|
||||
# Async runtime.
|
||||
tokio = { version = "1", features = ["full"] }
|
||||
|
||||
# Logging / tracing.
|
||||
tracing = "0.1"
|
||||
tracing-subscriber = { version = "0.3", features = ["env-filter", "fmt", "json"] }
|
||||
tracing-appender = "0.2"
|
||||
|
||||
# Time & IDs.
|
||||
jiff = { version = "0.2", features = ["serde"] }
|
||||
uuid = { version = "1", features = ["v4", "v7", "serde"] }
|
||||
|
||||
# Config & CLI.
|
||||
figment = { version = "0.10", features = ["toml", "env"] }
|
||||
clap = { version = "4", features = ["derive", "env"] }
|
||||
dirs = "5"
|
||||
|
||||
# Secrets.
|
||||
secrecy = { version = "0.10", features = ["serde"] }
|
||||
|
||||
# Testing.
|
||||
tempfile = "3"
|
||||
|
||||
[workspace.lints.rust]
|
||||
unsafe_code = "forbid"
|
||||
missing_docs = "warn"
|
||||
|
||||
[workspace.lints.clippy]
|
||||
all = { level = "warn", priority = -1 }
|
||||
# Pedantic is opt-in per-crate once code stabilises; too noisy for early skeleton.
|
||||
|
||||
[profile.release]
|
||||
lto = "thin"
|
||||
codegen-units = 1
|
||||
strip = "symbols"
|
||||
|
||||
[profile.dev]
|
||||
opt-level = 0
|
||||
debug = true
|
||||
@@ -0,0 +1,138 @@
|
||||
# ai-memory
|
||||
|
||||
> Long-term memory for AI coding agents. Quit Claude Code mid-task, start
|
||||
> OpenAI Codex in the same directory, continue without re-explaining the
|
||||
> architecture, the failed approaches, or the open questions.
|
||||
|
||||
[](docs/design-decisions.md)
|
||||
[](rust-toolchain.toml)
|
||||
[](#license)
|
||||
|
||||
## Why this exists
|
||||
|
||||
LLM coding agents lose all context when a session ends. Today's
|
||||
"memory" tools either (a) require the user to manually invoke `write_note`
|
||||
every time something matters, or (b) wrap a vector database in a chat
|
||||
shim and call it RAG.
|
||||
|
||||
This project takes a different bet, faithful to
|
||||
[Andrej Karpathy's "LLM Wiki"](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f)
|
||||
pattern: knowledge is **compiled** at ingest time into a structured,
|
||||
cross-linked, supersedeable wiki on disk — not retrieved over raw logs at
|
||||
query time. The wiki is plain markdown in a git repo, so you can `grep`
|
||||
it, open it in Obsidian, diff it, and back it up with `rsync`.
|
||||
|
||||
Capture is **automatic** via the agent CLI's lifecycle hooks; there is no
|
||||
`write_note` ceremony. Consolidation runs in the background when a session
|
||||
ends: the LLM reads recent observations and rewrites the relevant wiki
|
||||
pages atomically with full supersession history.
|
||||
|
||||
Read [`docs/research-karpathy-llm-wiki.md`](docs/research-karpathy-llm-wiki.md)
|
||||
for the pattern, and [`docs/design-decisions.md`](docs/design-decisions.md)
|
||||
for how this project implements it.
|
||||
|
||||
## Status
|
||||
|
||||
**Under construction.** Currently at milestone **M0** (workspace bootstrap +
|
||||
CI + config). Next: **M1** — SQLite substrate + file watcher + FTS5 search.
|
||||
|
||||
See [`docs/design-decisions.md`](docs/design-decisions.md) for the full
|
||||
roadmap. v1 ships when M0–M8 are all complete; vectors arrive in v0.2 (M9).
|
||||
|
||||
## Architecture in 60 seconds
|
||||
|
||||
A single Rust binary, optionally containerised. Runs as an
|
||||
[MCP](https://modelcontextprotocol.io/) server over stdio + HTTP. Owns a
|
||||
data directory containing:
|
||||
|
||||
```
|
||||
<data_dir>/
|
||||
├── wiki/ # markdown source of truth (git-versioned)
|
||||
├── raw/ # immutable session log archive
|
||||
├── db/ # SQLite (FTS5 + sqlite-vec) — derived index
|
||||
├── models/ # bundled embedding model (v0.2+)
|
||||
└── logs/ # rolling daily tracing output
|
||||
```
|
||||
|
||||
Agent lifecycle hooks fire-and-forget POST to the server's HTTP ingress.
|
||||
The server queues writes through a single SQLite writer (no
|
||||
`database is locked`). On session end, an optional LLM-driven pass
|
||||
rewrites 5–15 wiki pages atomically with supersession (`is_latest=false`
|
||||
+ `supersedes` chain). Retrieval is hierarchical: `index.md` first, then
|
||||
page-level FTS5, then optional graph-walk expansion.
|
||||
|
||||
Storage moves between machines via `git push` of the wiki dir +
|
||||
`sqlite3 .backup` of the DB, or just `rsync` of the data dir.
|
||||
|
||||
## Quick start (M0)
|
||||
|
||||
Requires Rust 1.95+. Currently exercises only `init` / `status` — the
|
||||
MCP server lands in M2.
|
||||
|
||||
```bash
|
||||
# Build.
|
||||
cargo build --workspace
|
||||
|
||||
# Create the data directory layout.
|
||||
./target/debug/ai-memory init
|
||||
|
||||
# Or override the location.
|
||||
AI_MEMORY_DATA_DIR=/srv/ai-memory ./target/debug/ai-memory init
|
||||
|
||||
# Inspect.
|
||||
./target/debug/ai-memory status --json
|
||||
```
|
||||
|
||||
After M2, the MCP server will be attachable via:
|
||||
|
||||
```bash
|
||||
claude mcp add ai-memory -- ai-memory serve --transport stdio
|
||||
```
|
||||
|
||||
After M2.5, the Docker quick-start will be:
|
||||
|
||||
```bash
|
||||
docker run -v ai-memory:/data -p 7777:7777 ghcr.io/akitaonrails/ai-memory:latest
|
||||
```
|
||||
|
||||
## Docs
|
||||
|
||||
Long-form research and design lives under [`docs/`](docs/):
|
||||
|
||||
| File | What it is |
|
||||
|---|---|
|
||||
| [`design-decisions.md`](docs/design-decisions.md) | **Read first.** The full spec: storage, MCP surface, hooks, lifecycle, mistakes-to-avoid checklist. |
|
||||
| [`research-karpathy-llm-wiki.md`](docs/research-karpathy-llm-wiki.md) | What Karpathy actually said + community extensions, with sources. |
|
||||
| [`research-agentmemory.md`](docs/research-agentmemory.md) | Deep-dive on the TypeScript predecessor; ideas to reuse and substrate to drop. |
|
||||
| [`research-basic-memory.md`](docs/research-basic-memory.md) | The manual-write-note model we explicitly diverge from. |
|
||||
| [`research-cognee.md`](docs/research-cognee.md) | Knowledge-graph pipeline ideas to adopt + dependency landmines to avoid. |
|
||||
| [`issues-agentmemory.md`](docs/issues-agentmemory.md) | Operational landmines from the upstream tracker. |
|
||||
| [`issues-basic-memory.md`](docs/issues-basic-memory.md) | File-watcher + capture-friction landmines. |
|
||||
| [`issues-cognee.md`](docs/issues-cognee.md) | LLM-gateway + multi-store landmines. |
|
||||
|
||||
[`CLAUDE.md`](CLAUDE.md) at the repo root holds the per-session operating
|
||||
rules; pinned to every Claude Code conversation that touches this repo.
|
||||
|
||||
## Influences and prior art
|
||||
|
||||
- **[Karpathy LLM Wiki](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f)** — the compile-not-retrieve pattern.
|
||||
- **[agentmemory](https://github.com/rohitg00/agentmemory)** — most of the right ideas; this project is the Rust successor.
|
||||
- **[basic-memory](https://github.com/basicmachines-co/basic-memory)** — the markdown-on-disk source-of-truth model.
|
||||
- **[cognee](https://github.com/topoteretes/cognee)** — pipeline composition and triplet embeddings.
|
||||
- **[A-MEM](https://arxiv.org/abs/2502.12110)** — Zettelkasten-style atomic notes with link evolution.
|
||||
|
||||
## Contributing
|
||||
|
||||
The project is intentionally narrow in v1 scope; see the non-goals in
|
||||
[`docs/design-decisions.md`](docs/design-decisions.md) §13. Issues and PRs
|
||||
welcome once we cut v1.0; for now, the cleanest way to follow along is to
|
||||
read the milestones in the design-decisions doc.
|
||||
|
||||
## License
|
||||
|
||||
Dual-licensed under MIT OR Apache-2.0.
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
This codebase is being built collaboratively with Claude Code (Anthropic
|
||||
Claude Opus 4.7) following the plan documented in `docs/design-decisions.md`.
|
||||
@@ -0,0 +1,32 @@
|
||||
[package]
|
||||
name = "ai-memory-cli"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "`ai-memory` binary entry point."
|
||||
|
||||
[[bin]]
|
||||
name = "ai-memory"
|
||||
path = "src/main.rs"
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
anyhow.workspace = true
|
||||
clap.workspace = true
|
||||
dirs.workspace = true
|
||||
figment.workspace = true
|
||||
serde.workspace = true
|
||||
serde_json.workspace = true
|
||||
tokio.workspace = true
|
||||
tracing.workspace = true
|
||||
tracing-appender.workspace = true
|
||||
tracing-subscriber.workspace = true
|
||||
|
||||
[dev-dependencies]
|
||||
tempfile.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,50 @@
|
||||
//! Command-line interface definition (clap derive).
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
use clap::{Args, Parser, Subcommand};
|
||||
|
||||
/// Top-level CLI for the `ai-memory` binary.
|
||||
#[derive(Debug, Parser)]
|
||||
#[command(name = "ai-memory", version, about, long_about = None)]
|
||||
pub struct Cli {
|
||||
/// Override the data directory.
|
||||
///
|
||||
/// Defaults to a platform path under `dirs::data_local_dir()`. Also
|
||||
/// settable via the `AI_MEMORY_DATA_DIR` environment variable.
|
||||
#[arg(long, env = "AI_MEMORY_DATA_DIR", global = true)]
|
||||
pub data_dir: Option<PathBuf>,
|
||||
|
||||
/// Path to an explicit config file (defaults to `<data_dir>/config.toml`).
|
||||
#[arg(long, global = true)]
|
||||
pub config: Option<PathBuf>,
|
||||
|
||||
/// Subcommand to run.
|
||||
#[command(subcommand)]
|
||||
pub command: Command,
|
||||
}
|
||||
|
||||
/// Top-level subcommands.
|
||||
#[derive(Debug, Subcommand)]
|
||||
pub enum Command {
|
||||
/// Initialise the data directory layout.
|
||||
Init(InitArgs),
|
||||
/// Print runtime status (counts, paths, version).
|
||||
Status(StatusArgs),
|
||||
}
|
||||
|
||||
/// Arguments for `init`.
|
||||
#[derive(Debug, Args)]
|
||||
pub struct InitArgs {
|
||||
/// Overwrite an existing `config.toml` if present.
|
||||
#[arg(long)]
|
||||
pub force: bool,
|
||||
}
|
||||
|
||||
/// Arguments for `status`.
|
||||
#[derive(Debug, Args)]
|
||||
pub struct StatusArgs {
|
||||
/// Emit the report as JSON instead of human-readable text.
|
||||
#[arg(long)]
|
||||
pub json: bool,
|
||||
}
|
||||
@@ -0,0 +1,93 @@
|
||||
//! `ai-memory init` — create the data directory layout.
|
||||
|
||||
use std::fs;
|
||||
use std::io::Write;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
|
||||
use crate::cli::InitArgs;
|
||||
use crate::config::Config;
|
||||
|
||||
const DEFAULT_CONFIG_TOML: &str = include_str!("../../templates/config.default.toml");
|
||||
|
||||
const SUBDIRS: &[&str] = &["wiki", "raw", "db", "models"];
|
||||
|
||||
/// Run the `init` subcommand.
|
||||
///
|
||||
/// Creates `<data_dir>/{wiki,raw,db,models}` (idempotent) and writes a default
|
||||
/// `config.toml` unless one already exists (use `--force` to overwrite).
|
||||
///
|
||||
/// # Errors
|
||||
/// Returns an error if directories cannot be created or the config file
|
||||
/// cannot be written.
|
||||
pub fn run(config: &Config, args: InitArgs) -> Result<()> {
|
||||
let root = &config.data_dir;
|
||||
fs::create_dir_all(root).with_context(|| format!("creating data root {}", root.display()))?;
|
||||
|
||||
for sub in SUBDIRS {
|
||||
let path = root.join(sub);
|
||||
fs::create_dir_all(&path).with_context(|| format!("creating {}", path.display()))?;
|
||||
tracing::info!(path = %path.display(), "ensured directory");
|
||||
}
|
||||
|
||||
let cfg_path = root.join("config.toml");
|
||||
if cfg_path.exists() && !args.force {
|
||||
tracing::info!(
|
||||
path = %cfg_path.display(),
|
||||
"config already exists; leaving untouched (pass --force to overwrite)",
|
||||
);
|
||||
} else {
|
||||
let mut f = fs::File::create(&cfg_path)
|
||||
.with_context(|| format!("creating {}", cfg_path.display()))?;
|
||||
f.write_all(DEFAULT_CONFIG_TOML.as_bytes())
|
||||
.with_context(|| format!("writing {}", cfg_path.display()))?;
|
||||
tracing::info!(path = %cfg_path.display(), "wrote default config");
|
||||
}
|
||||
|
||||
tracing::info!("init complete");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use tempfile::TempDir;
|
||||
|
||||
fn cfg_in(dir: &std::path::Path) -> Config {
|
||||
Config {
|
||||
data_dir: dir.to_path_buf(),
|
||||
bind: "127.0.0.1:7777".into(),
|
||||
log_level: "info".into(),
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn init_creates_subdirs_and_config() {
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let config = cfg_in(tmp.path());
|
||||
run(&config, InitArgs { force: false }).unwrap();
|
||||
for sub in SUBDIRS {
|
||||
assert!(tmp.path().join(sub).is_dir(), "missing {sub}");
|
||||
}
|
||||
assert!(tmp.path().join("config.toml").exists());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn init_is_idempotent() {
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let config = cfg_in(tmp.path());
|
||||
run(&config, InitArgs { force: false }).unwrap();
|
||||
// Touch the config to detect a clobber.
|
||||
let stamp = std::fs::metadata(tmp.path().join("config.toml"))
|
||||
.unwrap()
|
||||
.modified()
|
||||
.unwrap();
|
||||
std::thread::sleep(std::time::Duration::from_millis(20));
|
||||
run(&config, InitArgs { force: false }).unwrap();
|
||||
let stamp2 = std::fs::metadata(tmp.path().join("config.toml"))
|
||||
.unwrap()
|
||||
.modified()
|
||||
.unwrap();
|
||||
assert_eq!(stamp, stamp2, "second init clobbered the config");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
//! Subcommand implementations.
|
||||
|
||||
pub mod init;
|
||||
pub mod status;
|
||||
@@ -0,0 +1,43 @@
|
||||
//! `ai-memory status` — report runtime config and counts.
|
||||
//!
|
||||
//! M0 ships a placeholder that prints paths and version. M1 wires in the
|
||||
//! store and replaces the placeholder with real counts.
|
||||
|
||||
use std::path::Path;
|
||||
|
||||
use anyhow::Result;
|
||||
use serde::Serialize;
|
||||
|
||||
use crate::cli::StatusArgs;
|
||||
use crate::config::Config;
|
||||
|
||||
#[derive(Debug, Serialize)]
|
||||
struct Report<'a> {
|
||||
version: &'a str,
|
||||
data_dir: &'a Path,
|
||||
bind: &'a str,
|
||||
notes: &'static str,
|
||||
}
|
||||
|
||||
/// Run the `status` subcommand.
|
||||
///
|
||||
/// # Errors
|
||||
/// Returns an error if JSON serialization fails.
|
||||
pub fn run(config: &Config, args: StatusArgs) -> Result<()> {
|
||||
let report = Report {
|
||||
version: env!("CARGO_PKG_VERSION"),
|
||||
data_dir: &config.data_dir,
|
||||
bind: &config.bind,
|
||||
notes: "M0 placeholder; counts arrive in M1",
|
||||
};
|
||||
|
||||
if args.json {
|
||||
println!("{}", serde_json::to_string_pretty(&report)?);
|
||||
} else {
|
||||
println!("ai-memory {}", report.version);
|
||||
println!(" data-dir: {}", report.data_dir.display());
|
||||
println!(" bind: {}", report.bind);
|
||||
println!(" note: {}", report.notes);
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
@@ -0,0 +1,148 @@
|
||||
//! Runtime configuration loader.
|
||||
//!
|
||||
//! All settings are read exactly once at startup, merged into a single
|
||||
//! immutable [`Config`] value, and passed by reference everywhere. There is
|
||||
//! no second read path (lesson from agentmemory #456 / #469 — the dimension
|
||||
//! guard read `process.env` while the rest of the codebase used
|
||||
//! `getMergedEnv()`, masking the bug for weeks).
|
||||
|
||||
use std::path::{Path, PathBuf};
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use figment::{
|
||||
Figment,
|
||||
providers::{Env, Format, Serialized, Toml},
|
||||
};
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
/// Top-level runtime configuration.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
#[serde(deny_unknown_fields, default)]
|
||||
pub struct Config {
|
||||
/// Root data directory holding `wiki/`, `raw/`, `db/`, `models/`, `logs/`.
|
||||
pub data_dir: PathBuf,
|
||||
/// HTTP bind address used by `ai-memory serve`.
|
||||
pub bind: String,
|
||||
/// Per-subsystem log filter (overridable by `RUST_LOG`).
|
||||
pub log_level: String,
|
||||
}
|
||||
|
||||
impl Default for Config {
|
||||
fn default() -> Self {
|
||||
Self {
|
||||
data_dir: default_data_dir(),
|
||||
bind: "127.0.0.1:7777".into(),
|
||||
log_level: "info".into(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl Config {
|
||||
/// Load the merged configuration: defaults → file → env → CLI.
|
||||
///
|
||||
/// # Errors
|
||||
/// Returns an error if the config file is malformed or any required
|
||||
/// field is missing.
|
||||
pub fn load(config_path: Option<&Path>, cli_data_dir: Option<PathBuf>) -> Result<Self> {
|
||||
// Figure out where the config file *would* live so we can read it
|
||||
// before knowing the final data dir. CLI > env > default.
|
||||
let probe_data_dir = cli_data_dir.clone().unwrap_or_else(default_data_dir);
|
||||
let resolved_config_path = config_path
|
||||
.map(PathBuf::from)
|
||||
.unwrap_or_else(|| probe_data_dir.join("config.toml"));
|
||||
|
||||
let mut figment = Figment::from(Serialized::defaults(Self::default()));
|
||||
if resolved_config_path.exists() {
|
||||
figment = figment.merge(Toml::file(&resolved_config_path));
|
||||
}
|
||||
figment = figment.merge(Env::prefixed("AI_MEMORY_").split("__"));
|
||||
|
||||
let mut config: Config = figment.extract().with_context(|| {
|
||||
format!(
|
||||
"loading configuration (config file = {})",
|
||||
resolved_config_path.display()
|
||||
)
|
||||
})?;
|
||||
|
||||
// CLI override always wins (figment doesn't see it because clap has
|
||||
// already consumed the env var into `cli_data_dir`).
|
||||
if let Some(dir) = cli_data_dir {
|
||||
config.data_dir = dir;
|
||||
}
|
||||
|
||||
config.data_dir = canonicalise_or_keep(&config.data_dir);
|
||||
|
||||
Ok(config)
|
||||
}
|
||||
}
|
||||
|
||||
fn default_data_dir() -> PathBuf {
|
||||
dirs::data_local_dir()
|
||||
.unwrap_or_else(|| PathBuf::from("."))
|
||||
.join("ai-memory")
|
||||
}
|
||||
|
||||
fn canonicalise_or_keep(p: &Path) -> PathBuf {
|
||||
if let Ok(canon) = p.canonicalize() {
|
||||
return canon;
|
||||
}
|
||||
// Path may not exist yet (init hasn't run). Canonicalise the parent
|
||||
// and rejoin so logs and downstream comparisons still see the truth.
|
||||
if let (Some(parent), Some(name)) = (p.parent(), p.file_name())
|
||||
&& let Ok(canon_parent) = parent.canonicalize()
|
||||
{
|
||||
return canon_parent.join(name);
|
||||
}
|
||||
p.to_path_buf()
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use tempfile::TempDir;
|
||||
|
||||
#[test]
|
||||
fn defaults_have_canonical_endings() {
|
||||
let cfg = Config::default();
|
||||
assert!(cfg.data_dir.ends_with("ai-memory"));
|
||||
assert_eq!(cfg.bind, "127.0.0.1:7777");
|
||||
assert_eq!(cfg.log_level, "info");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cli_override_wins() {
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cli_dir = tmp.path().join("override");
|
||||
let cfg = Config::load(None, Some(cli_dir.clone())).unwrap();
|
||||
assert_eq!(
|
||||
cfg.data_dir,
|
||||
// We don't expect the directory to exist yet, so the
|
||||
// canonicalise-parent fallback will return parent + name.
|
||||
cli_dir
|
||||
.parent()
|
||||
.and_then(|p| p.canonicalize().ok())
|
||||
.map(|c| c.join(cli_dir.file_name().unwrap()))
|
||||
.unwrap_or(cli_dir)
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn config_file_overrides_defaults() {
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let cfg_path = tmp.path().join("config.toml");
|
||||
std::fs::write(
|
||||
&cfg_path,
|
||||
r#"
|
||||
bind = "0.0.0.0:9999"
|
||||
log_level = "debug"
|
||||
"#,
|
||||
)
|
||||
.unwrap();
|
||||
// Use the tmp dir as the data dir so the resolved config path
|
||||
// matches what `load` derives. Passing it explicitly keeps the test
|
||||
// free of any global env.
|
||||
let cfg = Config::load(Some(&cfg_path), Some(tmp.path().to_path_buf())).unwrap();
|
||||
assert_eq!(cfg.bind, "0.0.0.0:9999");
|
||||
assert_eq!(cfg.log_level, "debug");
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,55 @@
|
||||
//! Structured tracing setup.
|
||||
//!
|
||||
//! `RUST_LOG` honoured first; otherwise we fall back to the configured
|
||||
//! [`Config::log_level`]. The appender's own module is forced to `warn` to
|
||||
//! avoid the feedback loop that filled 137 GB of disk for agentmemory #519.
|
||||
//!
|
||||
//! [`Config::log_level`]: crate::config::Config::log_level
|
||||
|
||||
use std::fs;
|
||||
|
||||
use anyhow::{Context, Result};
|
||||
use tracing_appender::non_blocking::WorkerGuard;
|
||||
use tracing_appender::rolling::{RollingFileAppender, Rotation};
|
||||
use tracing_subscriber::EnvFilter;
|
||||
use tracing_subscriber::layer::SubscriberExt;
|
||||
use tracing_subscriber::util::SubscriberInitExt;
|
||||
|
||||
use crate::config::Config;
|
||||
|
||||
/// Initialise the global tracing subscriber.
|
||||
///
|
||||
/// Returns a guard whose drop flushes any pending log lines. Keep the guard
|
||||
/// alive for the duration of `main()`.
|
||||
///
|
||||
/// # Errors
|
||||
/// Returns an error if the log directory cannot be created.
|
||||
pub fn init(config: &Config) -> Result<WorkerGuard> {
|
||||
let log_dir = config.data_dir.join("logs");
|
||||
fs::create_dir_all(&log_dir)
|
||||
.with_context(|| format!("creating log directory {}", log_dir.display()))?;
|
||||
|
||||
let appender = RollingFileAppender::new(Rotation::DAILY, &log_dir, "ai-memory.log");
|
||||
let (file_writer, guard) = tracing_appender::non_blocking(appender);
|
||||
|
||||
let default_filter = format!("{},tracing_appender=warn", config.log_level);
|
||||
let env_filter =
|
||||
EnvFilter::try_from_default_env().unwrap_or_else(|_| EnvFilter::new(default_filter));
|
||||
|
||||
let stderr_layer = tracing_subscriber::fmt::layer()
|
||||
.with_target(true)
|
||||
.with_writer(std::io::stderr);
|
||||
|
||||
let file_layer = tracing_subscriber::fmt::layer()
|
||||
.with_target(true)
|
||||
.with_ansi(false)
|
||||
.with_writer(file_writer);
|
||||
|
||||
tracing_subscriber::registry()
|
||||
.with(env_filter)
|
||||
.with(stderr_layer)
|
||||
.with(file_layer)
|
||||
.init();
|
||||
|
||||
Ok(guard)
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
//! `ai-memory` binary entry point.
|
||||
//!
|
||||
//! Loads configuration once at startup, initialises tracing, then dispatches
|
||||
//! to the requested subcommand. Domain crates take `&Config` by reference;
|
||||
//! there is no global state, no `lazy_static`, no second config-read path
|
||||
//! (lesson from agentmemory #456 / #469).
|
||||
|
||||
#![doc(html_no_source)]
|
||||
|
||||
use std::sync::Arc;
|
||||
|
||||
use anyhow::Result;
|
||||
use clap::Parser;
|
||||
use tracing::info;
|
||||
|
||||
mod cli;
|
||||
mod commands;
|
||||
mod config;
|
||||
mod logging;
|
||||
|
||||
use cli::{Cli, Command};
|
||||
use config::Config;
|
||||
|
||||
#[tokio::main]
|
||||
async fn main() -> Result<()> {
|
||||
let cli = Cli::parse();
|
||||
|
||||
let config = Arc::new(Config::load(cli.config.as_deref(), cli.data_dir.clone())?);
|
||||
let _logging_guard = logging::init(&config)?;
|
||||
|
||||
info!(
|
||||
version = env!("CARGO_PKG_VERSION"),
|
||||
data_dir = %config.data_dir.display(),
|
||||
bind = %config.bind,
|
||||
"ai-memory starting",
|
||||
);
|
||||
|
||||
match cli.command {
|
||||
Command::Init(args) => commands::init::run(&config, args),
|
||||
Command::Status(args) => commands::status::run(&config, args),
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
# ai-memory configuration.
|
||||
#
|
||||
# All values can also be overridden via environment variables prefixed
|
||||
# `AI_MEMORY_` (e.g. `AI_MEMORY_BIND=0.0.0.0:7777`).
|
||||
#
|
||||
# `data_dir` is intentionally absent here: it's set by `--data-dir`
|
||||
# or `AI_MEMORY_DATA_DIR`, and the config file lives inside it.
|
||||
|
||||
# HTTP bind address used by `ai-memory serve`.
|
||||
bind = "127.0.0.1:7777"
|
||||
|
||||
# Default tracing filter (overridable by RUST_LOG).
|
||||
log_level = "info"
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-consolidate"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "Karpathy-style ingest / query / lint consolidation pipeline."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,5 @@
|
||||
//! Karpathy "LLM Wiki" consolidation pipeline.
|
||||
//!
|
||||
//! Implements the three operations from the gist: ingest (write fan-out),
|
||||
//! query (hierarchical retrieval), and lint (contradiction & orphan
|
||||
//! detection). Plus the retention/decay sweep. Lands in milestones M7 / M8.
|
||||
@@ -0,0 +1,19 @@
|
||||
[package]
|
||||
name = "ai-memory-core"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "Core domain types and errors for ai-memory."
|
||||
|
||||
[dependencies]
|
||||
serde.workspace = true
|
||||
serde_json.workspace = true
|
||||
thiserror.workspace = true
|
||||
uuid.workspace = true
|
||||
jiff.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,39 @@
|
||||
//! Workspace-wide error type.
|
||||
|
||||
use std::path::PathBuf;
|
||||
|
||||
use thiserror::Error;
|
||||
|
||||
/// Result alias used throughout the workspace.
|
||||
pub type MemoryResult<T> = Result<T, MemoryError>;
|
||||
|
||||
/// Top-level error type for the ai-memory domain.
|
||||
#[derive(Debug, Error)]
|
||||
#[non_exhaustive]
|
||||
pub enum MemoryError {
|
||||
/// A path was outside the configured data root (defense in depth).
|
||||
#[error("path {0:?} escapes the configured data root")]
|
||||
PathEscape(PathBuf),
|
||||
|
||||
/// A page identifier was malformed.
|
||||
#[error("invalid page path: {0}")]
|
||||
InvalidPagePath(String),
|
||||
|
||||
/// A persisted record could not be parsed.
|
||||
#[error("malformed record in store: {0}")]
|
||||
MalformedRecord(String),
|
||||
|
||||
/// Wraps any underlying I/O failure.
|
||||
#[error(transparent)]
|
||||
Io(#[from] std::io::Error),
|
||||
|
||||
/// Wraps a serde deserialization failure.
|
||||
#[error("serde: {0}")]
|
||||
Serde(String),
|
||||
}
|
||||
|
||||
impl From<serde_json::Error> for MemoryError {
|
||||
fn from(value: serde_json::Error) -> Self {
|
||||
Self::Serde(value.to_string())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,183 @@
|
||||
//! Strongly-typed identifiers for the domain.
|
||||
//!
|
||||
//! The 3-tuple ([`WorkspaceId`], [`ProjectId`], [`PagePath`]) is the universal
|
||||
//! identity coordinate for any memory. It is baked in from M0 even though v1
|
||||
//! ships single-workspace, so we never inherit basic-memory's v0.20 retrofit
|
||||
//! pain (issues #783, #834, #802 and friends — see
|
||||
//! `docs/issues-basic-memory.md`).
|
||||
|
||||
use std::fmt;
|
||||
use std::str::FromStr;
|
||||
|
||||
use serde::{Deserialize, Serialize};
|
||||
use uuid::Uuid;
|
||||
|
||||
use crate::error::MemoryError;
|
||||
|
||||
macro_rules! id_newtype {
|
||||
($vis:vis $name:ident, $doc:literal) => {
|
||||
#[doc = $doc]
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||
#[serde(transparent)]
|
||||
$vis struct $name(pub Uuid);
|
||||
|
||||
impl $name {
|
||||
/// Generate a fresh v7 (time-ordered) identifier.
|
||||
#[must_use]
|
||||
pub fn new() -> Self {
|
||||
Self(Uuid::now_v7())
|
||||
}
|
||||
}
|
||||
|
||||
impl Default for $name {
|
||||
fn default() -> Self {
|
||||
Self::new()
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Debug for $name {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
f.debug_tuple(stringify!($name)).field(&self.0).finish()
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Display for $name {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
self.0.fmt(f)
|
||||
}
|
||||
}
|
||||
|
||||
impl FromStr for $name {
|
||||
type Err = MemoryError;
|
||||
fn from_str(s: &str) -> Result<Self, Self::Err> {
|
||||
Uuid::from_str(s)
|
||||
.map(Self)
|
||||
.map_err(|e| MemoryError::MalformedRecord(format!("invalid uuid: {e}")))
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
id_newtype!(pub WorkspaceId, "Workspace identifier (top of the 3-tuple).");
|
||||
id_newtype!(pub ProjectId, "Project identifier (middle of the 3-tuple).");
|
||||
id_newtype!(pub SessionId, "Identifier for a single agent run.");
|
||||
id_newtype!(pub ObservationId, "Identifier for a single observation captured during a session.");
|
||||
|
||||
/// Relative path of a page within the wiki tree.
|
||||
///
|
||||
/// Always uses `/` as the separator (POSIX-style), normalised on construction.
|
||||
/// Never starts with a slash; never contains `..` or `.` components. This
|
||||
/// invariant lets the store treat paths as flat keys without re-validating.
|
||||
#[derive(Clone, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||
#[serde(transparent)]
|
||||
pub struct PagePath(String);
|
||||
|
||||
impl PagePath {
|
||||
/// Construct from a raw string. Rejects empty, leading-slash, and
|
||||
/// dot-segment paths.
|
||||
///
|
||||
/// # Errors
|
||||
/// Returns [`MemoryError::InvalidPagePath`] when the input is empty or
|
||||
/// contains a path component that would escape or alias the wiki root.
|
||||
pub fn new(raw: impl Into<String>) -> Result<Self, MemoryError> {
|
||||
let raw = raw.into();
|
||||
if raw.is_empty() {
|
||||
return Err(MemoryError::InvalidPagePath("empty path".into()));
|
||||
}
|
||||
if raw.starts_with('/') {
|
||||
return Err(MemoryError::InvalidPagePath(format!(
|
||||
"leading slash: {raw}"
|
||||
)));
|
||||
}
|
||||
for segment in raw.split('/') {
|
||||
if segment.is_empty() || segment == "." || segment == ".." {
|
||||
return Err(MemoryError::InvalidPagePath(format!(
|
||||
"invalid segment in {raw}"
|
||||
)));
|
||||
}
|
||||
}
|
||||
Ok(Self(raw))
|
||||
}
|
||||
|
||||
/// Borrow the inner string.
|
||||
#[must_use]
|
||||
pub fn as_str(&self) -> &str {
|
||||
&self.0
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Debug for PagePath {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
f.debug_tuple("PagePath").field(&self.0).finish()
|
||||
}
|
||||
}
|
||||
|
||||
impl fmt::Display for PagePath {
|
||||
fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
|
||||
f.write_str(&self.0)
|
||||
}
|
||||
}
|
||||
|
||||
/// Discriminator for the agent CLI that captured an observation or handoff.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
||||
#[serde(rename_all = "kebab-case")]
|
||||
pub enum AgentKind {
|
||||
/// Anthropic Claude Code CLI.
|
||||
ClaudeCode,
|
||||
/// OpenAI Codex CLI.
|
||||
Codex,
|
||||
/// OpenCode (open-source coding agent).
|
||||
OpenCode,
|
||||
/// Anything else (manual capture, future agents).
|
||||
Other,
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn page_path_accepts_simple() {
|
||||
assert_eq!(PagePath::new("foo/bar.md").unwrap().as_str(), "foo/bar.md");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn page_path_rejects_empty() {
|
||||
assert!(PagePath::new("").is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn page_path_rejects_leading_slash() {
|
||||
assert!(PagePath::new("/foo").is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn page_path_rejects_dot_segments() {
|
||||
assert!(PagePath::new("a/./b").is_err());
|
||||
assert!(PagePath::new("a/../b").is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn ids_are_unique() {
|
||||
let a = WorkspaceId::new();
|
||||
let b = WorkspaceId::new();
|
||||
assert_ne!(a, b);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn id_round_trips_through_string() {
|
||||
let id = SessionId::new();
|
||||
let s = id.to_string();
|
||||
let parsed: SessionId = s.parse().unwrap();
|
||||
assert_eq!(id, parsed);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn agent_kind_serde() {
|
||||
let k = AgentKind::ClaudeCode;
|
||||
let s = serde_json::to_string(&k).unwrap();
|
||||
assert_eq!(s, "\"claude-code\"");
|
||||
let back: AgentKind = serde_json::from_str(&s).unwrap();
|
||||
assert_eq!(back, k);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,11 @@
|
||||
//! Core domain types and errors for ai-memory.
|
||||
//!
|
||||
//! This crate is the closure of the project's vocabulary: identifiers, agent
|
||||
//! kinds, and the workspace-wide error type. Nothing in here performs I/O,
|
||||
//! which keeps it trivially unit-testable and free of platform concerns.
|
||||
|
||||
pub mod error;
|
||||
pub mod ids;
|
||||
|
||||
pub use error::{MemoryError, MemoryResult};
|
||||
pub use ids::{AgentKind, ObservationId, PagePath, ProjectId, SessionId, WorkspaceId};
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-hooks"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "Hook payload schemas, sanitisation, and HTTP ingress for agent lifecycle hooks."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,5 @@
|
||||
//! Agent lifecycle hook plumbing.
|
||||
//!
|
||||
//! Typed payload schemas for Claude Code / Codex / OpenCode hook events,
|
||||
//! the `Sanitized<Observation>` boundary, and the HTTP ingress handler.
|
||||
//! Implementation lands in milestone M3.
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-llm"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "LLM provider trait with typed Anthropic, OpenAI and OpenAI-compat clients."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,6 @@
|
||||
//! LLM provider abstraction.
|
||||
//!
|
||||
//! Native typed HTTP clients per provider, never a generic gateway. Lesson
|
||||
//! from cognee #2840: `LiteLLM` + `instructor` silently drop unknown kwargs
|
||||
//! and the wrapper layer ends up papering over wire-protocol drift forever.
|
||||
//! Implementation lands in milestone M6.
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-mcp"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "MCP server transport and tool router."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,4 @@
|
||||
//! MCP server for ai-memory.
|
||||
//!
|
||||
//! Wraps the official `rmcp` SDK with the workspace tool surface. Read-only
|
||||
//! tools land in milestone M2; write tools and hook ingress follow in M3.
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-store"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "SQLite storage layer with single-writer actor and FTS5/sqlite-vec indices."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,5 @@
|
||||
//! SQLite storage layer for ai-memory.
|
||||
//!
|
||||
//! Hosts the single-writer actor and the read-only connection pool. Schema
|
||||
//! migrations land via `refinery`. Implementation lands in milestone M1; this
|
||||
//! crate is intentionally empty in M0.
|
||||
@@ -0,0 +1,15 @@
|
||||
[package]
|
||||
name = "ai-memory-wiki"
|
||||
version.workspace = true
|
||||
edition.workspace = true
|
||||
rust-version.workspace = true
|
||||
license.workspace = true
|
||||
repository.workspace = true
|
||||
authors.workspace = true
|
||||
description = "Markdown wiki read/write, file watcher and git versioning."
|
||||
|
||||
[dependencies]
|
||||
ai-memory-core.workspace = true
|
||||
|
||||
[lints]
|
||||
workspace = true
|
||||
@@ -0,0 +1,4 @@
|
||||
//! Wiki filesystem layer.
|
||||
//!
|
||||
//! Owns the markdown-on-disk source of truth: atomic writes, watcher with
|
||||
//! reconciliation pass, git versioning. Implementation lands in milestone M1.
|
||||
@@ -0,0 +1,42 @@
|
||||
# cargo-deny configuration.
|
||||
# Goal: reject GPL/AGPL transitive deps (license trap that cost cognee #2807)
|
||||
# and pin to the official crates.io registry.
|
||||
|
||||
[graph]
|
||||
all-features = false
|
||||
no-default-features = false
|
||||
|
||||
[output]
|
||||
feature-depth = 1
|
||||
|
||||
[advisories]
|
||||
version = 2
|
||||
ignore = []
|
||||
|
||||
[licenses]
|
||||
version = 2
|
||||
allow = [
|
||||
"MIT",
|
||||
"Apache-2.0",
|
||||
"Apache-2.0 WITH LLVM-exception",
|
||||
"BSD-2-Clause",
|
||||
"BSD-3-Clause",
|
||||
"ISC",
|
||||
"Unicode-DFS-2016",
|
||||
"Unicode-3.0",
|
||||
"Zlib",
|
||||
"MPL-2.0",
|
||||
"CC0-1.0",
|
||||
]
|
||||
confidence-threshold = 0.9
|
||||
exceptions = []
|
||||
|
||||
[bans]
|
||||
multiple-versions = "allow"
|
||||
wildcards = "deny"
|
||||
deny = []
|
||||
|
||||
[sources]
|
||||
unknown-registry = "deny"
|
||||
unknown-git = "deny"
|
||||
allow-registry = ["https://github.com/rust-lang/crates.io-index"]
|
||||
@@ -0,0 +1,229 @@
|
||||
# ai-memory — Design Decisions (Synthesis)
|
||||
|
||||
> Distills the four research reports (`research-*.md`) and three issue-tracker
|
||||
> reports (`issues-*.md`) into the concrete decisions this project will make.
|
||||
> Read this first; the research files are the receipts.
|
||||
|
||||
## 1. Product shape
|
||||
|
||||
A self-contained Rust binary that:
|
||||
|
||||
1. Runs as an **MCP server** (stdio + HTTP/SSE) for coding-agent CLIs (Claude Code, OpenAI Codex, OpenCode, future).
|
||||
2. Captures the agent's session **automatically** — no `write_note` ceremony — via hook scripts that the agent CLIs invoke (Claude Code lifecycle hooks, Codex hooks, OpenCode equivalents). Optional transcript-tail fallback for agents without hook APIs.
|
||||
3. Maintains a **Karpathy-style wiki**: incrementally-compiled markdown pages with cross-links, supersession, an `index.md` and a `log.md`.
|
||||
4. Serves retrieval via the MCP `tools/list` to coding agents: a handful of *narrow* tools, not 50.
|
||||
5. Ships a **Docker image** (`docker run -v ai-memory-data:/data -p 7777:7777 ai-memory`) so it can move between desktop and homelab.
|
||||
6. Is *self-healing*: schema migrations on startup, vector-index dim/provider check, write-ahead durability, periodic integrity audit, single-writer queue to avoid `database is locked`.
|
||||
|
||||
## 2. Hard requirements (extracted from the prompt)
|
||||
|
||||
- Rust, clean architecture, modular, unit-tested.
|
||||
- Cargo-format clean.
|
||||
- Docker-deployable, easy backup, easy move desktop↔homelab.
|
||||
- MCP server for coding agents.
|
||||
- **Automatic** memory capture/fetch — minimal manual tool invocations.
|
||||
- Differentiates **short-term** vs **long-term** memory temporally (like agentmemory).
|
||||
- Self-healing memory management.
|
||||
- Helps with handoffs between agent CLIs (resume from Codex where Claude Code left off).
|
||||
- Iteratively planned — each feature working before the next starts. No dead code.
|
||||
|
||||
## 3. Storage model — the biggest architectural decision
|
||||
|
||||
Three options surveyed:
|
||||
|
||||
| Option | Source-of-truth | DB used for | Pros | Cons |
|
||||
|---|---|---|---|---|
|
||||
| **A. DB-primary** | SQLite | Everything | Single transaction boundary, fast search, no FS race conditions | Opaque to humans; harder backup story |
|
||||
| **B. Markdown-in-git** primary | Files in repo | Derived index | Diff-able, grep-able, portable, Karpathy-faithful | Watcher correctness (basic-memory #580/#758/#798), inode races (#765), startup cost |
|
||||
| **C. DB-primary with on-demand export** | SQLite | Everything | Best of both | Two formats to keep coherent; user must remember to export |
|
||||
|
||||
**Decision: Option B — markdown in a git repo is source of truth, SQLite is derived index.**
|
||||
|
||||
**Why:**
|
||||
- Backup/move story is trivial — `git clone` or `rsync` a directory. The user explicitly asked for this.
|
||||
- Karpathy's pattern *is* the wiki on disk. Faking it with an export step loses the inspect-in-Obsidian property.
|
||||
- DB is rebuildable from files — corruption is recoverable.
|
||||
- Cross-tool compatibility for free: any agent that reads `~/.ai-memory/wiki/*.md` works without an MCP integration.
|
||||
|
||||
**How we avoid basic-memory's watcher pain:**
|
||||
- Watcher has a heartbeat + reconciliation pass (full diff every 30s to catch missed events).
|
||||
- We *own* writes through the MCP server's `wiki_write` path; the watcher is a *safety net* for external edits, not the primary input.
|
||||
- Inode-locking advisory + psutil-style live-process check before destructive ops (`reset`, `purge`). Lesson from basic-memory #765/#776.
|
||||
- Hidden-directory paths handled explicitly (basic-memory #798).
|
||||
|
||||
**How we avoid the "files-and-DB drift" overhead:**
|
||||
- DB stores `(path, mtime, size, sha256, indexed_at, provider, model, dim)` per page. On startup, fast scan vs. cached SHAs; only changed files re-parsed.
|
||||
- Embeddings keyed by `sha256(content) + provider + model + dim`. Re-embed only when content changes.
|
||||
|
||||
## 4. Database choice — single SQLite file
|
||||
|
||||
**Decision: one SQLite file with FTS5 + `sqlite-vec` extension + JSON columns for graph edges.**
|
||||
|
||||
Why not Postgres/pgvector? Cognee's #2717 and basic-memory's #830/#831 show Postgres is a real-deployment-only pain. v1 ships embedded.
|
||||
|
||||
Why not LanceDB/Qdrant/Kuzu/CozoDB/SurrealDB?
|
||||
- LanceDB: cognee #2702/#2720 (file-format drift, filter propagation failures). Pyarrow underneath.
|
||||
- Kuzu / Ladybug: cognee #2098/#2768 (upstream archived, fork-risk realized).
|
||||
- CozoDB: small bus factor.
|
||||
- SurrealDB: heavy, multi-mode storage; we'd inherit a lot of surface we don't need.
|
||||
- Embedded `sqlite-vec` is well-maintained, single dependency, fits in one file with FTS5 + relational tables.
|
||||
|
||||
**The graph is just SQL tables.** A `wiki_pages` table, a `wiki_links (from_id, to_id, link_type)` table, optional `wiki_concepts (page_id, concept)`. Graph queries are recursive CTEs in SQLite. Petgraph in-memory for batch traversals. Avoids the entire "embedded graph DB" footgun cognee fell into.
|
||||
|
||||
**Crates** (research-backed picks):
|
||||
- `rusqlite` + `rusqlite-extension` for sqlite-vec loading. `bundled-sqlcipher` if we want encryption later.
|
||||
- `sqlx` for migrations (`sqlx::migrate!`). Async, type-checked.
|
||||
- `tantivy` *not* used initially — sqlite FTS5 is sufficient at the corpus sizes we expect (hundreds to low-thousands of pages per project). Revisit only if FTS5 ranking proves inadequate.
|
||||
- `petgraph` for in-memory graph algorithms during consolidation.
|
||||
|
||||
## 5. Embedding & LLM
|
||||
|
||||
**Embeddings:**
|
||||
- Default: **local model via `ort` (ONNX Runtime) crate or `fastembed-rs`** running `bge-small-en-v1.5` (384 dim) or `bge-small-en-v1.5-q` quantized. Same model basic-memory uses.
|
||||
- Persist `{provider, model, dim}` next to every vector. Refuse to load on mismatch with clear remediation (agentmemory #469 lesson).
|
||||
- Cache path: `<data_dir>/models/`, never `/tmp` (basic-memory #741).
|
||||
- Trait-based: `trait Embedder { ... }` with implementations `LocalOrtEmbedder`, `OpenAIEmbedder`, `VoyageEmbedder`. User configures one.
|
||||
|
||||
**LLM for consolidation passes:**
|
||||
- **Off by default**, behaves like agentmemory after #138's fix. Without a provider, the system still works: synthetic compression (rule-based), no LLM-generated summaries, no `memory_consolidate` page-rewrite.
|
||||
- With a provider, scheduled consolidation runs (1× per session-end + optional 6h timer).
|
||||
- Provider trait `LlmProvider { complete(...); complete_structured(...) }`. Implementations: `AnthropicProvider`, `OpenAIProvider`, `OllamaProvider`, `OpenAICompatProvider`.
|
||||
- **Native HTTP per provider** — no LiteLLM-equivalent. The cognee tracker (#2412/#2430/#2537/#2608/#2749/#2782/#2840/#2842) showed silent-kwarg-drop in a generic gateway is the #1 source of provider bugs. Each provider's typed JSON, errors on unknown fields. Hand-coded but correct.
|
||||
- **Structured output via JSON schema, not XML, not Instructor-style wrapping.** Use each provider's native JSON-mode where available; for Anthropic, request a tool-use response with a typed schema. Validate with `serde_json` + `schemars`-derived schemas.
|
||||
|
||||
## 6. Capture model — auto, never `write_note`
|
||||
|
||||
Three capture surfaces, in priority order:
|
||||
|
||||
1. **Lifecycle hooks** (Claude Code, Codex, OpenCode). These are fast, reliable, structured. We ship hook scripts the user installs once. Lessons from agentmemory:
|
||||
- Hooks must be **fire-and-forget** (#221). No `await fetch()` blocking session start.
|
||||
- Sub-second hard timeouts on the writer side (`tokio::time::timeout`).
|
||||
- All hooks → single HTTP/Unix-socket POST → server queues → returns 202 immediately.
|
||||
- Privacy strip at the hook boundary, not later (agentmemory `stripPrivateData`).
|
||||
|
||||
2. **Transcript tail** (universal fallback). Watch `~/.claude/projects/`, `~/.codex/`, `~/.config/opencode/sessions/`. Lossier but works for any agent. Required for the basic-memory #669/#687/#730 demand the tracker has been asking for.
|
||||
|
||||
3. **Manual MCP tool** (`memory_remember`) — only for ad-hoc explicit captures from the user ("remember this"). Not the primary path; not what the agent reaches for by default.
|
||||
|
||||
## 7. Memory model (temporal)
|
||||
|
||||
Adopt agentmemory's tier model **but** keep the surface narrow:
|
||||
|
||||
| Tier | What it is | Lifetime | Decay |
|
||||
|---|---|---|---|
|
||||
| **Working** | Current session: last N observations, last user prompt, current files | Until session end | Drop on session end (kept in DB for forensics, but excluded from default recall) |
|
||||
| **Episodic** | Per-session summaries with concept tags, files-touched, decisions made | 30 days hot, 180 days cold, then evict if cold-score < threshold | Salience × exp(-λΔt) + Σ(σ/days_since_access) — agentmemory's formula, validated |
|
||||
| **Semantic** | Distilled facts/preferences/architecture notes — the wiki pages themselves | Indefinite, supersedeable | Versioned in place: old `is_latest=false`, new `supersedes=old_id` |
|
||||
| **Procedural** | Repeated patterns extracted from episodic clusters (`pattern` type with frequency ≥ 2) | Indefinite | Frequency-decay if not re-observed in N days |
|
||||
|
||||
**Implementation note:** the four tiers map to one `pages` table with a `tier` enum column + an `observations` table for raw working/episodic, not four separate tables. Keeps schema migrations sane.
|
||||
|
||||
## 8. Consolidation (the Karpathy bit)
|
||||
|
||||
Three scheduled MCP operations:
|
||||
|
||||
- **`memory_ingest`** (auto-called by hooks): one observation → write-fan-out to ~5–15 wiki pages. New page if no match; supersede + version if the page already exists. No-LLM fallback: append to a per-day digest page if no provider configured.
|
||||
- **`memory_query`** (called by agent on demand): hierarchical — search `index.md` first, then page-level FTS+vector, then optional graph-walk expansion. RRF-fused. Agentmemory hit 95.2% R@5 with this pattern.
|
||||
- **`memory_lint`** (scheduled hourly + on session-end): scans for contradictions, orphan pages, broken links, stale claims, low-confidence + zero-reinforcement entries. Pure LLM with strict JSON output.
|
||||
|
||||
Decay/forget runs as a separate `memory_forget_sweep` job: applies the retention formula; soft-deletes via `is_latest=false` + `superseded_at`; hard-deletes only after 180 days *and* zero accesses. Never silently destroys anything user-pinned.
|
||||
|
||||
## 9. Cross-agent handoff
|
||||
|
||||
A first-class typed protocol, not just shared state:
|
||||
|
||||
```rust
|
||||
struct Handoff {
|
||||
from_agent: String, // "claude-code", "codex"
|
||||
to_agent: Option<String>,
|
||||
project_id: ProjectId,
|
||||
cwd: PathBuf,
|
||||
summary: String,
|
||||
open_questions: Vec<String>,
|
||||
files_touched: Vec<PathBuf>,
|
||||
next_steps: Vec<String>,
|
||||
model: String,
|
||||
created_at: DateTime,
|
||||
}
|
||||
```
|
||||
|
||||
MCP tools `memory_handoff_begin` (writes a handoff page tagged `state=open`) and `memory_handoff_accept` (acknowledges, returns the handoff content, marks `accepted_by`). The user can stop Claude Code, start Codex, and Codex's session-start hook fetches the latest open handoff for the cwd.
|
||||
|
||||
agentmemory has this informally (`/handoff` skill); we make it explicit from day one because every research report flagged cross-agent as the v0.1 weak spot.
|
||||
|
||||
## 10. MCP tool surface — narrow on purpose
|
||||
|
||||
basic-memory has ~25 tools, agentmemory has 53. Both have user confusion as a result. Ship **at most 10 v1 tools**:
|
||||
|
||||
| Tool | Purpose | Annotation |
|
||||
|---|---|---|
|
||||
| `memory_remember` | Manual capture (rare) | destructive, idempotent |
|
||||
| `memory_query` | Search + retrieve, auto-routed | read-only |
|
||||
| `memory_recent` | Recent activity, optionally for-this-project | read-only |
|
||||
| `memory_handoff_begin` | Mark session boundary, write handoff | destructive |
|
||||
| `memory_handoff_accept` | Fetch+ack open handoff | destructive |
|
||||
| `memory_forget` | Explicit user forget | destructive |
|
||||
| `memory_session_summary` | Internal — used by stop hook | destructive (internal) |
|
||||
| `memory_consolidate` | Internal — used by scheduler | destructive (internal) |
|
||||
| `memory_lint` | Internal — used by scheduler | destructive (internal) |
|
||||
| `memory_status` | Health, counts, last-consolidation-at | read-only |
|
||||
|
||||
Internal tools are gated by `tools/list` annotation; agents see only the user-visible ones unless `expose=all` is set. Every tool has MCP `readOnlyHint`/`destructiveHint`/`idempotentHint` (basic-memory pattern; lesson from #818 — be careful with `bool | None` aliases on tool params).
|
||||
|
||||
Tool param aliases: accept `query|q|search`, `project|workspace`, `dir|directory` — basic-memory's `AliasChoices` pattern works for LLM resilience.
|
||||
|
||||
## 11. Identity & project scoping (3-tuple from day one)
|
||||
|
||||
Lesson from basic-memory's v0.20 trauma: `(workspace, project, page_path)`. Even if v1 ships single-workspace, the schema and every API/tool param encodes the full 3-tuple. No retrofits.
|
||||
|
||||
Project resolution chain: explicit param → server's default → cwd-based heuristic (match repo root) → error.
|
||||
|
||||
## 12. Operability
|
||||
|
||||
- **Single binary**, statically-linked where possible. Distroless Docker image. **Absolute data path** by default (`dirs::data_local_dir().join("ai-memory")`); log it loudly on startup (agentmemory #303 lesson).
|
||||
- **Atomic config**: one `Config::load()` → typed struct, every reader takes `&Config`. No `process.env` double-read paths (agentmemory #456/#469).
|
||||
- **Write durability**: every observation lands in SQLite *and* is appended to a `log.md` line *before* the hook gets its 202. No background-task indexing-after-return (basic-memory #763/#578/#839).
|
||||
- **Migrations**: `sqlx::migrate!` runs on startup; never inline DDL (basic-memory #727).
|
||||
- **Schema versioning**: one source of truth for the schema; derived clients/docs. No "update 7 files" checklists (agentmemory AGENTS.md smell).
|
||||
- **Backup/move**: `ai-memory export <dir>` dumps wiki/ + sqlite snapshot. `ai-memory import <dir>` consumes. Default data dir is portable. Optional: `auto_git_commit = true` config flag → commits the wiki directory on every `memory_lint` run.
|
||||
- **Self-healing**: startup checks (`memory_diagnose`): vector dim/provider drift, FTS index corruption, orphan pages, broken links, zombie sessions. `memory_heal` auto-fixes the safe subset.
|
||||
- **Logging**: structured `tracing` with rotating files, capped at N MB. No feedback loops (agentmemory #519).
|
||||
|
||||
## 13. What we are explicitly NOT doing in v1
|
||||
|
||||
To stay scoped:
|
||||
|
||||
- No multi-tenant auth/RBAC (single-user homelab).
|
||||
- No web UI / dashboard (use `sqlite3` + `glow`/Obsidian).
|
||||
- No Postgres backend (revisit if a real homelab user hits scale walls).
|
||||
- No remote/cloud sync (use git remote on the wiki dir).
|
||||
- No alternative embedded vector backends (sqlite-vec only).
|
||||
- No alternative graph DB (SQL recursive CTEs only).
|
||||
- No multimodal (text only).
|
||||
- No "skills" / slash-command bundle in v1 (agentmemory plugin format) — focus on hooks + MCP first.
|
||||
- No LongMemEval-style benchmark harness in v1 — add in v0.4.
|
||||
|
||||
## 14. Mistakes-to-avoid checklist (from issue research)
|
||||
|
||||
Top-line rules carved into the codebase:
|
||||
|
||||
1. One config-read path (agentmemory #456/#469).
|
||||
2. Indexes in the same txn as the source-of-truth row (agentmemory #204/#309, basic-memory #763/#578).
|
||||
3. JSON-schema structured outputs, no XML (agentmemory #492/#539; cognee #2840).
|
||||
4. Hooks fire-and-forget (agentmemory #221, #143).
|
||||
5. No background-task index-after-return; either sync or `index_status: pending` (basic-memory #763).
|
||||
6. 3-tuple identity from day one (basic-memory #783/#834).
|
||||
7. Vector index records `{provider, model, dim}`; refuse on mismatch (agentmemory #469).
|
||||
8. Embedding cache path absolute, not `/tmp` (basic-memory #741).
|
||||
9. Watcher heartbeat + reconciliation pass (basic-memory #580/#758/#798).
|
||||
10. Live-process check before destructive ops (basic-memory #765).
|
||||
11. Per-provider typed HTTP client; no LiteLLM equivalent (cognee #2840).
|
||||
12. Idempotent ingest with deterministic id derivation (cognee #2510/#2557/#2633).
|
||||
13. Single transactional boundary; no implicit graph/vector/relational sync (cognee Section B).
|
||||
14. Filter propagation tests (cognee #2720 was a recall correctness bug).
|
||||
15. Default data dir is an absolute canonical platform path (agentmemory #303).
|
||||
16. No `lru_cache` on configs (cognee #2228/#2853).
|
||||
17. Datasets/projects are query-time filters, not orchestration-mode-conditional (cognee #2867).
|
||||
18. LLM features off by default; opt-in via env (agentmemory #138/#143).
|
||||
19. `cargo deny` for transitive license audits (cognee #2807 — FastEmbed removed for license).
|
||||
20. Pin upstream native deps; ship a lockfile (agentmemory #555/#540).
|
||||
@@ -0,0 +1,87 @@
|
||||
# agentmemory — Issue & PR Pain-Point Synthesis
|
||||
|
||||
> Source: GitHub `rohitg00/agentmemory`, captured 2026-05-21.
|
||||
> Repo health: 15.7k stars, very active. ~50 merged PRs in last week.
|
||||
> Architecture: TypeScript MCP server over a native Rust `iii-engine` KV.
|
||||
|
||||
## Top recurring pain points (ranked)
|
||||
|
||||
### 1. Install / ops — the single largest bucket
|
||||
- **`iii-engine` is a separate native binary** with its own version, config file, and storage layout. Pinning is fragile: `iii-sdk@0.11.6` broke routing because `package.json` used `^0.11.2` (#555, fixed by PR #567 pinning exact). Migrating to `iii-database`/iii 0.11.7 is blocking the SQLite migration (#309 comment). The whole stack is held back by an upstream you don't control.
|
||||
- **Distroless engine + Docker named volumes** = silent permission denied. UID 65532 can't write to root-owned `/data`; engine logs the error but the wrapper buffers in RAM and looks fine until restart wipes everything (#301 — still OPEN despite an earlier 0.9.7 "fix").
|
||||
- **Engine writes `data/` to caller's `cwd`**, so launching from different directories produces different state stores. On Windows users believed memories vanished — they were stranded in `E:\文档\New project\data\state_store.db` while the dashboard read `C:\Users\Lenovo\data\` (#303, still OPEN even after PR #314 added `--data-dir`).
|
||||
- **Hooks break on Windows when username has spaces** because `hooks.json` doesn't quote `${CLAUDE_PLUGIN_ROOT}` (#477).
|
||||
- **Runaway log feedback loop** — `iii::workers::observability` warns about "subscriber lagged", that warn is captured by the same subscriber: 137 GB `daemon.log.new`, system at 98% (#519, still OPEN; `RUST_LOG` not honoured).
|
||||
|
||||
### 2. Data integrity / silent loss
|
||||
- **State persistence buffered behind a 5s `IndexPersistence` debounce**. When `state::set` times out at 30s, the uncaught `IIIInvocationError` crashes the Node process, losing every in-memory BM25/vector update since the last debounce flush (#204).
|
||||
- **BM25 index `mem%3Aindex%3Abm25.bin` stays at ~96 bytes** because every `state::set` times out at 180s on 10k-observation corpora; each restart pays a 5-minute rebuild (#309, OPEN).
|
||||
- **Sessions never end on Ctrl-C / SSH-drop / laptop sleep**, so the consolidation + graph-extraction pipeline never fires, then the eviction sweep deletes the entire session (#308, "graph would-have-extracted content is permanently lost").
|
||||
- **`AGENTMEMORY_DROP_STALE_INDEX=true` in `.env` did nothing** — the dimension guard read `process.env` directly while everything else reads `getMergedEnv()`. **Two config-read paths in the same codebase** (#456). Combined with #469 (vector index 2048-dim on disk vs 384-dim from provider), stranded users with no working recovery path.
|
||||
- **Sessions never created, but observations were** — separate KV scopes; OpenClaw plugin only wrote observations, so `GET /sessions` returned `[]` (#522, OPEN). Masked by `postJson({fallback_on_error:true})` swallowing the 4xx.
|
||||
- **`session.summary` and `session.firstPrompt` set to the same truncated title** (#276, OPEN, labelled CRITICAL).
|
||||
|
||||
### 3. LLM compression / token-cost
|
||||
- **`AGENTMEMORY_AUTO_COMPRESS=true` was the default in v0.8.7**. User olcor1: *"My allocation is busted within 20 minutes."* — the PostToolUse hook called the LLM on every tool call (#138, brown-paper-bag fix in v0.8.8 flipping default to false). Maintainer: *"this is a real bug in the tool's design, not your setup."*
|
||||
- **`SessionStart` hook was injecting ~1-2 K tokens into every new session**. Maintainer initially blamed Claude Pro caps, then retracted and gated injection behind `AGENTMEMORY_INJECT_CONTEXT=false` default in v0.8.10 (#143). His own retraction: *"I pattern-matched without verifying against the docs."*
|
||||
- **`mem::compress` silently failed on ~47% of Claude Code tool calls** because `post-tool-use.mjs` read `data.tool_output` but Claude Code sends `tool_response` (#539, fixed PR #561).
|
||||
- **Graph extraction parser drops self-closing `<entity .../>` tags** (#492). Same family: #338.
|
||||
|
||||
### 4. Retrieval quality / API contract
|
||||
- **MCP `memory_recall` was aliased to `smart_search` and dropped the `format` param**, so full content was unreachable via MCP no matter what the caller asked for (#440 + #507, fixed PR #516). Six weeks of users getting only compact-mode hits.
|
||||
- **Viewer/status show 0 memories on real corpora** because `/agentmemory/memories?latest=true` and `/agentmemory/export` materialize the entire list, time out on >8k memories (#544, OPEN).
|
||||
- **Event-loop starvation from pure-JS dot products** — VectorIndex.search hangs the loop on 100k vectors; sqlite-vec swap took search from 200-250 ms to 20-40 ms (#195).
|
||||
|
||||
### 5. Agent integration / MCP surface
|
||||
- **Claude Code requested protocol version 2025-03-26 but the shim pinned 2024-11-05**; Claude Code discarded the tools list as a result (#510, OPEN). Same root in #553 (OpenCode 8/51 tools) and #400.
|
||||
- **Codex worktrees treated as separate projects**, fragmenting lessons/sessions per ephemeral worktree path (#515, OPEN).
|
||||
- **Hooks blocked startup** — `await fetch(... 5000 ms timeout)` on session-start; 10 parallel `claude -p` jobs OOM-killed the engine (#221, fixed PR #222 making hooks fire-and-forget).
|
||||
|
||||
## Design choices that caused the most issues
|
||||
|
||||
1. **Embedding the search indexes in the Node process while persisting through a remote KV with a 30s timeout.** Drives #204, #309, the rebuild-on-boot cost, and 5-second window of data loss on every crash.
|
||||
2. **LLM compression on every observation, on by default** (#138, #143, #539).
|
||||
3. **Two config-read paths (`process.env` vs `getMergedEnv()`)** caused #456/#469.
|
||||
4. **XML as the compression/extraction wire format**, hand-parsed: drops self-closing tags (#492), accepts only specific casings, fails `CompressOutputSchema` on schema drift (#539).
|
||||
5. **Hooks that `await` REST round-trips during agent startup** (#221).
|
||||
6. **Relative paths in the bundled `iii-config.yaml`** (#303), distroless engine without a chown init container (#301), no log rotation (#519).
|
||||
7. **`fallback_on_error: true` everywhere with swallowed errors** (#522, #539).
|
||||
8. **Unpinned upstream native dep** (`iii-sdk: ^0.11.2`) → #555. No lockfile shipped → #540.
|
||||
|
||||
## Maintainer fixes reveal regrets
|
||||
|
||||
- **Auto-compress flipped from default-on to default-off** in v0.8.8 the same day #138 was filed — *"the closest thing to an admission that the headline feature was the headline misfeature"*. *"This is exactly the kind of brown-paper-bag issue."*
|
||||
- **Context injection moved behind `AGENTMEMORY_INJECT_CONTEXT=false`** in v0.8.10 (#143) — maintainer publicly retracted a wrong first diagnosis. Two defaults reversed in two minor versions.
|
||||
- **`VECTOR_BACKEND=sqlite-vec` introduced behind a flag**, not flipped on, explicitly because *"some Windows / Alpine Docker users will hit install issues we can't preempt"* (#195).
|
||||
- **Multiple GH-packages mirror experiments reverted within hours** — PRs #545 → #547 → #548. Lots of try/revert.
|
||||
- **PR #500: rebuildIndex made non-blocking on boot**, PR #504: batch-embed in rebuildIndex (25 h → 3 h on large corpora). Admits the original boot path made the daemon unusable for hours.
|
||||
- **OpenCode plugin (#236) shipped as separate subsystem** because the Claude Code hook abstraction didn't fit other agents.
|
||||
|
||||
## Still-open architectural debts
|
||||
|
||||
- **#309 in-memory BM25/graph → SQLite/FTS5** — blocked on iii v0.11.7. Biggest debt in the repo.
|
||||
- **#519 daemon.log feedback loop** — offending warn is inside the closed-source `iii` binary.
|
||||
- **#303 cwd-relative state store on Windows** — still leaks the wrong dir despite #314.
|
||||
- **#301 distroless docker volume permissions** — still OPEN.
|
||||
- **#510 / #553 / #400 MCP protocol-version negotiation** — three reports, same root cause.
|
||||
|
||||
## Seven "do not repeat" lessons for the Rust rewrite
|
||||
|
||||
1. **Keep search indexes and durable storage in the same transaction boundary.** Don't buffer index writes in memory behind a debounce that loses 5s of data on crash (#204, #309). In Rust, use a single `sqlx` transaction per observation; SQLite FTS5 + `sqlite-vec` in one file solves all three search types and removes the entire "rebuild on boot" pathology.
|
||||
|
||||
2. **LLM compression must be opt-in, with a visible token-cost banner.** Default to a zero-LLM synthetic compression (extract title/files/narrative from raw tool I/O) (#138, #143).
|
||||
|
||||
3. **One config-read path.** Have `Config::load()` resolve env + file + CLI once into a typed struct; every reader takes `&Config` (#456 + #469).
|
||||
|
||||
4. **Never use XML as a wire format for LLM extraction.** Use JSON-mode / structured outputs (#492, #539).
|
||||
|
||||
5. **Hooks must be fire-and-forget by contract.** No `.await` on an HTTP round-trip during agent startup (#143, #221). Budget the response time hard (`tokio::time::timeout` with sub-second ceilings).
|
||||
|
||||
6. **Persist provider metadata next to the index.** A vector index file must record `{provider, model, dim}` and refuse to load on mismatch with a *single* clear error and an in-process re-embed migration path (#469).
|
||||
|
||||
7. **Don't depend on an unpinned native sidecar.** Statically link the engine (`tantivy` for BM25, `sqlite-vec` via `rusqlite`, `petgraph` for the concept graph) (#301, #519, #555). Half the install issues in the tracker exist because `iii-engine` is a separate binary the wrapper can't fix.
|
||||
|
||||
**Bonus:** Default the data directory to a canonicalized absolute platform path (`dirs::data_local_dir().join("ai-memory")`) and log it loudly on startup — single change would have prevented #303 entirely.
|
||||
|
||||
### Driver issues
|
||||
#138 (auto-compress default), #143 (context-injection default), #195 (CPU-bound JS), #204 (uncaught SDK timeout), #221 (blocking hooks), #274 (lesson discard), #276 (corrupt session fields), #301 (distroless volume), #303 (cwd state), #308 (sessions never end), #309 (in-memory BM25), #338/#492 (XML parser), #440/#507 (MCP recall aliased), #456/#469 (dim mismatch + two env paths), #477 (Windows quoting), #510/#553/#400 (MCP protocol version), #515 (Codex worktrees), #519 (log feedback loop), #522 (silent error swallow), #539 (tool_response vs tool_output), #540 (no lockfile), #544 (unbounded list endpoints), #555 (iii-sdk semver).
|
||||
@@ -0,0 +1,105 @@
|
||||
# basic-memory — Issue & PR Pain-Point Synthesis
|
||||
|
||||
> Source: GitHub `basicmachines-co/basic-memory`. Captured 2026-05-21.
|
||||
> Tracker is unusually high-signal — small team, deep technical replies.
|
||||
> Dominant themes: sync correctness, multi-project routing, embedding install hell.
|
||||
|
||||
## Top recurring pain points (ranked)
|
||||
|
||||
### 1. Sync correctness & file-watcher reliability — the single biggest theme
|
||||
|
||||
- **#580** — watch service can go stale while process stays alive (no heartbeat, no liveness signal, `awatch()` blocks forever on macOS FSEvents buffer overflow). Closed with partial fix.
|
||||
- **#758** — watch service ignored the `--project` constraint, N concurrent MCP processes produced N overlapping watchers and races. The fix found *three independent bugs* while fixing one.
|
||||
- **#798** — watch service silently dropped events under hidden-directory parents (e.g., projects under `~/.claude/...`) because gitignore-style globs treated *every* component starting with `.` as hidden.
|
||||
- **#765** — "Stale FTS index entries persist after `reset --reindex`": fixed by detecting **zombie MCP processes holding the old SQLite inode** after `unlink`. > *"That process keeps its connection to the old `memory.db` inode (now unlinked but not freed)... newly-spawned MCP processes attach to the new file. But any tool call that gets routed to a zombie MCP process queries the **old** inode."* Now `bm reset` refuses to run while MCP processes are alive (PR #776).
|
||||
- **#763** — `write_note` returns *before* semantic indexing completes. Maintainer defends this as intentional architecture but admits the CLI path loses embeddings on process exit.
|
||||
- **#839** (OPEN) — CLI `write-note` prints `CancelledError` traceback because `_log_task_failure` doesn't handle task cancellation on process exit. Same root cause as #763.
|
||||
- **#578** — *new entities silently skip embedding generation* after a single sqlite-vec load failure. Background task errors fire-and-forget.
|
||||
- **#634** — schema-validate uses stale `entity_metadata` when files edited externally.
|
||||
- **#481** — `alembic/env.py` was unconditionally setting `BASIC_MEMORY_ENV="test"` at module import time, which silently disabled the watch service in production.
|
||||
|
||||
### 2. Multi-project / multi-workspace routing — the v0.20 trauma cluster
|
||||
|
||||
A wave of nearly identical bugs in April–May 2026: **#782, #783, #788, #793, #799, #800, #802, #803, #804, #805, #810, #820, #834**. Every one of them is the same shape: an MCP tool ignores or mis-resolves a project/workspace identifier. Maintainer on #783: > *"Permalinks are unique per project but project is not unique per workspace... any tool that holds only a permalink string cannot distinguish between them."* **Architecture catching up to ambition**: permalinks were designed as a 2-tuple (project, path) and later had to grow a third dimension (workspace) under load.
|
||||
|
||||
### 3. SQLite-vec / embedding provider install hell
|
||||
|
||||
- **#735 / #767** — duplicates: "no such module: vec0" on Windows in worker connections. Fixed by ensuring `_ensure_sqlite_vec_loaded` is called on *every* session that touches vec0.
|
||||
- **#829 / #658** — still-open variants.
|
||||
- **#741 / #681** — FastEmbed cache defaulted to `/tmp/fastembed_cache`, wiped in sandboxed runtimes like Codex CLI, causing `ONNXRuntimeError: NO_SUCHFILE` on every subsequent semantic search.
|
||||
- **#830** (OPEN) — `docker-compose-postgres.yml` ships plain `postgres:17` instead of `pgvector/pgvector:pg17`; semantic search silently fails.
|
||||
- **#831** (OPEN) — `IndexError: pop from an empty deque` during async engine dispose on Postgres/asyncpg.
|
||||
|
||||
### 4. Parser fragility (markdown-as-source-of-truth)
|
||||
|
||||
- **#738** — Parser captured Obsidian callout syntax (`> [!note]`) as observation categories.
|
||||
- **#721** — `edit_note` fails on notes with long text around inline wikilinks because `relation_type` exceeds `MaxLen`.
|
||||
- **#528** — Cloud sync prepended duplicate frontmatter to files that already had YAML.
|
||||
- **#408** — YAML frontmatter parsing fails with unquoted colons in title.
|
||||
- **#256** — *"Editing notes causes them to disappear from index"* — the search index DELETE was missing a `project_id` filter, so an edit in one project nuked search rows in every project with the same permalink.
|
||||
|
||||
### 5. Search quality / pagination
|
||||
|
||||
- **#693** — `read_note` pagination params were ignored at the API endpoint.
|
||||
- **#354** — `tag:tagname` syntax silently treated as literal text.
|
||||
- **#686** (OPEN) — user hitting MCP response size limits at 57 pages.
|
||||
- **#666, #618, #603** (all OPEN) — reranking, time-decay, length normalization. Search ranking is admitted-weak by maintainers.
|
||||
|
||||
### 6. Backup / undo / git
|
||||
|
||||
- **#124** (OPEN since June 2025) — git-based undo. Maintainer wrote design spec but issue remains open: > *"there hasn't been an obvious way to handle when to commit. Commit on every change? Commit every so often, how often? Push to a remote? How to handle conflicts?"*
|
||||
- **#59** — leverage `git diff` to prevent knowledge corruption/deletion. OPEN since March 2025.
|
||||
|
||||
### 7. Manual capture friction — present but indirect
|
||||
|
||||
**Nobody filed "I'm tired of telling the agent to remember."** But:
|
||||
- **#297** — Cursor + basic-memory produce sensationalized progress logs (`BREAKTHROUGH - Live Data Packets Detected!`) that pollute search results. Maintainer closes as **not basic-memory's fault, it's the LLM's**: > *"It is just the tools. Your LLM is in charge of how to use it."* That's the philosophical commitment that drives the manual `write_note` workflow — transferring all the noise problems onto the user.
|
||||
- **#669, #730, #687** (all OPEN) — three separate proposals to add a *sidecar that watches session transcripts and builds the knowledge graph automatically*. **The strongest signal that manual capture is friction.** The maintainer is interested but hasn't started.
|
||||
|
||||
## Design choices that caused the most issues
|
||||
|
||||
| Design choice | Cited in | Symptom |
|
||||
|---|---|---|
|
||||
| Background async `create_task` for embedding sync | #763, #578, #839 | Lost embeddings on CLI exit; silent skips; CancelledError tracebacks |
|
||||
| Single permalink space per project (not per workspace) | #783, #802, #834 | Teams launch blocker; cross-workspace name collisions |
|
||||
| sqlite-vec extension loaded per-session, not globally | #735, #767, #829, #658 | Vector ops fail on connections that didn't load it |
|
||||
| Defaulting FastEmbed cache to `/tmp` | #741, #681 | Models re-downloaded every run in sandboxed environments |
|
||||
| File-watcher with no liveness heartbeat | #580, #758, #798 | Silent watcher death, no recovery, missed events under hidden dirs |
|
||||
| Schema-validation `MaxLen` on relation_type | #721 | Valid markdown fails to edit |
|
||||
| FastMCP `AliasChoices` on `bool \| None` params | #818 | Broken JSON schema; external clients silently drop the bool |
|
||||
| `alembic/env.py` setting `BASIC_MEMORY_ENV` at import | #481 | Watch service silently disabled in production |
|
||||
| Inline DDL `ALTER TABLE` at runtime instead of migration | #727 | Postgres deadlocks under concurrent vector sync |
|
||||
| Search index DELETE without `project_id` filter | #256 | Edits in one project erase another project's index rows |
|
||||
|
||||
## What the maintainers' fixes reveal
|
||||
|
||||
- **Defaults flipped under fire**: FastEmbed cache (#741), `bm reset` behavior with live MCP processes (#765 → PR #776 adds psutil guard), `force_full=True` removed from cloud sync (#706, then again #804).
|
||||
- **Multi-project params added retroactively to nearly every tool**: PR #777, #789, #803, #807. The MCP tool surface grew a workspace dimension after launch.
|
||||
- **Aliases were added to make tools "training-data-friendly"** (#766: `find_text`/`old_text`/`search` aliases) and then immediately *broke* `overwrite` (#818). Reverted in #841.
|
||||
- **Docs heavily expanded after confusion**: Postgres setup (#830 still open), `bm cloud setup` (#779 was pointing users at a non-existent command).
|
||||
- **Out-of-scope back-outs**: #720 (visible_project_ids filter) closed not-fix because *"would leak multi-tenant visibility concerns into a local-first single-user product"*.
|
||||
|
||||
## Still-unsolved open issues — and why they're hard
|
||||
|
||||
- **#124 git-based undo** — open 11+ months. Hard because commit cadence is ill-defined and conflicts in markdown trees are nasty.
|
||||
- **#834 local project_id routing in mixed cloud mode** — same root cause keeps spawning new symptoms.
|
||||
- **#382 / #686 large-context handling** — pagination of search results into LLM context windows is still unsolved.
|
||||
- **#740 startup time** — 4.6s for `--help` because FastMCP/onnxruntime/fastembed are imported eagerly. Multi-file lazy-import refactor not shipped.
|
||||
- **#830 / #831 / #829 Postgres + sqlite-vec footguns** — install path still surprises users.
|
||||
- **#669 / #687 transcript-watching sidecar** — the holy grail; nobody has built it.
|
||||
|
||||
## Concrete "do not repeat" lessons for the Rust rewrite
|
||||
|
||||
1. **Do not background-task the indexing pipeline behind the tool reply.** `write_note → return → embed later` makes the tool a liar; the next `search_notes` may miss the entity. (#763, #578, #839, #685). Either make indexing synchronous and bounded, or return a structured `index_status: pending|complete` so the caller can `--wait`.
|
||||
|
||||
2. **Bake the identity dimension in from day one.** `(workspace, project, permalink)` is a 3-tuple. Retrofitting it caused 12+ bugs (#782/#783/#788/#793/#799/#800/#802/#803/#804/#805/#810/#820/#834). The Rust schema should encode the full coordinate at every layer, even when you only ship single-workspace mode.
|
||||
|
||||
3. **Liveness > correctness assumptions for the watcher.** A long-lived file watcher *will* go stale. Build heartbeats, watchdog timers, and "did we miss any events" reconciliation passes from the start. (#580, #758, #798). Treat every notify-rs-style loop as suspect.
|
||||
|
||||
4. **Treat the embedding/vector backend as a fallible plugin, not a default-on assumption.** sqlite-vec, pgvector, FastEmbed, ONNX — every one of them has bitten basic-memory (#735, #767, #741, #681, #830, #831, #658, #829, #578). In Rust, isolate the embedder behind a trait, fail loudly at startup if the backend can't load (don't silently degrade like #578), and don't default cache paths to `/tmp`.
|
||||
|
||||
5. **Make `reset` safe.** Unlinking SQLite while a sibling process holds the inode = mystery phantom search results (#765). Acquire an exclusive advisory lock, or psutil-style live-process check before any destructive op.
|
||||
|
||||
6. **No inline DDL at runtime, ever.** #727's Postgres deadlock from runtime `ALTER TABLE`. Migrations are migrations; never let "ensure table" code paths execute schema changes.
|
||||
|
||||
7. **For the manual-capture problem**: the absence of *explicit* complaints in the tracker is a trap. The complaints are encoded as (a) repeated proposals for a transcript-watching sidecar (#669, #687, #730), (b) Cursor-pollution complaints maintainers close as not-our-problem (#297), and (c) the "what should we commit" indecision in #124. **A Rust rewrite that listens to a Claude Code/Codex transcript directory and writes notes without an `@-mention` would convert the loudest implicit pain into the headline feature.**
|
||||
@@ -0,0 +1,163 @@
|
||||
# cognee — Issue & PR Pain-Point Synthesis
|
||||
|
||||
> Source: GitHub `topoteretes/cognee`. Captured 2026-05-21.
|
||||
> Tracker character: ~40% feature requests (high comment counts), ~40% bugs
|
||||
> closed fast by next release, ~20% truly hard still-open bugs at architectural seams.
|
||||
|
||||
## Top recurring pain points (ranked)
|
||||
|
||||
### A. LLM-adapter brittleness (highest volume, highest churn)
|
||||
|
||||
Every provider has had a wire-level bug in 2026.
|
||||
- Anthropic adapter dropped `max_tokens`, every call HTTP-422 (#2749, #2782 — **two consecutive releases shipped broken Anthropic**).
|
||||
- Ollama / LlamaCpp adapters missing `@observe` decorator (#2820).
|
||||
- vLLM hangs because system message sent after user message (#2537).
|
||||
- vLLM "custom" provider doesn't forward `LLM_ENDPOINT` to LiteLLM (#2412, #2430, #2842).
|
||||
- LiteLLM `model_cost` lookup overrides user's `LLM_MAX_COMPLETION_TOKENS` (#2608, fixed by #2613, #2582).
|
||||
- HTTP/2 stream stalls cause 60s timeouts on sequential Anthropic calls (#2607).
|
||||
- Preflight LLM connection test hangs forever for non-OpenAI providers (#2752, #2123, #2380).
|
||||
- Local OpenAI-compatible LM Studio / Ollama hangs on macOS (#2119, #1743, #1742).
|
||||
- **OPEN, unsolved**: "Severe Performance Degradation Due to Thinking Tokens + Instructor Incompatibility" (#2840). User patch shows the root cause: when `response_model=str`, instructor wraps `str` in a JSON/tool schema that llama.cpp doesn't honor; the LLM returns plain text, instructor fails to parse it, then **tenacity retries sleep 8-128s per attempt**. LiteLLM silently drops non-standard top-level kwargs like `chat_template_kwargs` and `reasoning_effort`, requiring an `extra_body` shim.
|
||||
|
||||
**Root design choice causing this**: LiteLLM + Instructor as the universal LLM gateway. Both churn fast; both silently drop kwargs they don't recognize; both have OpenAI-specific assumptions baked into the structured-output path.
|
||||
|
||||
### B. Multi-store coordination & data integrity bugs
|
||||
|
||||
The "graph + vector + relational" tripartite store is the cognee architectural commitment, and it is the source of consistent, high-severity bugs.
|
||||
- `EntityAlreadyExistsError` — `Entity("institution")` and `EntityType("institution")` collide on UUID (#2510). Second `cognify` blew up.
|
||||
- `add_data_points` parallel DB ops cause SQLite `database is locked` deadlock; **still reproducible in 1.0.2**: `multiple PipelineRunErrored ... elapsed 561s before crash` (#2717, OPEN).
|
||||
- `cognify` exceeds asyncpg bind-argument limit during `upsert_edges` for large batches (~4356 edges) (#2829, fixed via batching in #2798, #2586).
|
||||
- `add_data_points` crashes with `asyncpg.CharacterNotInRepertoireError` on null bytes (0x00) (#2612, OPEN).
|
||||
- `index_data_points`: shallow copy of metadata dict — only the first `index_field` is embedded (#2529, OPEN).
|
||||
- `get_graph_from_model + copy_model` drops user-defined DataPoint id (#2633, OPEN).
|
||||
- Edge deduplication in `retrieve_existing_edges` is non-functional (#2557).
|
||||
- `delete_dataset` fails in non-"public" Postgres schema (#2291).
|
||||
- **Shared data: deleting from one dataset wipes the same data from other datasets that share it** (#2732, OPEN).
|
||||
- N+1 query pattern in `/api/v1/cognify` (#2532, OPEN).
|
||||
- KuzuAdapter writes visible in-memory but not persisted to disk on 0.5.1 (#1981).
|
||||
- Per-collection distance normalization in `brute_force_triplet_search` produces incorrect ranking (#2030, fix #2451 removed normalization, **then #2720 surfaced downstream**).
|
||||
|
||||
### C. Recall quality regressions — the bug a memory server cannot afford
|
||||
|
||||
- **#2720 (OPEN)**: "Graph-completion retrieval returns identical subgraph regardless of query". User-built reproducer shows direct LanceDB queries return different top-K for different queries, but cognee's `/api/v1/search` returns ~identical answers — LanceDB IDs are not propagating into the graph projection. User attributes to fallout from #2451: downstream thresholds in `brute_force_triplet_search.py` still expect pre-#2451 [0,1] scale and silently fall back to an unfiltered graph.
|
||||
- `SearchType.CHUNKS` silently ignores `node_name` filtering (#2815).
|
||||
- `CHUNKS/SUMMARIES/GRAPH_COMPLETION` ignores `datasets=` filter when `ENABLE_BACKEND_ACCESS_CONTROL=false` (#2867 — maintainer answer is essentially "datasets only work with access control on").
|
||||
- Cognee-mcp GRAPH_COMPLETION discarded all but first dataset (#2617).
|
||||
- `TemporalRetriever` is event-only, blocks ontology-wide temporal filtering (#2429, closed as superseded by an internal Q2 redesign — not shipped).
|
||||
- GRAPH_COMPLETION does not search custom DataPoint vector collections (#2495).
|
||||
|
||||
### D. Dependency-installation hell
|
||||
|
||||
- macOS arm64 + Python 3.14: `kuzu` wheel does not exist; quick-start fails (#2753).
|
||||
- `ModuleNotFoundError: No module named 'kuzu' in Docker since v1.0.4` (#2775) — kuzu removed but image not rebuilt.
|
||||
- `fastembed` removed from the Docker image entirely because one of its transitive deps was not Apache/MIT (#2807).
|
||||
- LiteLLMEmbeddingEngine truncated `BAAI/bge-m3` to `bge-m3` (#1915).
|
||||
- `embedding_dimensions` defaulted to 3072 regardless of model — every non-3072-dim embedder broke (#2751, fix #2757).
|
||||
- LanceDB lance-file writer schema drift / "contained null values" RuntimeError bypassed auto-migration (#2702, #2768).
|
||||
- Pydantic v1/v2 friction: upper bound conflicts with openai-agents (#2019); `.json()` deprecated (#2042); generic validation issues (#1198).
|
||||
- Mistral client import error (#2481).
|
||||
- `lru_cache` hash invalidation bugs for Vector and Graph configs (#2357), and **`lru_cache` was eventually disabled outright** in PR #2853 — "refactor: reduce lru cache".
|
||||
|
||||
### E. MCP server bugs vs FastAPI
|
||||
|
||||
The MCP wrapper consistently lags the core API.
|
||||
- MCP `cognify` with valid local file path returns success but creates no Data item (#2250, OPEN).
|
||||
- cognee-mcp Quick Start fails on macOS arm64 + Py 3.14 (#2753).
|
||||
- cognee-mcp `cognify(data=str)` silently dropped all writes after first due to hardcoded `data.txt` filename (#2747).
|
||||
- MCP recall in default config fails: `'NoneType' object has no attribute 'id'` — wrapper does not pass user to cognee.recall (#2855).
|
||||
- cognee-cli `--api-url` doesn't support remember/recall/improve/forget (#2809, OPEN).
|
||||
- Frontend Docker build broken (#2832), Turbopack import case mismatch (#2605), `cognee-cli -ui` v0.5.5 missing 3 npm deps (#2413), UI compile errors (#2709).
|
||||
|
||||
### F. Auth / multi-tenancy regressions
|
||||
|
||||
- Token refresh mechanism literally not implemented; maintainer admits *"we had to reimplement it for our cloud deployment. At this point, we can't allocate resources"* (#2065).
|
||||
- Request-scoped LLM config impossible because `get_llm_config()` and `get_embedding_config()` use `@lru_cache` — singletons (#2228).
|
||||
- Auth disable needs *two* flags: `ENABLE_BACKEND_ACCESS_CONTROL=false` AND `REQUIRE_AUTHENTICATION=False` (#2808, fix #2836).
|
||||
- `cognee.search` ignored ACL when resolving dataset by name for non-owners (#2845); `cognee.add` silently created a new owner-scoped dataset when a non-owner reused a name (#2846); `cognify` silently skipped data added by non-owners (#2847).
|
||||
- Agent display name leaked user ID (#2811).
|
||||
- `GRAPH_DATASET_TO_DATABASE_HANDLER` (user typo of `GRAPH_DATASET_DATABASE_HANDLER`) was silently ignored and defaulted to kuzu (#2697). No validation on env-var names.
|
||||
|
||||
## Design choices that caused the most issues
|
||||
|
||||
1. **Singleton config via `@lru_cache`.** Breaks multi-tenancy (#2228), invalidation bugs (#2357), eventually reverted (#2853).
|
||||
2. **LiteLLM + Instructor as the universal LLM/structured-output layer.** Source of #2412, #2430, #2537, #2608, #2613, #2749, #2782, #2820, #2840, #2842.
|
||||
3. **SQLite as the default relational backend with greenlet parallelism.** #2717: `OperationalError: database is locked` under parallel cognify, OPEN. Maintainer hedges: *"sqlite is not there for production use-cases"*.
|
||||
4. **Kuzu as the default embedded graph DB.** Kuzu was archived upstream (#2098), maintainers chose to replace with **Ladybug** (PR #2755) — a *fork* of Kuzu. Immediate post-replacement bugs: #2768, #2775, WAL corruption (PR #2838). **The forked-DB risk has played out.**
|
||||
5. **LanceDB as the default vector store** with schema migration assumed to auto-handle drift. Reality: #2702 null-values bypasses migration, #2720 retrieval pipeline drops filter on the path from vector hits to graph subgraph.
|
||||
6. **Tripartite store with implicit sync (graph + vector + relational + optional ontology).** Source of the entire integrity-bug class in section B. Orchestration is handled inside cognee, not by any transactional layer.
|
||||
7. **`@lru_cache` + ContextVar mixed model** for tenant isolation. #2228 explains: db config uses ContextVar pattern, but LLM and embedding configs do not.
|
||||
8. **Per-collection distance normalization in `brute_force_triplet_search`** (#2030) — fix #2451 removed normalization, breaking downstream threshold assumptions, surfacing as #2720.
|
||||
9. **Backend access control is the orchestration plane.** When `ENABLE_BACKEND_ACCESS_CONTROL=false`, *all* dataset-scoped retrieval silently degrades. (#2867, #2845, #2846, #2847, #2808.)
|
||||
10. **Default-LLM cost coupling**: `embedding_dimensions` defaults to 3072 (text-embedding-3-large), and litellm's `model_cost` table silently overrode user `max_completion_tokens`. Two bugs from "assume OpenAI defaults" (#2751, #2608).
|
||||
|
||||
## What the maintainers' fixes reveal
|
||||
|
||||
- **LRU cache for configs has been quietly retreated from** (#2853, #2851).
|
||||
- **Subprocess mode + Redis** was added explicitly to escape the SQLite-greenlet trap (#2803, #2812).
|
||||
- **Ladybug is a Kuzu fork** they own. Already shipped "fix: resolve issue with WAL file corruption for ladybug" (#2838).
|
||||
- **Auto-migrate LanceDB on schema drift** had to be added (#2703) after lance-file writer crashed workers.
|
||||
- **Anthropic adapter broken for a full version cycle** — #2749 then re-broken in 1.0.5 as #2782.
|
||||
- **Defaults flipped**: `fastembed` removed from core (#2807), result-cache logging disabled by default (#2851), embedding dimensions auto-derived not defaulted (#2757), auth gated by single switch (#2836).
|
||||
- **Feature deprecated, not fixed**: `TemporalRetriever` (#2429).
|
||||
|
||||
## Open issues maintainers haven't solved
|
||||
|
||||
- **#2717 SQLite deadlock under parallel cognify** — Reproducible across versions.
|
||||
- **#2720 LanceDB filter not propagating to graph projection** — *Correctness* bug in core retrieval path. No assignee.
|
||||
- **#2840 Thinking-token + Instructor incompatibility.**
|
||||
- **#2532 N+1 query in `/api/v1/cognify`.**
|
||||
- **#2612 null-byte crash in asyncpg.**
|
||||
- **#2529 shallow-copy bug in `index_data_points`.**
|
||||
- **#2228 request-scoped LLM/embedding config.** Architectural.
|
||||
- **#2065 token refresh.** Punted to community.
|
||||
|
||||
**What's hard about these**: they all sit at architectural seams (config plane, retrieval pipeline, async orchestration). Not one-PR fixes.
|
||||
|
||||
## Specific dependency culprits (Rust must bet differently)
|
||||
|
||||
| Lib | Bug | Issue |
|
||||
|---|---|---|
|
||||
| litellm | drops `extra_body` kwargs silently; `model_cost` overrides user setting | #2608, #2613, #2840 |
|
||||
| instructor | wraps `response_model=str` in JSON schema local LLMs don't honor | #2840 |
|
||||
| tenacity | 8-128s backoff multiplies instructor's parse failures | #2840 |
|
||||
| asyncpg | bind-arg limit; `CharacterNotInRepertoireError` on `\0` | #2829, #2612 |
|
||||
| lancedb / lance-file | null-values RuntimeError bypasses migration | #2702 |
|
||||
| pyarrow (under lance) | upstream of #2720 / #2702 schema drift | #2702 |
|
||||
| kuzu | upstream archived, wheel gap on Py 3.14 / arm64 | #2098, #2753 |
|
||||
| ladybug (fork of kuzu) | version-mapping crashes on every fresh DB; WAL corruption | #2768, PR#2838 |
|
||||
| sqlite/sqlalchemy/greenlet | database-is-locked under parallel cognify | #2717 |
|
||||
| anthropic SDK | `max_tokens` required, two consecutive releases broken | #2749, #2782 |
|
||||
| fastembed | transitive dep not Apache/MIT; removed from core | #2807 |
|
||||
| pydantic | v1/v2 friction, deprecated `.json()`, upper bound conflicts | #1198, #2019, #2042 |
|
||||
| mistralai | client import error | #2481 |
|
||||
| openai-agents | pin conflict with pydantic | #2019 |
|
||||
| HF tokenizers | every word triggered HF request in chunking | #729 |
|
||||
| Turbopack / npm | frontend builds repeatedly broken | #2605, #2413, #2709, #2832 |
|
||||
|
||||
## Do-not-repeat lessons for the Rust rewrite
|
||||
|
||||
1. **Do not put the LLM call behind a generic Python-style gateway that silently drops kwargs.** Each provider gets a typed Rust client that errors on unknown fields rather than dropping them. (#2840, #2608, #2782.)
|
||||
|
||||
2. **Don't use SQLite for write-parallel pipeline state.** Use Postgres or — for embedded — a single-writer actor in front of an LMDB/Sled/SQLite-with-WAL serialized via a queue. (#2717.)
|
||||
|
||||
3. **Don't pin a forked embedded graph DB as the default.** Either commit to a battle-tested external store (Postgres + AGE, Neo4j) or build the graph primitives directly on the relational store. The Kuzu→Ladybug pivot cost real users (#2098, #2768, #2775, #2753, PR#2838).
|
||||
|
||||
4. **Treat retrieval filter propagation as a first-class invariant with property tests.** Show test cases like `assert different_queries_yield_different_subgraphs`. The exact bug in #2720 is what kills a memory product.
|
||||
|
||||
5. **Configuration must be per-request from day one.** No global singletons, no `lru_cache` over config. Use a request-scoped context type passed explicitly. (#2228, #2357, PR#2853.)
|
||||
|
||||
6. **Never default `embedding_dimensions` to a constant.** Derive from `(provider, model)` at startup; refuse to start if collection-dim and model-dim mismatch. (#2751, #2757.)
|
||||
|
||||
7. **Idempotent ingestion with explicit id derivation.** Node IDs must be a function of `(category, name, dataset)` not just `name`. Property tests on cross-run determinism. (#2510, #2557, #2633.)
|
||||
|
||||
8. **Re-ingestion / pipeline-run status must be a state machine, not a flag.** "PipelineRunAlreadyCompleted" prevented re-ingest of deleted files (#2097).
|
||||
|
||||
9. **Dataset isolation must work whether or not "access control" is on.** Dataset is a hard query-time filter on every retriever. (#2867, #2845, #2846, #2847.)
|
||||
|
||||
10. **Batch every cross-store mutation.** asyncpg bind-arg limit, SQLite locks, lance writer flushes — all rooted in unbounded fan-out (#2829, #2717, #2702).
|
||||
|
||||
11. **Single transactional boundary or a documented eventual-consistency contract.** Cognee's silent skips between graph/vector/relational are the deepest class of bug. Pick one.
|
||||
|
||||
12. **Audit logs and result caches with retention from day one.** Cognee's relational DB grew unbounded — 42k cached results in 9 days (#2548). They eventually disabled by default; do it before the first user hits it.
|
||||
|
||||
**Calibration**: the single most repeated lesson, weighted by both severity and recurrence: **the LLM/structured-output layer (LiteLLM + Instructor) is fragile, and the multi-store sync (graph + vector + relational) is the deepest source of correctness bugs.** A Rust rewrite that gets either of those wrong inherits cognee's tracker.
|
||||
@@ -0,0 +1,131 @@
|
||||
# agentmemory — Research Report
|
||||
|
||||
> Source project: `~/Projects/agentmemory` (TypeScript, MCP server built on `iii-engine`).
|
||||
> This is the user's own existing project. The Rust effort in this repo is a
|
||||
> spiritual successor: keep the *ideas*, replace the *substrate*.
|
||||
|
||||
## 1. Purpose & Scope
|
||||
|
||||
agentmemory is **persistent memory infrastructure for AI coding agents**. The core pitch: an agent silently captures what you do during a coding session (tool calls, prompts, decisions, errors), compresses those raw observations into searchable memory, and re-injects relevant context into the *next* session so the user never has to re-explain architecture, preferences, or past bugs. README pitches a tagline ("Your coding agent remembers everything. No more re-explaining.") with claimed retrieval R@5 of 95.2% on LongMemEval-S vs. 86.2% BM25-only fallback, and ~1,900 tokens/session vs. ~22K for raw CLAUDE.md.
|
||||
|
||||
The author explicitly frames it as the *implementation* of Karpathy's "LLM Wiki" pattern, extended with confidence scoring, lifecycle, knowledge graph, and hybrid search — the project page touts a viral gist with 1200+ stars that articulated the design.
|
||||
|
||||
Caveat about `DESIGN.md`: that file in agentmemory is a Lamborghini-inspired *visual* design system for the marketing site, not architecture. The real architecture docs are in `AGENTS.md` and `README.md`.
|
||||
|
||||
`ROADMAP.md` confirms the trajectory: Q2 2026 "Depth" (multimodal), Q3 "Breadth" (more agents, OpenSSF), Q4 "Trust" (SSO/RBAC), Q1 2027 v1.0 freeze. Candidate item for Q1 2027: *"Reference implementation in a second language (Rust or Go)"* — directly relevant to this project.
|
||||
|
||||
## 2. Architecture
|
||||
|
||||
- **Stack**: TypeScript (ESM, Node ≥ 20), packaged as `@agentmemory/agentmemory`. Build via `tsdown`.
|
||||
- **Not a standalone server**. Everything is built on top of **iii-engine**, a separately-installed Rust binary that runs on `ws://localhost:49134` and provides Worker/Function/Trigger primitives (`AGENTS.md:5`). The Node process registers functions; the engine routes them. This is the project's central architectural bet.
|
||||
- **Storage**: a single **file-based SQLite KV store**, owned by iii-engine's StateModule, not by the Node process. From `iii-config.yaml:11-16`:
|
||||
```yaml
|
||||
- name: iii-state
|
||||
config:
|
||||
adapter: { name: kv, config: { store_method: file_based, file_path: ./data/state_store.db } }
|
||||
```
|
||||
The Node code only sees a tiny shim (`src/state/kv.ts:6-46`) wrapping five RPCs (`state::get/set/list/update/delete`). No direct SQLite, no Postgres, no Qdrant, no graph DB — *all* memory types are stored as JSON values under namespaced "scopes" (e.g. `mem:memories`, `mem:semantic`, `mem:graph:nodes`).
|
||||
- **Indices** live in-process: BM25 (`src/state/search-index.ts`), an in-RAM vector index with cosine similarity (`src/state/vector-index.ts`), and persisted snapshots written back into KV via `IndexPersistence` (`src/state/index-persistence.ts`). Hybrid search uses RRF-style fusion of BM25 + vector + graph (`src/state/hybrid-search.ts`).
|
||||
- **Surfaces**: REST API on `:3111` (124 endpoints — `src/triggers/api.ts`), MCP server on stdio via `npx @agentmemory/mcp` (`src/mcp/server.ts`, ~62KB, with 53 tools in `src/mcp/tools-registry.ts`), live WebSocket stream on `:3112`, and a real-time viewer HTML on `:3113`.
|
||||
- **KV scope catalogue**: `src/state/schema.ts:3-50` lists ~40 scopes (sessions, observations per-session, memories, summaries, semantic, procedural, graph:nodes/edges, insights, lessons, crystals, sketches, sentinels, actions, leases, routines, signals, checkpoints, mesh, slots, retention, accessLog, audit, imageRefs, etc.). The breadth is striking — but it's all one SQLite file.
|
||||
|
||||
## 3. Memory Model
|
||||
|
||||
The system has many memory *types*, organized roughly into a four-tier consolidation hierarchy declared explicitly in `types.ts:429`:
|
||||
```ts
|
||||
export type ConsolidationTier = "working" | "episodic" | "semantic" | "procedural";
|
||||
```
|
||||
|
||||
- **Raw observation** (`RawObservation`, `types.ts:29-42`): captured by hooks on every tool call.
|
||||
- **Compressed observation** (`types.ts:44-62`): structured XML output from an LLM (`mem::compress`) with `type`, `title`, `facts`, `narrative`, `concepts`, `files`, `importance` 1–10. *Critically* this LLM compression is OFF by default (`src/index.ts:245-253`, issue #138). Default path uses **synthetic** compression (`src/functions/compress-synthetic.ts`) — zero LLM calls — to keep token bills sane. Set `AGENTMEMORY_AUTO_COMPRESS=true` to opt in.
|
||||
- **Memory** (`types.ts:81-101`): consolidated long-term entry with `type` (pattern/preference/architecture/bug/workflow/fact), `strength`, `version`, `supersedes`, `isLatest`. Versioned in place — old memory `isLatest=false`, new one keeps `parentId`/`supersedes` chain (`src/functions/consolidate.ts:159-191`).
|
||||
- **SemanticMemory** (`types.ts:435-446`): individual *facts* with confidence + access counts.
|
||||
- **ProceduralMemory** (`types.ts:448-462`): named procedures with steps + trigger condition.
|
||||
- **Lesson / Insight / Crystal**: a higher tier of distillation. Crystals (`crystallize.ts`) summarize chains of completed Actions into narrative + key outcomes + lessons; Lessons feed Reflect which produces Insights from concept clusters.
|
||||
- **MemorySlot** (`types.ts:222-232`, `functions/slots.ts:13-83`): pinned editable text blocks (persona, user_preferences, project_context, guidance, pending_items, etc.) — Karpathy-wiki-style human-editable section. Always injected into context via `src/functions/context.ts:43-61`.
|
||||
|
||||
**Automatic operations (no manual writes needed)**:
|
||||
- `mem::observe` runs from `PostToolUse` hook on every tool call — privacy-stripped, dedup'd, optionally LLM-compressed, indexed, streamed live (`functions/observe.ts:42-280`).
|
||||
- `setInterval` timers in `src/index.ts:491-531`: auto-forget every 1h, lesson decay every 24h, insight decay every 24h, consolidation pipeline every 2h.
|
||||
- `mem::auto-forget` (`functions/auto-forget.ts`): deletes TTL-expired memories, soft-deletes contradiction pairs (Jaccard similarity > 0.9), purges 180-day-old observations with importance ≤ 2.
|
||||
- `mem::retention-score` (`functions/retention.ts:80-94`): retention = `salience * exp(-λ·Δt) + reinforcementBoost(accessLog, σ)`. Below the "cold" threshold (0.15), entries become evictable. The reinforcement boost (`computeReinforcementBoost`) sums `1/daysSinceAccess` — classic spaced-repetition.
|
||||
|
||||
## 4. Reorganization / Consolidation (the Karpathy bit)
|
||||
|
||||
This is the most interesting part. The consolidation pipeline runs on a 2h cron (`src/index.ts:523-531`) when `CONSOLIDATION_ENABLED=true`. `functions/consolidation-pipeline.ts:50-269` orchestrates four tiers:
|
||||
|
||||
1. **Semantic tier**: takes the 20 most recent `SessionSummary` items, asks the LLM to extract `<fact confidence="x">…</fact>` items. Existing facts (matched case-insensitively) get `accessCount++` and `confidence = max(old, new)`; new ones become `SemanticMemory` rows (`consolidation-pipeline.ts:91-122`).
|
||||
2. **Reflect tier**: `mem::reflect` (`functions/reflect.ts`) walks the knowledge graph, builds concept clusters via BFS-by-degree (`buildGraphClusters`, falls back to Jaccard clustering at line 106 if no graph), feeds each cluster's facts + lessons + crystals to the LLM with the `REFLECT_SYSTEM` prompt, expects `<insight>` XML back. Existing insights (fingerprinted on content) get `reinforcements++` and `confidence += 0.1*(1-confidence)` (reflect.ts:26-35); new ones are stored with `decayRate=0.05/week`.
|
||||
3. **Procedural tier**: finds Memory rows of type `pattern` with `frequency >= 2`, extracts named `<procedure>` blocks with steps.
|
||||
4. **Decay tier**: applies geometric decay `strength *= 0.9^decayPeriods` after a configurable inactivity window (`applyDecay` at consolidation-pipeline.ts:21-43).
|
||||
|
||||
Separately, `mem::consolidate` (`functions/consolidate.ts:65-225`) groups observations by concept, picks the top-N most important per concept, and for each cluster either *creates* a Memory or *evolves* an existing one. Evolution = mark old `isLatest=false`, write a new row pointing at it via `supersedes` and `parentId` (lines 161-189). The old memory isn't deleted — it remains versioned-but-shadowed, with `isLatest` filtering at read time.
|
||||
|
||||
`mem::insight-decay-sweep` (`functions/reflect.ts:425-476`) runs weekly: `newConfidence = confidence - decayRate * weeksSince`. Below 0.1 with zero reinforcements, soft-deleted.
|
||||
|
||||
This is genuinely Karpathy-wiki-shaped: **memories aren't append-only — they get rewritten in place via versioned supersession, and unused entries quietly fade.**
|
||||
|
||||
## 5. Agent Integration
|
||||
|
||||
Three surfaces, depending on what the host supports:
|
||||
|
||||
- **Hooks (Claude Code, Codex)**: 12 hook scripts in `src/hooks/` — standalone Node scripts that read JSON from stdin and POST to `/agentmemory/observe` over HTTP with a 3s timeout. `plugin/hooks/hooks.json` registers all 12 (SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, PostToolUseFailure, PreCompact, SubagentStart/Stop, Notification, TaskCompleted, Stop, SessionEnd). Codex gets a 6-hook subset (`README.md:419`).
|
||||
- **MCP tools**: 53 tools in `src/mcp/tools-registry.ts` — `memory_recall`, `memory_smart_search`, `memory_save`, `memory_sessions`, `memory_consolidate`, `memory_action_create`, `memory_lease`, `memory_signal_send`, `memory_crystallize`, etc. Only 8 are visible by default (`AGENTMEMORY_TOOLS=all` exposes the rest, AGENTS.md:114).
|
||||
- **Skills** (slash commands): `plugin/skills/{recall,remember,handoff,recap,forget,session-history,commit-context,commit-history}/SKILL.md` — each is a markdown file with frontmatter the host parses.
|
||||
- **System-prompt injection**: pre-tool-use and session-start hooks can write up to 4000 chars of memory context to stdout, which Claude Code prepends to the next turn. **Off by default** (`AGENTMEMORY_INJECT_CONTEXT=false`) because of token-burn complaints (#143, see `src/hooks/pre-tool-use.ts:9-22`).
|
||||
|
||||
## 6. Cross-Agent Handoff
|
||||
|
||||
Two mechanisms:
|
||||
|
||||
- **`/handoff` skill** (`plugin/skills/handoff/SKILL.md`): finds the most recent session whose `cwd` matches your current directory (with proper path-boundary check), surfaces any unanswered question first, then fetches a recall of top concepts. This works *because all agents write to the same `:3111` server* — a Claude Code session and a Codex session in the same project share KV. Cross-agent recall is implicit.
|
||||
- **Signals** (`functions/signals.ts`): a typed inter-agent message bus with `from`/`to`/`threadId`/`type: handoff|request|response|alert|info` and TTLs. `mem::signal-send` + `mem::signal-read` MCP tools.
|
||||
- **Mesh sync** (`functions/mesh.ts`) does LWW-merge replication between agentmemory *servers* on different machines, but cross-agent on a single box is just the shared SQLite.
|
||||
|
||||
Roadmap explicitly flags "Cross-agent shared memory namespace" as Q3 2026 candidate and "Agent-to-agent memory handoff protocol" for Q4 2026 — so the current support is informal.
|
||||
|
||||
## 7. Self-Healing / Operations
|
||||
|
||||
- **Backup**: `mem::snapshot-create` (`functions/snapshot.ts:44-150`) dumps sessions + memories + graphNodes + observations + accessLogs to a `state.json` and `git commit`s it into a configurable directory. Configurable interval cron. Each snapshot has a SnapshotMeta with commitHash. There's a `mem::snapshot-diff` too. Git-as-backup is clever.
|
||||
- **Migration**: `mem::migrate` (`functions/migrate.ts`) reads legacy `better-sqlite3` DBs and replays them through the KV layer, with a path-allowlist guard (only under `~/.agentmemory/`).
|
||||
- **Schema versioning**: an explicit `ExportData.version` union with 50+ string literals (`types.ts:296`) and `supportedVersions` in `export-import.ts`. Updating it requires editing 6 files (AGENTS.md:28-34).
|
||||
- **Vector index corruption guard**: at boot, `src/index.ts:367-410` *refuses to start* if persisted vectors have mixed dimensions, with a structured remediation message. Beautiful defensive code.
|
||||
- **Diagnostics**: `mem::diagnose` (`functions/diagnostics.ts`) runs 8 category checks (orphan leases, blocked actions with all deps done, dead sentinels...) and `mem::heal` auto-fixes the fixable ones.
|
||||
- **Audit**: every state-changing operation records to `KV.audit` with a typed operation union (`types.ts:493-539`, ~45 op types). `AGENTS.md:39-41` makes adding new ones mandatory.
|
||||
- **Resilience**: top-level `unhandledRejection` swallow with throttled logging (`src/index.ts:118-128`) to survive iii-engine timeout spikes (#204). Circuit-breaker + fallback-chain providers (`src/providers/`).
|
||||
|
||||
## 8. What's Good / What's Missing — Honest Take
|
||||
|
||||
**What's clever and worth reusing:**
|
||||
|
||||
- **Two-tier compression (synthetic default, LLM opt-in)**: the #138 fix is the right call. Default zero-LLM keeps token bills sane while preserving BM25/vector search. Critical lesson — don't make the LLM call mandatory.
|
||||
- **Versioned in-place memory evolution** (`isLatest` + `supersedes` chain): exactly the Karpathy wiki-rewrite pattern. Cheaper than maintaining a separate "graveyard", and `parentId` keeps history queryable.
|
||||
- **Retention-as-formula** rather than rules: `salience * exp(-λΔt) + Σ(σ/daysSinceAccess)`. Tunable, principled, batchable. Worth borrowing wholesale.
|
||||
- **Triple-stream RRF retrieval** (BM25 + vector + graph): graph-walk-as-third-signal is unusual and the benchmark numbers (95.2% R@5) suggest it pays off.
|
||||
- **Git-as-snapshot-backup**: dump state.json, `git commit`. Diffable, restorable, no DB-specific tooling needed.
|
||||
- **Slots**: pinned, human-editable, always-injected wiki-style sections. The user can hand-edit `project_context` and it survives forever — a clean escape hatch from purely-emergent memory.
|
||||
- **Hook scripts as standalone Node files** (no SDK import, just HTTP): fast startup, fault-tolerant. The 800ms-1500ms hard timeouts (`src/hooks/session-start.ts:27-28`) are explicit — there's a war story about #221 where unbounded hook fan-out OOM-killed iii-engine.
|
||||
- **Privacy filter at observe boundary** (`stripPrivateData` before any persistence): defense in depth.
|
||||
|
||||
**What feels overengineered:**
|
||||
|
||||
- **~40 KV scopes, 53 MCP tools, 124 REST endpoints, 50+ iii functions** is a *lot* of surface for v0.9. Many feel speculative (sentinels, sketches, frontier, leases, routines, checkpoints, facets, crystals, mesh, branch-aware, flow-compress, vision-search). The AGENTS.md "you must update ALL of the following" 7-step checklists are a smell — the system is so wide that ordinary changes ripple far.
|
||||
- **iii-engine dependency**: every install requires a separate Rust binary, pinned to a specific version (0.11.2; 0.11.6 broke them), with no canonical Windows installer (`README.md:549`). For a Rust rewrite this is a *strong argument to be self-contained*: embed SQLite directly, drop the engine.
|
||||
- **All-JSON storage under one big SQLite file**: every "memories" or "graph nodes" list operation is `state::list` → return *all* JSON, parse in Node, filter in memory. `auto-forget.ts:67` literally caps at 1000 latest memories because the O(N²) Jaccard loop. Won't scale past ~10K memories without proper indexing.
|
||||
- **XML-in-LLM-output everywhere**: `<memory>...</memory>`, `<temporal_graph>`, `<insight>`. Fragile regex parsing (`parseCompressionXml`, `parseTemporalGraphXml`). A schema-validated JSON / structured-outputs path would be more robust.
|
||||
- **DESIGN.md is for the website**, not the architecture. Architecture is scattered across AGENTS.md, README.md, and inline comments. No real `ARCHITECTURE.md`.
|
||||
|
||||
**What's missing that a Rust competitor should improve:**
|
||||
|
||||
1. **First-class embedded store.** Use `rusqlite` (or `sqlx` + SQLite/Postgres) with *proper indices and FTS5* instead of JSON-blob-in-KV. The retention/auto-forget logic deserves real WHERE clauses, not in-memory filter passes. Look at `sqlite-vec` or `lancedb` for vectors so you're not maintaining a Map<String, Vec<f32>>.
|
||||
2. **Self-contained binary.** No separate engine. A single `agentmemory` binary that *is* the MCP server, REST server, and storage. The iii-engine indirection costs operability for ~nothing the user sees.
|
||||
3. **Native MCP transport.** Use a proper Rust MCP SDK; you don't need an HTTP shim between hooks and storage — hooks can speak MCP-over-stdio or a Unix socket. Cuts the 3s HTTP timeout and the auth bearer dance.
|
||||
4. **First-class structured outputs.** Don't parse `<memory><type>...</type></memory>` regex-strings. Use JSON-schema constrained generation; `serde_json` deserialization with strict validation.
|
||||
5. **Smaller, sharper tool surface.** Pick the 8–12 tools that demonstrably matter (recall, save, sessions, smart_search, handoff, consolidate, forget, governance_delete) and ship those well. The 53-tool surface invites confusion; the README admits 8 are visible by default for a reason.
|
||||
6. **Real graph store for temporal edges.** The temporal-graph design (`tvalid`, `tvalidEnd`, `supersededBy`, edge-history) is *good* — but implementing it on JSON blobs in KV means every "what does Alice prefer as of 2024-06-15" query is a full table scan. Consider `petgraph` in memory backed by a real SQL graph table, or a proper graph DB if you're going big.
|
||||
7. **Cross-agent handoff as a designed protocol**, not just shared SQLite. The Q4 2026 roadmap candidate is exactly this — get there day one. Define a `Handoff` type with from-agent, to-agent, context, open-questions, files-touched, model-used; expose `mcp::handoff/begin` and `mcp::handoff/accept`.
|
||||
8. **Reproducible eval harness.** They ran LongMemEval-S and got 95.2% R@5 — bake the harness into CI from the start so regressions are caught.
|
||||
9. **Single source of truth for schema.** The AGENTS.md "update all 7 files" checklists indicate the schema is duplicated across types, tools-registry, REST, MCP, tests, README, plugin. In Rust, one `proto`-or-derive-driven definition can generate all of those.
|
||||
10. **A clear `ARCHITECTURE.md`.** Don't repeat their mistake of scattering architecture across AGENTS, README, and inline comments.
|
||||
|
||||
The big takeaway: **agentmemory's *concepts* are unusually thoughtful** — versioned supersession, retention formulas, slot pinning, hybrid retrieval, opt-in LLM compression, session-cwd-based handoff. The *implementation* is constrained by the iii-engine bet and the all-JSON-in-KV storage choice. A Rust rewrite has a real opportunity to keep the ideas and drop the substrate.
|
||||
@@ -0,0 +1,107 @@
|
||||
# basic-memory — Research Report
|
||||
|
||||
> Source project: `~/Projects/basic-memory` (Python, MCP-native, markdown-on-disk).
|
||||
> Studied as inspiration *and* as the manual-write-note model we explicitly
|
||||
> diverge from.
|
||||
|
||||
## 1. Purpose & Scope
|
||||
|
||||
Basic Memory is a **local-first, MCP-native, Markdown-based personal knowledge graph**. Its tagline is "Your AI never forgets again": notes live as plain Markdown files on disk, both humans (in Obsidian, VS Code) and LLMs (over MCP) read and write them, and a SQLite/Postgres index keeps a knowledge graph in sync.
|
||||
|
||||
**The model is explicitly *manual*.** The user (or, more often, the agent in response to the user) calls `write_note`, `edit_note`, `move_note`, etc. There is **no implicit capture** of conversation content. The "How it works" example in `README.md:317-356` is literally *"Ask the LLM to capture it: 'Make a note on coffee brewing methods.'"* The README even prescribes user prompts ("Create a note about our project architecture decisions") as the activation step. Nothing in the codebase auto-saves conversation turns. The closest thing to automation is the `continue_conversation` *prompt* (`src/basic_memory/mcp/prompts/continue_conversation.py:18-90`), which only *retrieves* — it asks the model to search recent activity and load context; it never writes.
|
||||
|
||||
## 2. Storage Model
|
||||
|
||||
**Both Markdown and SQL, with the files being source-of-truth.** Each note is one Markdown file with YAML frontmatter:
|
||||
|
||||
```yaml
|
||||
---
|
||||
title: Coffee Brewing Methods
|
||||
type: note
|
||||
permalink: coffee-brewing-methods
|
||||
tags: [coffee, brewing]
|
||||
---
|
||||
```
|
||||
|
||||
The grammar is documented in `docs/NOTE-FORMAT.md` and parsed at `src/basic_memory/markdown/entity_parser.py:1-27` using `markdown-it` plus custom `observation_plugin` and `relation_plugin`. Three semantic primitives:
|
||||
|
||||
- **Entity** (one per file): `src/basic_memory/models/knowledge.py:28-149`. Holds `title`, `note_type`, `permalink`, `file_path`, `checksum`, `mtime`, `size`, `entity_metadata` (JSON for custom frontmatter), `external_id` (stable UUID).
|
||||
- **Observation**: `- [category] text #tag (context)` lines, indexed in the `observation` table (`knowledge.py:220-263`).
|
||||
- **Relation**: `- relation_type [[Other Entity]]` lines, indexed in `relation` (`knowledge.py:265-311`). Bare `[[X]]` becomes `links_to`. Relations can be *unresolved* (`to_id` NULL until the target exists), then auto-resolved on sync.
|
||||
|
||||
**Search** is dual-stack and selected by `database_backend` config (`src/basic_memory/config.py:222-226`):
|
||||
|
||||
- **SQLite**: FTS5 virtual table `search_index` with custom tokenizer `'unicode61 tokenchars 0x2F'` so `/` is searchable (`src/basic_memory/models/search.py:62-94`). Semantic vectors via `sqlite-vec` virtual table `search_vector_embeddings` (`search.py:146-153`).
|
||||
- **Postgres**: real `search_index` table with `tsvector` GIN + `pgvector` (`search.py:17-58`) and `pg_trgm` for fuzzy link resolution (migration `f8a9b2c3d4e5`).
|
||||
|
||||
Hybrid search defaults: vector candidates `semantic_vector_k=100`, similarity threshold `0.55`, model `bge-small-en-v1.5` via FastEmbed (`config.py:233-313`). Search type is `hybrid` when semantic is on, else `text` (`config.py:313-318`).
|
||||
|
||||
There is also a `NoteContent` table (`knowledge.py:152-217`) that **materializes the markdown body in the DB** with a `file_write_status` state machine (`pending|writing|synced|failed|external_change_detected`) and `db_version`/`file_version` for conflict resolution between AI writes and human file edits.
|
||||
|
||||
## 3. MCP Tools Exposed
|
||||
|
||||
All tools are registered via the `@mcp.tool` decorator and exported from `src/basic_memory/mcp/tools/__init__.py:9-65`. Every tool is annotated with MCP `readOnlyHint`/`destructiveHint`/`idempotentHint`/`openWorldHint` so agents can pick safely. Every single tool requires **explicit invocation** — none are triggered by the server itself.
|
||||
|
||||
| Tool | Hint | File |
|
||||
|---|---|---|
|
||||
| `write_note` | destructive, not idempotent | `tools/write_note.py:21` |
|
||||
| `edit_note` (append/prepend/find_replace/replace_section/insert_*) | not destructive | `tools/edit_note.py:225` |
|
||||
| `read_note`, `view_note`, `read_content` | read-only | `tools/read_note.py:64`, `view_note.py:13`, `read_content.py:157` |
|
||||
| `delete_note` | destructive | `tools/delete_note.py:184` |
|
||||
| `move_note` | not destructive (updates links) | `tools/move_note.py:346` |
|
||||
| `search_notes` (advanced: tags, status, metadata_filters, after_date, search_type=text/vector/hybrid) | read-only | `tools/search.py:564` |
|
||||
| `recent_activity` | read-only | `tools/recent_activity.py:28` |
|
||||
| `build_context` (resolves `memory://` URIs, walks relation graph N hops, `depth=1..3`) | read-only | `tools/build_context.py:114` |
|
||||
| `list_directory`, `canvas` (Obsidian canvas), `list_workspaces` | mixed | various |
|
||||
| `list_memory_projects`, `create_memory_project`, `delete_project` | mixed | `tools/project_management.py` |
|
||||
| `schema_infer`, `schema_validate`, `schema_diff` (Picoschema over frontmatter) | read-only | `tools/schema.py:206-440` |
|
||||
| ChatGPT-compat `search` / `fetch` | read-only | `tools/chatgpt_tools.py:107-171` |
|
||||
| `cloud_info`, `release_notes` | read-only | |
|
||||
|
||||
Two MCP **prompts** ship: `continue_conversation` and `recent_activity` (both retrieve-only). There is also a `view_note` UI artifact path.
|
||||
|
||||
## 4. Memory Lifecycle
|
||||
|
||||
**There is no lifecycle.** Grep'd the entire `src/` tree for `decay|aging|consolidat|summariz|forget|expire|ttl|prune|archive_old` — zero matches outside of unrelated OAuth-token expiry and SQLAlchemy `expire_on_commit`. Notes are **append-only forever** unless a human or agent explicitly calls `delete_note`, `move_note`, or `edit_note`. There is no auto-summarization, no auto-merge of duplicates, no recency weighting in search ranking (only an `after_date` filter), no "cold storage" tier. `recent_activity` (`tools/recent_activity.py:28`) just queries by `created_at`/`updated_at`; it doesn't shape memory.
|
||||
|
||||
The only background process is the `WatchService` (`src/basic_memory/sync/watch_service.py:81-145`) + `SyncService` (`src/basic_memory/sync/sync_service.py:153-188`), which reconcile file changes with the DB after a `sync_delay` of 1000 ms (`config.py:338`). That's housekeeping, not lifecycle.
|
||||
|
||||
## 5. Cross-Project / Cross-Agent
|
||||
|
||||
Projects are a first-class concept. Config holds a `Dict[str, ProjectEntry]` (`config.py:184-195`), each with a `path`, a `mode` (`LOCAL` or `CLOUD` — per-project routing), optional `workspace_id`, and bisync state. A `default_project` is auto-set to the first project (`config.py:705-711`).
|
||||
|
||||
**Project resolution is a unified three-tier chain** (`docs/ARCHITECTURE.md:298-324`): explicit `project` argument → default → single-project fallback. Every MCP tool accepts `project` and `project_id` (UUID) parameters. `get_project_client(project, ...)` (`mcp/project_context.py`) routes to local ASGI or cloud HTTP per-project, so you can mix.
|
||||
|
||||
For agent handoffs, there is no special handshake — basic-memory treats all MCP clients identically. The "handoff" story is: agent A writes notes to project `foo`, agent B (any other MCP client pointed at the same project) reads them. The Markdown files on disk are the lingua franca. There is no session/agent identifier, no per-agent scratchpad. `created_by`/`last_updated_by` columns exist (`knowledge.py:99-102`) but are populated only in cloud (user_profile_id), null for local/CLI.
|
||||
|
||||
## 6. Backup & Portability
|
||||
|
||||
- **Default location**: `~/basic-memory` for notes (per `BASIC_MEMORY_HOME` env, README:197), `~/.basic-memory/` for SQLite DB + config (`config.py:24, 67-80`). Config in `~/.basic-memory/config.json`, chmod 0600.
|
||||
- **Portability**: Excellent for files (they're just Markdown — `git clone`, `rsync`, Syncthing, rclone all work). The DB is a derived index — `bm sync` rebuilds it from the files.
|
||||
- **Schema migrations**: 22 Alembic migrations in `src/basic_memory/alembic/versions/`, auto-run on startup via `get_or_create_db` (`services/initialization.py:23-38`). Migrations cover both SQLite and Postgres.
|
||||
- **Importers** for Claude conversations, ChatGPT exports, and `memory.json` (the original MCP "memory" server format) live in `src/basic_memory/importers/`.
|
||||
- **Cloud sync** uses `rclone bisync` (`config.py:136-138, 164-167`), not a custom protocol — another portable choice.
|
||||
|
||||
## 7. Strengths Worth Borrowing
|
||||
|
||||
1. **Files are source of truth, DB is derived index.** Survives any DB corruption, plays nice with git, version control, and grep. Bidirectional human/AI editing via file-watcher + checksum.
|
||||
2. **MCP behavior annotations on every tool** (`readOnlyHint`, `destructiveHint`, etc., `tools/write_note.py:23`). Agents can plan multi-step actions without trial-and-error.
|
||||
3. **Aggressive `AliasChoices` aliasing** of parameter names (`write_note.py:31` accepts `directory|folder|dir|path`; `search.py:574` accepts `query|q|search|text`). LLMs use whatever name their training reaches for; the tool absorbs the variance.
|
||||
4. **`memory://` URI scheme + `build_context` graph walker** (`tools/build_context.py:114-247`) — turns wiki-links into navigable context. Cleaner than dumping the whole graph.
|
||||
5. **Unresolved relations as first-class state** (`knowledge.py:282`) — `to_id` nullable, resolves later when target appears. Forward references just work.
|
||||
6. **Per-project routing** with mixed local/cloud modes per project, not per-server.
|
||||
7. **Composition root + typed-client pattern** (`docs/ARCHITECTURE.md:14-256`) keeps MCP/CLI/API entrypoints clean and testable.
|
||||
8. **The `NoteContent` table** with `file_write_status` state machine (`knowledge.py:155-217`) handles the race between agent writes and on-disk human edits — worth replicating if you keep a DB cache.
|
||||
|
||||
## 8. Weaknesses / Friction (Avoid)
|
||||
|
||||
1. **The manual `write_note` ceremony is the headline friction.** Users must explicitly tell the model "make a note about this," and the model must decide *title*, *directory*, *tags*, *note_type*, and the semantic observation/relation grammar — all on every call. The `write_note` signature has 11 parameters (`write_note.py:25-45`). Skip this entirely: capture should be ambient (post-turn summarization, automatic salience scoring, etc.) — not a tool the model must remember to call.
|
||||
2. **Append-only forever.** No decay, no consolidation, no automatic deduplication. Long-running graphs accumulate cruft. Recent v0.20 added a guard so `write_note` *errors* on conflict instead of silently upserting (`write_note.py:240-262`), which protects data but pushes the burden back onto the agent to manage identity. For agent long-term memory, decay/merge/consolidation is essential and absent here.
|
||||
3. **The semantic grammar is human-authored convention.** `- [category] text #tag (context)` and `- relation_type [[Target]]` are intuitive for humans in Obsidian but require the LLM to *generate* this format correctly every time. Drift is normal. A Rust server can store edges natively and let the LLM emit prose.
|
||||
4. **Note identity is brittle.** `permalink` is derived from title/path; renames create work (`update_permalinks_on_move` defaults to `False`, `config.py:349-352`). The 11-column `entity` table + separate `note_content` is a lot of machinery to keep in sync with files.
|
||||
5. **No agent/session model.** No notion of "who wrote this," "which session," or "what was the user's intent." `created_by` exists only for cloud auth (`knowledge.py:99-102`). For multi-agent handoffs, you'd want provenance baked in.
|
||||
6. **Search ranking is keyword-or-vector with a fixed `min_similarity=0.55`** (`config.py:307-312`). No recency/importance reranking, no usage feedback, no per-query learning.
|
||||
7. **The Markdown source-of-truth tradeoff.** Filesystem latency, checksum recomputation, FTS5 rebuilds, and circuit-breaker retry tracking (`sync_service.py:179-281`) are a lot of infrastructure just to keep a DB consistent with files. For agent memory where the files exist only because LLMs wrote them, this is overhead with no payoff — keep the DB as primary and only export Markdown on demand.
|
||||
8. **Tool surface is wide (~25 tools).** Agents have to pick among `write_note`/`edit_note`/`move_note`/`delete_note`/`read_note`/`view_note`/`read_content`/`search`/`search_notes`/`build_context`/`recent_activity`/`list_directory`/`list_memory_projects`/`canvas`/`schema_*`. Even with hints, that's a lot of context burned describing tools. Aim for a smaller, more orthogonal set.
|
||||
|
||||
**Bottom line**: borrow the *graph + observations + wiki-link* primitives, the *memory:// URI* navigation, the *typed-client + composition-root* layering, the *MCP annotations*, and the *file-as-portable-export* idea — but invert the capture model (ambient, not invoked), add a lifecycle (decay/consolidation/summarization), keep the DB primary, add agent/session provenance, and ship a much narrower MCP tool surface.
|
||||
@@ -0,0 +1,158 @@
|
||||
# cognee — Research Report
|
||||
|
||||
> Source project: `~/Projects/cognee` (Python, knowledge-graph + vector + relational, MCP server).
|
||||
|
||||
## 1. Purpose & Scope
|
||||
|
||||
Cognee bills itself as **"the brain behind your agents" — a memory control plane** (`README.md:40,67`). It ingests heterogeneous data (text, PDF, CSV, code, web pages), then continuously builds a hybrid **knowledge graph + vector index + relational catalog** so that agents can retrieve context by both meaning (embeddings) and structure (graph relationships). The current public SDK is intentionally minimal — four verbs: `remember`, `recall`, `forget`, `improve` (README L127). Internally, those map to the older `add`/`cognify`/`search`/`prune` primitives. Cognee positions itself between a RAG retriever and a "company brain," with multi-tenant isolation, ontology grounding, and a Claude Code plugin for capturing agent traces.
|
||||
|
||||
## 2. The Cognify Pipeline (heart of the system)
|
||||
|
||||
The pipeline is composed as a **list of `Task` objects** executed by `run_pipeline`. Canonical definition at `cognee/api/v1/cognify/cognify.py:316-344`:
|
||||
|
||||
```python
|
||||
default_tasks = [
|
||||
Task(classify_documents), # EXTRACT
|
||||
Task(extract_chunks_from_documents, ...), # EXTRACT
|
||||
Task(extract_graph_and_summarize, graph_model=..., ...), # COGNIFY (LLM)
|
||||
Task(add_data_points, embed_triplets=embed_triplets, ...), # LOAD
|
||||
Task(extract_dlt_fk_edges), # LOAD (relational FK edges)
|
||||
]
|
||||
```
|
||||
|
||||
Step-by-step:
|
||||
|
||||
1. **Classify documents** (`cognee/tasks/documents/classify_documents.py`). Wraps each raw `Data` row in a typed `Document` subclass (`PdfDocument`, `TextDocument`, `DltRowDocument`, etc.) so downstream chunkers know how to read content.
|
||||
|
||||
2. **Chunk documents** (`cognee/tasks/documents/extract_chunks_from_documents.py` + `cognee/modules/chunking/TextChunker.py:1-60`). The default `TextChunker` uses `chunk_by_paragraph` to pack paragraphs up to `max_chunk_size` tokens (calculated as `min(embedding_max_completion_tokens, llm_max_completion_tokens // 2)`). Each chunk becomes a `DocumentChunk` DataPoint with `metadata={"index_fields": ["text"]}` — those `index_fields` are how the storage layer later knows what to embed.
|
||||
|
||||
3. **Extract graph + summarize** (`cognee/tasks/graph/extract_graph_and_summarize.py:21-37`). Two LLM tasks concurrently on every chunk via `asyncio.gather`:
|
||||
- `extract_graph_from_data` (`cognee/tasks/graph/extract_graph_from_data.py:128-222`) calls `extract_content_graph(chunk.text, graph_model, custom_prompt)` for each chunk — an `instructor`/`litellm`-backed structured call that returns a Pydantic `KnowledgeGraph` of `Node`/`Edge`. Edges with missing source/target are filtered (L181-188). Entity nodes are then validated against an ontology via `expand_with_nodes_and_edges(..., ontology_resolver, ...)` (L110-112). Provenance is stamped onto every DataPoint (`_stamp_provenance_deep`, L30-53). Existing edges are deduplicated via `retrieve_existing_edges`.
|
||||
- `summarize_text` produces a `TextSummary` for each chunk used by later "SUMMARIES" search.
|
||||
|
||||
4. **Persist nodes, edges & embeddings** (`cognee/tasks/storage/add_data_points.py:30-149`). `get_graph_from_model` recursively walks the Pydantic graph into `(nodes, edges)` tuples, then `deduplicate_nodes_and_edges` removes duplicates. The pipeline branches on `EngineCapability.HYBRID_WRITE`:
|
||||
- Hybrid backend (e.g. Postgres+pgvector): `graph_engine.add_nodes_with_vectors(nodes)` in one transaction.
|
||||
- Otherwise, parallel writes: `graph_engine.add_nodes(nodes)` + `index_data_points(...)` to the vector engine.
|
||||
- If `embed_triplets=True`, builds `(source -› relation -› target)` text strings and embeds them as additional `Triplet` DataPoints (L184-265). This is what makes graph-walk retrieval finable by similarity.
|
||||
|
||||
5. **DLT FK edges** — adds deterministic foreign-key edges for ingested SQL/DLT-backed datasets without LLM cost.
|
||||
|
||||
There's also a parallel **temporal** pipeline (`cognify.py:347-392`) that swaps the graph extractor for `extract_events_and_timestamps` → `extract_knowledge_graph_from_events` to build a time-aware graph.
|
||||
|
||||
## 3. Storage Backends — Triple-Stack, Pluggable
|
||||
|
||||
Cognee always uses three stores in parallel, abstracted by a `UnifiedStoreEngine` (`cognee/infrastructure/databases/unified/unified_store_engine.py:11-66`):
|
||||
|
||||
| Layer | Default | Other supported |
|
||||
|---|---|---|
|
||||
| **Graph** | `ladybug` (Kuzu fork, file-based) | `neo4j`, `postgres` (AGE-style), `kuzu`, `ladybug-remote`, `neptune`, `neptune_analytics` (hybrid) |
|
||||
| **Vector** | `lancedb` (file-based, subprocess-isolated) | `pgvector`, `chromadb`, `neptune_analytics` |
|
||||
| **Relational** | `sqlite` (aiosqlite) | `postgres` via SQLAlchemy async |
|
||||
|
||||
Selection is **environment-driven**: `GRAPH_DATABASE_PROVIDER`, `VECTOR_DB_PROVIDER`, `DB_PROVIDER`, with normalized credentials. There is also a `USE_UNIFIED_PROVIDER=pghybrid` short-circuit that makes a single Postgres instance back graph+vector+relational simultaneously.
|
||||
|
||||
Two `EngineCapability` flags (`HYBRID_WRITE`, `HYBRID_SEARCH`) let the pipeline branch on whether nodes+vectors can land in one transaction (saves a round trip) or need two writes.
|
||||
|
||||
The default local-only stack is therefore: **SQLite + LanceDB + Ladybug/Kuzu**, all file-based — no servers required. For homelab parity, that's the lean setup.
|
||||
|
||||
## 4. Search / Retrieval
|
||||
|
||||
Recall is **multi-strategy and auto-routed**. `SearchType` enum (`cognee/modules/search/types/SearchType.py`) lists 16 strategies: `GRAPH_COMPLETION` (default), `GRAPH_COMPLETION_COT`, `GRAPH_COMPLETION_CONTEXT_EXTENSION`, `GRAPH_COMPLETION_DECOMPOSITION`, `GRAPH_SUMMARY_COMPLETION`, `RAG_COMPLETION`, `TRIPLET_COMPLETION`, `CHUNKS`, `CHUNKS_LEXICAL`, `SUMMARIES`, `CYPHER`, `NATURAL_LANGUAGE`, `TEMPORAL`, `FEELING_LUCKY`, `CODING_RULES`, `AGENTIC_COMPLETION`.
|
||||
|
||||
The `recall()` API (`cognee/api/v1/recall/recall.py:314-513`) picks one via `route_query(query_text)` — a **rule-based classifier** in `query_router.py` (regex patterns for "when/before/after" → `TEMPORAL`, keyword fragments → `CHUNKS_LEXICAL`, multi-hop wording → `GRAPH_COMPLETION_COT`, etc.). Default fallback is `GRAPH_COMPLETION`.
|
||||
|
||||
The flagship retriever is `GraphCompletionRetriever` (`cognee/modules/retrieval/graph_completion_retriever.py`). It uses **`brute_force_triplet_search`** (`cognee/modules/retrieval/utils/brute_force_triplet_search.py:216-355`), which:
|
||||
|
||||
1. Embeds the query, runs **vector search across multiple collections in parallel**: `Entity_name`, `TextSummary_text`, `EntityType_name`, `DocumentChunk_text`, `EdgeType_relationship_name` (L281-290).
|
||||
2. Projects results into a `CogneeGraph` memory fragment, scoring triplets by combined node+edge similarity with `triplet_distance_penalty` and `feedback_influence` (L318-333).
|
||||
3. Optionally expands a `neighborhood_depth` hop-out from top seed nodes.
|
||||
4. Resolves the top-K edges to natural-language sentences (`resolve_edges_to_text`).
|
||||
5. Feeds them into an LLM completion with `graph_context_for_question.txt` as the user prompt.
|
||||
|
||||
Recall also supports **session-cache first-pass** (keyword match against recent QA entries in a relational session table) with **fall-through to graph** when no session hit (`recall.py:382-397, 447-457`). That's the "hybrid working-memory + long-term memory" pattern.
|
||||
|
||||
## 5. Memory Lifecycle — Cognee Does Reorganize
|
||||
|
||||
Cognee has explicit **memory enrichment** beyond one-shot ingest. The `improve()` API (`cognee/api/v1/improve/improve.py:36-232`) runs up to five stages:
|
||||
|
||||
1. **Feedback weights**: session entries with thumbs-up/down ratings adjust `feedback_weight` on the **specific graph nodes/edges that were used** to answer (`apply_feedback_weights_pipeline`, L284-299). Tracked via `used_graph_element_ids` recorded at retrieval time. Higher-rated answers boost their source nodes; lower-rated ones decrease them.
|
||||
2. **Persist session Q&A**: cognifies session transcripts into permanent graph tagged `node_set="user_sessions_from_cache"`.
|
||||
3. **Triplet enrichment / memify**: `cognee/memify_pipelines/create_triplet_embeddings.py` builds and embeds new triplet datapoints.
|
||||
4. **Global context index**: `global_context_index_pipeline` builds bucket+root summaries over all text summaries.
|
||||
5. **Sync graph→session**: incrementally copies recently-added graph edges into session caches as JSON-lines so live agents pick up new knowledge without a re-query.
|
||||
|
||||
There's also a **`consolidate_entity_descriptions.py`** pipeline (`cognee/memify_pipelines/consolidate_entity_descriptions.py`) that walks Entity nodes, fetches neighbors+edges, and rewrites their `description` field via an LLM (`NodeDescription` Pydantic). And `apply_frequency_weights.py` ages knowledge by usage frequency. Deduplication happens at write time in `add_data_points.py` and at extraction time in `retrieve_existing_edges`.
|
||||
|
||||
So: **not one-shot.** There's a clear feedback loop, summary consolidation, and edge-aging concept. There is no automatic decay/TTL, but `feedback_weight` and `frequency_weight` give you the levers.
|
||||
|
||||
## 6. MCP Integration
|
||||
|
||||
Yes, in `cognee-mcp/` as a sibling project with its own pyproject. The server (`cognee-mcp/src/server.py`) uses the official `mcp` SDK over stdio/SSE/HTTP. The **publicly exposed tools are deliberately minimal** (L1076-1222):
|
||||
|
||||
- `remember(data, dataset_name?, session_id?, custom_prompt?)` — with `session_id`, fast session-cache write; without, runs full `add + cognify` pipeline.
|
||||
- `recall(query, search_type?, datasets?, session_id?, top_k=10)` — auto-routes search type.
|
||||
- `forget(dataset?, everything=False)` — deletes across all three stores.
|
||||
|
||||
There are *internal* tools registered but not exposed in API mode: `cognify`, `search`, `list_data`, `delete`, `prune`, `improve`, plus a UI bundle (`visualize_graph_ui`, `upload_file_ui`, `open_cognee_workspace`, `cognify_file`) that opens an embedded HTML workspace from `src/app_bundles/visualize-graph.html`.
|
||||
|
||||
**Notable design choice**: per-MCP-client auto-named datasets (`cursor_vscode_memory`, `claude_code_memory`) so different agents don't pollute each other's memory unless they opt in by passing `dataset_name="main_dataset"`. Toggle with `COGNEE_MCP_AGENT_SCOPED=false`.
|
||||
|
||||
MCP integration is **manual** in the sense that the agent invokes `remember`/`recall` explicitly. The Claude Code plugin (separate repo `cognee-integrations`) automates it via Claude Code lifecycle hooks (`SessionStart`, `PostToolUse`, `UserPromptSubmit`, `PreCompact`, `SessionEnd`) — see README L204.
|
||||
|
||||
## 7. Operational Concerns
|
||||
|
||||
**Deployment**: A top-level `Dockerfile` builds the FastAPI server; `cognee-mcp/Dockerfile` builds the MCP server. `docker-compose.yml` ships the API on port 8000 with **resource limits of 4 CPUs and 8 GB RAM** — non-trivial. Deployment targets include Cognee Cloud, Modal, Railway, Fly.io, Render, Daytona (`distributed/deploy/`).
|
||||
|
||||
**LLM cost per ingest**: heavy. For *every* chunk, cognify runs:
|
||||
- 1 structured-extraction call (entity/relationship → Pydantic KnowledgeGraph)
|
||||
- 1 summarization call
|
||||
- N embedding calls (one per Entity/EntityType/DocumentChunk/TextSummary/EdgeType row, plus optional Triplet rows)
|
||||
|
||||
A 100-chunk document easily generates 200+ LLM calls plus hundreds of embeddings. `chunks_per_batch` defaults to 100; `LLM_RATE_LIMIT_*` env vars exist but are off by default.
|
||||
|
||||
**What needs API keys**: `LLM_API_KEY` is mandatory unless you point at a local Ollama/HuggingFace model via the `ollama`/`huggingface` extras. Default provider is OpenAI (`openai>=1.80.1` is a core dep). Defaults to its own embedding model via `litellm`.
|
||||
|
||||
**Local-only stack**: SQLite + LanceDB + Ladybug means **no external services** for the storage layer, but the LLM is the cost driver. Replacing OpenAI with Ollama makes it CPU/GPU-heavy but free.
|
||||
|
||||
## 8. Strongest Ideas to Adopt
|
||||
|
||||
1. **Task-list pipeline composition.** `cognify.py:316-344` reads like a data-flow recipe. Each `Task` is a thin wrapper with a `batch_size` config and a context object. This is the cleanest pattern to port to Rust — represent the pipeline as a typed `Vec<Box<dyn Task>>` and let tasks declare their batching behavior.
|
||||
|
||||
2. **Triplet embeddings as first-class citizens.** Embedding `(source -› relation -› target)` text directly (`add_data_points.py:184-265`) is the trick that lets a graph become searchable by semantic similarity, not just by walking. This is genuinely the bridge between "vector RAG" and "graph RAG."
|
||||
|
||||
3. **Capability-flag-driven storage facade.** `UnifiedStoreEngine.has_capability(HYBRID_WRITE)` lets the same pipeline target separate `(Kuzu, LanceDB)` or fused `(Postgres+pgvector)` without conditionals scattered everywhere.
|
||||
|
||||
4. **`improve()` lifecycle with feedback weights.** The `feedback_weight` on graph elements + `used_graph_element_ids` from the retrieval trace is a sharp idea — it gives you graded knowledge without retraining anything.
|
||||
|
||||
5. **Multi-collection vector search at recall time.** Querying `Entity_name`, `TextSummary_text`, `DocumentChunk_text`, and edge collections in parallel and merging by triplet score, rather than picking one index, is the core retrieval move.
|
||||
|
||||
6. **Provenance stamping** (`_stamp_provenance_deep`) — every DataPoint carries `source_pipeline` + `source_task`, making the graph debuggable.
|
||||
|
||||
## 9. Realistic Downsides for a Lean Homelab
|
||||
|
||||
- **Dependency weight**: 40+ core dependencies — `openai`, `litellm`, `instructor`, `sqlalchemy`, `aiosqlite`, `lancedb`, `pylance`, `ladybug`, `networkx`, `pypdf`, `fastapi`, `fastapi-users`, `rdflib`, `langdetect`, `datamodel-code-generator`, `tiktoken`, `tenacity`, `aiolimiter`, `diskcache`, `fakeredis`, plus optional extras (neo4j, chromadb, postgres, anthropic, ollama, huggingface, scraping, dlt, monitoring). Cold install is huge — `poetry.lock` is **1.4 MB**.
|
||||
- **Python async everywhere + LRU-cached engine handles** (`closing_lru_cache`) — debugging stale-adapter issues required adding `_GraphEngineHandle` (see the 50-line docstring at `get_graph_engine.py:56-105`). That's complexity you inherit.
|
||||
- **Per-cognify LLM cost** is high. There's no obvious caching of "I've already extracted entities from this chunk hash" — re-ingestion repeats work. `incremental_loading=True` helps but only at the dataset level.
|
||||
- **Three databases to back up**, three to migrate. The "wipe `.cognee_system/` when you flip `ENABLE_BACKEND_ACCESS_CONTROL`" caveat (cognee-mcp/README L518-525) shows the multi-store coordination is fragile.
|
||||
- **8 GB RAM limit** in docker-compose is the recommended floor. LanceDB + Kuzu + embedding model loaded simultaneously is heavy on a low-end NAS.
|
||||
- **Bus factor on `ladybug`** — a 0.16.0 pinned Kuzu fork by the same org. If they stop maintaining it, you're on a single-vendor graph DB.
|
||||
- **Two-tier MCP API**: public 3 tools + 10+ "internal" tools means the API surface is in motion; not stable.
|
||||
|
||||
For a Rust homelab MCP server, the **lean equivalents** would be: SQLite (relational) + a single embedded vector store (e.g. `lancedb-rs` or embedded `sqlite-vec`) + an embedded graph DB (CozoDB, SurrealDB, or just SQL tables with proper indices) — then port the *task pipeline + triplet embedding + capability flag pattern* and skip the FastAPI/users/migrations/UI bundle stack entirely.
|
||||
|
||||
### Key file references
|
||||
|
||||
- Pipeline definition: `cognee/api/v1/cognify/cognify.py:316-344`
|
||||
- Graph extraction (LLM call site): `cognee/tasks/graph/extract_graph_from_data.py:128-222`
|
||||
- Triple-store write path: `cognee/tasks/storage/add_data_points.py:62-149`
|
||||
- Triplet embeddings: `cognee/tasks/storage/add_data_points.py:184-265`
|
||||
- Vector backend factory: `cognee/infrastructure/databases/vector/create_vector_engine.py:150-318`
|
||||
- Graph backend factory: `cognee/infrastructure/databases/graph/get_graph_engine.py:241-457`
|
||||
- Unified facade: `cognee/infrastructure/databases/unified/unified_store_engine.py:11-66`
|
||||
- Retrieval (graph+vector hybrid): `cognee/modules/retrieval/utils/brute_force_triplet_search.py:216-355`
|
||||
- Graph completion retriever: `cognee/modules/retrieval/graph_completion_retriever.py`
|
||||
- Recall + auto-router: `cognee/api/v1/recall/recall.py:314-513`, `cognee/api/v1/recall/query_router.py`
|
||||
- Improve / lifecycle: `cognee/api/v1/improve/improve.py:36-411`
|
||||
- Memify pipelines: `cognee/memify_pipelines/`
|
||||
- MCP server: `cognee-mcp/src/server.py:1076-1222`
|
||||
- Deps and extras: `pyproject.toml:22-160`
|
||||
@@ -0,0 +1,107 @@
|
||||
# Karpathy's "LLM Wiki" — Research Report
|
||||
|
||||
> The pattern this project is trying to implement faithfully. Primary source
|
||||
> below; related/competing ideas listed for honest contrast.
|
||||
|
||||
## 1. What Karpathy Actually Said
|
||||
|
||||
The canonical primary source is Karpathy's April 2026 gist [`llm-wiki.md`](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f), which he calls an **"idea file"** — explicitly *not* a library or app, but a pattern designed to be copy-pasted into an agent (Claude Code, Codex, OpenCode) so the agent can instantiate it for the user's domain.
|
||||
|
||||
The original framing came from an X thread on April 2, 2026, paraphrased as: *"using LLMs to build personal knowledge bases for various topics of research interest"*. He followed up two days later with the gist. He then [boosted "Farzapedia"](https://x.com/karpathy/status/2040572272944324650) as a good example of the pattern in the wild.
|
||||
|
||||
### The core argument (verbatim from the gist)
|
||||
|
||||
> "Most people's experience with LLMs and documents looks like RAG: you upload a collection of files, the LLM retrieves relevant chunks at query time, and generates an answer. This works, but the LLM is rediscovering knowledge from scratch on every question. There's no accumulation."
|
||||
|
||||
> "Instead of just retrieving from raw documents at query time, the LLM **incrementally builds and maintains a persistent wiki** — a structured, interlinked collection of markdown files that sits between you and the raw sources. When you add a new source, the LLM doesn't just index it for later retrieval. It reads it, extracts the key information, and integrates it into the existing wiki — updating entity pages, revising topic summaries, noting where new data contradicts old claims, strengthening or challenging the evolving synthesis. The knowledge is compiled once and then *kept current*, not re-derived on every query."
|
||||
|
||||
> "The wiki is a persistent, compounding artifact. The cross-references are already there. The contradictions have already been flagged."
|
||||
|
||||
> "The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping... LLMs don't get bored, don't forget to update a cross-reference, and can touch 15 files in one pass."
|
||||
|
||||
He explicitly links the idea to Vannevar Bush's 1945 Memex — "a personal, curated knowledge store with associative trails between documents" — arguing the part Bush couldn't solve was *who does the maintenance*; LLMs solve that.
|
||||
|
||||
## 2. The Core Principles
|
||||
|
||||
From the gist itself (Karpathy's, not paraphrased):
|
||||
|
||||
1. **Compilation, not retrieval.** Knowledge is compiled at ingest time, not re-synthesized at query time. The wiki is the artifact; raw sources are the source of truth.
|
||||
2. **Three-layer architecture.**
|
||||
- **Raw sources** — immutable; LLM reads only.
|
||||
- **Wiki** — markdown files; LLM owns and maintains entirely.
|
||||
- **Schema** (CLAUDE.md / AGENTS.md) — conventions that turn "a generic chatbot into a disciplined wiki maintainer."
|
||||
3. **Three operations: Ingest / Query / Lint.**
|
||||
- *Ingest*: one source typically touches **10–15 wiki pages**.
|
||||
- *Query*: "good answers can be filed back into the wiki as new pages... explorations compound in the knowledge base just like ingested sources do."
|
||||
- *Lint*: periodic health check for contradictions, stale claims, orphan pages, missing cross-references, data gaps.
|
||||
4. **Cross-linking is the synthesis.** The wiki is interlinked like Wikipedia or a fan wiki (he cites [Tolkien Gateway](https://tolkiengateway.net/wiki/Main_Page)); the graph *is* the consolidated knowledge.
|
||||
5. **Two navigation files: `index.md` (content catalog) and `log.md` (chronological append-only ledger).** The log uses a fixed prefix so unix tools (`grep "^## \["`) can parse it.
|
||||
6. **Division of labor.** Human curates sources and asks good questions; the LLM does "the summarizing, cross-referencing, filing, and bookkeeping." Or in his metaphor: *"Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase."*
|
||||
|
||||
### What Karpathy did *not* explicitly say (honest caveats)
|
||||
|
||||
The community frequently attributes a few additional ideas to him that are **paraphrase / extension**, not in his gist:
|
||||
|
||||
- **Episodic vs. semantic memory tiers** — neuroscience framing. Not in the gist. This comes from extensions like [LLM Wiki v2](https://gist.github.com/rohitg00/2067ab416f7bbe447c1977edaaa681e2) and the broader memory-research literature.
|
||||
- **"Sleep-like" consolidation passes** — also extension framing, not Karpathy's. His closest analog is the *Lint* operation (periodic health-check), which is rule-based rather than dream-like.
|
||||
- **Confidence scoring, Ebbinghaus decay, supersession semantics** — these are LLM Wiki v2's additions, not the original.
|
||||
- **The numbered "1. Explicit. 2. ..." Farzapedia points** — community summaries say Karpathy listed several advantages of *explicit memory artifacts* over the "AI that allegedly gets better the more you use it" status quo. Flagged as **community paraphrase**.
|
||||
|
||||
## 3. Implementation Hints from the Gist
|
||||
|
||||
- **Trigger**: human-in-the-loop on ingest ("I prefer to ingest sources one at a time and stay involved"), though batch ingestion is allowed.
|
||||
- **Format**: plain markdown in a git repo. Optional YAML frontmatter for Dataview queries.
|
||||
- **Retrieval at small scale**: `index.md` is "surprisingly good... at ~100 sources, ~hundreds of pages" — no embeddings needed.
|
||||
- **Retrieval at larger scale**: shell out to a local hybrid search tool (he names [`qmd`](https://github.com/tobi/qmd), BM25 + vector + LLM re-rank, available as CLI and MCP).
|
||||
- **Tooling**: Obsidian as the viewer (graph view to spot orphan/hub pages), Web Clipper to capture sources, version-controlled in git.
|
||||
|
||||
## 4. Related / Competing Ideas
|
||||
|
||||
- **MemGPT / Letta** ([letta.com](https://www.letta.com/blog/benchmarking-ai-agent-memory)): treats the context window as virtual memory; the *agent itself* decides what to page in/out across core, recall, and archival tiers. Stronger on long-horizon episodic coherence; higher lock-in (owns the agent loop).
|
||||
- **Mem0** ([tokenmix.ai comparison](https://tokenmix.ai/blog/ai-agent-memory-mem0-vs-letta-vs-memgpt-2026)): lightweight memory layer with `extract / store / retrieve`. Extracts memories *passively* from conversations rather than letting the agent self-edit. Low lock-in.
|
||||
- **A-MEM** ([arXiv 2502.12110](https://arxiv.org/abs/2502.12110), NeurIPS 2025): explicitly Zettelkasten-inspired. Each memory is an atomic note with structured attributes, keywords, tags; new memories trigger *evolution* of existing notes' representations. This is the closest published research analog to Karpathy's wiki — atomic notes + automatic linking + revision propagation.
|
||||
- **ReadAgent** (Google DeepMind, 2024): "gist memory" — compresses long contexts into a tree of summaries with pointers back to detail. Different angle (long-document reading), but shares the "compile, don't re-retrieve" instinct.
|
||||
- **LLM Wiki v2** (Rohit Ghumare): explicit extension with confidence scores, supersession, Ebbinghaus decay, four consolidation tiers (working → episodic → semantic → procedural), event-driven hooks, audit trails. This is basically the agentmemory model.
|
||||
- **Rowboat / knowledge-graph extension** ([dailydoseofds.com](https://blog.dailydoseofds.com/p/the-next-step-after-karpathys-wiki)): argues the wiki of summaries breaks down for evolving work contexts (deadlines, commitments) and proposes a *typed-entity knowledge graph* (decisions, people, projects as nodes).
|
||||
|
||||
## 5. Design Implications for a Rust MCP Server for Coding Agents
|
||||
|
||||
Translating Karpathy faithfully — a "Karpathy-style" backend looks very different from naive vector RAG:
|
||||
|
||||
**What it is, concretely:**
|
||||
|
||||
- **Storage = markdown files in a git repo**, not opaque vector blobs. The wiki must be human-inspectable and grep-able. Embeddings can index it but never replace it.
|
||||
- **Three directories enforced by the MCP server**: `raw/` (append-only, immutable), `wiki/` (LLM-writable, structured), and a schema doc (`AGENTS.md`-style) the server injects into every session.
|
||||
- **MCP tools mirror the three operations**: `memory_ingest`, `memory_query`, `memory_lint` — plus low-level primitives (`wiki_read`, `wiki_write`, `wiki_link`, `wiki_supersede`). Not `vector_search` as the headline tool.
|
||||
- **Ingest must be a *write fan-out*, not just an insert.** A new observation should *touch ~10–15 existing pages* — updating an entity page, a concept page, a decisions log, a gotchas page. This is the single biggest deviation from vector RAG, which only ever appends.
|
||||
- **`index.md` and `log.md` as first-class files.** The log is the audit trail and the consolidation trigger source. Use the prefix convention (`## [YYYY-MM-DD] action | title`) so it's grep-able.
|
||||
- **Retrieval is hierarchical, not just nearest-neighbor.** Read `index.md` → narrow to candidate pages → read them → optionally fall back to hybrid search (BM25 + vector, RRF-fused) for novel queries. The index *is* the synthesis; the embeddings are a backstop.
|
||||
- **Consolidation is an explicit, scheduled MCP operation**, not a side-effect. `memory_consolidate` is invoked on session-end (Claude Code's stop hook, Codex's session-end), on compaction events, or on a timer. It is LLM-driven (needs a provider key); if absent, it runs no-op as agentmemory does.
|
||||
- **Cross-agent shared state.** Because the wiki is plain text, Claude Code, Codex, and OpenCode all read/write the *same* artifact. The MCP server is the gatekeeper; the markdown is the contract. No vendor lock.
|
||||
- **Coding-specific page types**: library gotchas, architectural decisions (ADR-style), failed approaches, repo conventions, environment quirks. Karpathy's example domains were personal/research; for coding agents the high-value pages are *failure modes* and *decisions*, because those are exactly what gets dropped on context compaction.
|
||||
|
||||
**What it deliberately is *not*:**
|
||||
|
||||
- Not a vector database with a chat wrapper. Vectors are a retrieval *aid* over markdown, not the source of truth.
|
||||
- Not a chronological transcript. The log exists, but it's metadata. The semantic content lives in synthesized pages.
|
||||
- Not opaque. Every memory the agent has must be openable in Obsidian, diff-able in git, and explainable in prose.
|
||||
|
||||
**Honest tension** worth resolving in design: Karpathy's gist is optimized for *human-curated research wikis* ingested one source at a time with the user watching. A coding agent ingests *continuously and unsupervised* from tool calls. This project inherits Karpathy's structure but needs the lifecycle layer (decay, supersession, confidence) that LLM Wiki v2 proposes — otherwise the wiki will fill with stale, low-signal observations from autonomous runs.
|
||||
|
||||
## Sources
|
||||
|
||||
- [Karpathy — `llm-wiki.md` gist (primary source)](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f)
|
||||
- [Karpathy — Farzapedia tweet](https://x.com/karpathy/status/2040572272944324650)
|
||||
- [Yuchen Jin's summary tweet quoting Karpathy](https://x.com/Yuchenj_UW/status/2040482771576197377)
|
||||
- [AkitaOnRails — AI Agent Memory: Karpathy LLM Wiki and agentmemory in Practice](https://akitaonrails.com/en/2026/05/18/ai-agent-memory-karpathy-llm-wiki-agentmemory/)
|
||||
- [Rohit Ghumare — LLM Wiki v2 (gist)](https://gist.github.com/rohitg00/2067ab416f7bbe447c1977edaaa681e2)
|
||||
- [A-MEM: Agentic Memory for LLM Agents (NeurIPS 2025)](https://arxiv.org/abs/2502.12110)
|
||||
- [Mem0 vs Letta vs MemGPT comparison (TokenMix, 2026)](https://tokenmix.ai/blog/ai-agent-memory-mem0-vs-letta-vs-memgpt-2026)
|
||||
- [Benchmarking AI Agent Memory: Is a Filesystem All You Need? (Letta)](https://www.letta.com/blog/benchmarking-ai-agent-memory)
|
||||
- [The Next Step After Karpathy's Wiki Idea — Avi Chawla](https://blog.dailydoseofds.com/p/the-next-step-after-karpathys-wiki)
|
||||
- [Gamgee: Why the Future of AI Memory Isn't RAG](https://gamgee.ai/blogs/karpathy-llm-wiki-memory-pattern/)
|
||||
- [Beyond RAG: How Karpathy's LLM Wiki Pattern Builds Knowledge That Compounds (Plaban Nayak, Level Up Coding)](https://levelup.gitconnected.com/beyond-rag-how-andrej-karpathys-llm-wiki-pattern-builds-knowledge-that-actually-compounds-31a08528665e)
|
||||
- [Analytics Vidhya — LLM Wiki Revolution](https://www.analyticsvidhya.com/blog/2026/04/llm-wiki-by-andrej-karpathy/)
|
||||
- [Agentpedia — Karpathy's LLM Wiki: Complete Guide to His Idea File](https://agentpedia.codes/blog/karpathy-llm-wiki-idea-file)
|
||||
- [Tolkien Gateway (the fan-wiki Karpathy cites)](https://tolkiengateway.net/wiki/Main_Page)
|
||||
- [qmd — local hybrid markdown search (referenced in the gist)](https://github.com/tobi/qmd)
|
||||
@@ -0,0 +1,4 @@
|
||||
[toolchain]
|
||||
channel = "1.95"
|
||||
components = ["rustfmt", "clippy"]
|
||||
profile = "minimal"
|
||||
Reference in New Issue
Block a user