mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
docs(readme): lead with the differentiators, move detail into linked docs
The README had grown to 1,125 lines and buried the case for the project under installation detail. Rebuilt it at a third of the size: - New "Why ai-memory" section up top - the five differentiators the September 2026 landscape research validated (cross-agent handoffs as a typed protocol, cross-machine continuity, team sharing with auth and attribution, plain-markdown source of truth, silent zero-LLM capture), written for a human reader rather than as a feature list, plus an honest-ops note with the measured ~700/s write ceiling. - New "How it works" capture->consolidate->recall->handoff flow. - Support matrix compacted to two columns; the full 30+-row per-agent matrix moved verbatim to docs/support-matrix.md. - Quick start trimmed to AUR + docker + the two claude-code wiring commands; stale "enable per_session for two sessions" advice replaced with the v1.39 works-out-of-the-box story. - Use cases, LLM providers, and security hardening moved verbatim to docs/use-cases.md, docs/llm-providers.md, docs/security.md, each linked from a short pointer section and from the Docs index. Verbatim moves only - no prose was rewritten in the extracted docs. Packaging tests (36) still pass; README keeps zero mutable-main install URLs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
This commit is contained in:
co-authored by
Claude Fable 5
parent
204483cbfa
commit
b1595623c0
@@ -333,6 +333,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
|
||||
### Docs
|
||||
|
||||
- Reworked the README around the September 2026 research pass: a humanized
|
||||
"Why ai-memory" differentiator section and "How it works" flow up top, a
|
||||
compact support matrix, and a shorter quick start. Detail sections moved
|
||||
verbatim into linked docs: `docs/support-matrix.md` (full per-agent matrix),
|
||||
`docs/use-cases.md`, `docs/llm-providers.md`, and `docs/security.md`.
|
||||
Added `docs/research-2026-landscape.md` and follow-up pointers in the four
|
||||
May 2026 research documents.
|
||||
- `docs/windows.md` gains **Scenario E**, a persistent-server story for native
|
||||
Windows. Scenarios C and D both ended at `ai-memory serve` in a foreground
|
||||
terminal, while Linux got `Restart=on-failure` from the packaged systemd
|
||||
|
||||
@@ -0,0 +1,153 @@
|
||||
# LLM Providers
|
||||
|
||||
> Provider configuration for consolidation and embeddings. Moved verbatim from the README front page.
|
||||
|
||||
ai-memory runs without an LLM: hooks still capture sessions, search uses
|
||||
FTS5 + declared entities + graph neighbors, and summaries fall back to
|
||||
rule-based output. Add an LLM provider
|
||||
when you want LLM consolidation (on PreCompact, on demand via
|
||||
`memory_consolidate`, or opt-in at session end with
|
||||
`AI_MEMORY_CONSOLIDATE_ON_SESSION_END`), richer linting, and bootstrap.
|
||||
Substantive session ends always write a rule-based summary page + handoff either
|
||||
way. A session containing only `SessionStart` / `SessionEnd` boundaries is
|
||||
closed without a page, handoff, or provider job. When that empty session had
|
||||
accepted startup context, its session-bound handoff is returned to the open
|
||||
pool for the next receiver instead of being lost.
|
||||
When the session-end opt-in is enabled, provider work is durably queued after
|
||||
those deterministic writes and handled by one bounded server worker, so hook
|
||||
drain latency does not cancel it. Failed jobs retry with backoff and survive a
|
||||
server restart. A resumed native session is ended again only after its
|
||||
observation generation advances; the persisted generation watermark makes
|
||||
duplicate SessionEnd delivery and system clock skew converge without repeated
|
||||
provider work. The end watermark and automatic handoff commit atomically, and
|
||||
an interrupted keyed replay finishes the wiki commit, queue insert, and key
|
||||
completion without duplicating that handoff. On the next SessionStart, the
|
||||
newest cwd-eligible automatic handoff wins; accepting it expires older eligible
|
||||
automatic handoffs without consuming manual or sibling-directory work. A new
|
||||
automatic handoff also expires prior open automatic handoffs from its exact
|
||||
cwd, so repeated SessionEnds cannot accumulate there before a receiver starts.
|
||||
|
||||
To keep consolidation style project-specific, write
|
||||
`_prompts/consolidation.md` in that project's wiki. Its body can express
|
||||
preferences such as "prefer Portuguese titles" or "omit routine CI noise".
|
||||
Automatic, single-page, and multi-page consolidation use the page; a manual
|
||||
`memory_consolidate` call can pass `instructions` to override it once. ai-memory
|
||||
sanitizes and caps the value at 2,000 characters, JSON-encodes it in the user
|
||||
message, and treats it as untrusted advisory data. It cannot supply facts,
|
||||
request tool use or disclosure, or override the consolidation schema and
|
||||
faithfulness rules. TTL-expired preference pages are ignored. With no active
|
||||
page or argument, no preference block is appended.
|
||||
|
||||
Recommended defaults:
|
||||
|
||||
| Provider | Default | Use when |
|
||||
|---|---|---|
|
||||
| `anthropic` | `claude-haiku-4-5` | Best default for consolidation quality and rule classification. |
|
||||
| `anthropic-oauth` | `claude-sonnet-4-6` | Use a Claude Pro/Max subscription via `claude setup-token`, no API key. |
|
||||
| `openai` | `gpt-5.4-mini` | Cheaper and faster hosted option. |
|
||||
| `openai-oauth` | `gpt-5.5` | ChatGPT Pro/Plus/Codex backend via `ai-memory auth login openai-oauth`; no Platform API key. |
|
||||
| `copilot` | `gpt-5.5` | GitHub Copilot Chat backend via `ai-memory auth login copilot` or `COPILOT_GITHUB_TOKEN`; requires a Copilot subscription. |
|
||||
| `gemini` | `gemini-3.5-flash` | Google-hosted option with a generous free tier. |
|
||||
| `openai-compat` | no default | OpenRouter, Atlas Cloud, OrcaRouter, Ollama, vLLM, LM Studio, and other compatible endpoints. |
|
||||
|
||||
`openai-oauth` stores a refresh token in `<data_dir>/auth.json` and talks to
|
||||
the ChatGPT/Codex Responses backend, not `api.openai.com`. For Docker quick
|
||||
starts, run `ai-memory auth login openai-oauth` with the wrapper so the token
|
||||
lands in the same `ai-memory-data` volume as the server.
|
||||
|
||||
`anthropic-oauth` hits the same `/v1/messages` endpoint as `anthropic` but
|
||||
authenticates with an OAuth bearer token instead of an API key. Run
|
||||
`claude setup-token` once, then set `AI_MEMORY_LLM_PROVIDER=anthropic-oauth` and
|
||||
`ANTHROPIC_OAUTH_TOKEN=<token>` (or `CLAUDE_CODE_OAUTH_TOKEN`, which `claude
|
||||
setup-token` writes automatically). No `ANTHROPIC_API_KEY` is needed. The Docker
|
||||
wrappers forward either token by name to short-lived helper commands such as
|
||||
`llm-test`; configure the long-lived server container separately as shown in the
|
||||
installation guide.
|
||||
|
||||
For both Anthropic providers, ai-memory omits `temperature` for Claude
|
||||
4.7 and later models and Claude Mythos Preview because those models reject
|
||||
non-default sampling parameters. `llm-test` sends the same representative 0.2
|
||||
value as the normal pipeline before the provider applies that compatibility
|
||||
rule.
|
||||
|
||||
**⚠️ Unofficial and against Anthropic's usage policies — use at your own risk;
|
||||
it may get your account rate-limited or banned. See
|
||||
[the warning in `docs/install.md`](docs/install.md#anthropic-via-claude-subscription-oauth).**
|
||||
|
||||
`copilot` stores a GitHub user token in the same auth file, exchanges it for a
|
||||
short-lived Copilot API token via GitHub's `/copilot_internal/v2/token`, and
|
||||
uses the Copilot Chat endpoint with `vscode-chat` integration headers. You can
|
||||
also set `COPILOT_GITHUB_TOKEN`, `GH_TOKEN`, or `GITHUB_TOKEN` on the server.
|
||||
|
||||
> [!TIP]
|
||||
> **For the OAuth/subscription backends (`anthropic-oauth`, `openai-oauth`,
|
||||
> `copilot`), pick a small, fast model** via `AI_MEMORY_LLM_MODEL` — e.g.
|
||||
> `claude-haiku-4-5` or `gpt-5-mini`. ai-memory's LLM work (consolidation,
|
||||
> lint, explore) is summarisation, not hard reasoning, so a Haiku/mini-class
|
||||
> model is plenty and is much easier on subscription rate limits. Save the
|
||||
> high-effort thinking models for your coding agent.
|
||||
|
||||
> [!TIP]
|
||||
> **OpenAI-compatible structured output is schema-constrained by default.**
|
||||
> ai-memory sends each operation's JSON Schema through
|
||||
> `response_format=json_schema`, which recent Ollama, vLLM, LM Studio, and
|
||||
> llama.cpp releases honour. It falls back to the tolerant parser when an
|
||||
> endpoint explicitly rejects that field or returns a malformed shape. Set
|
||||
> `AI_MEMORY_LLM_COMPAT_STRICT=false` only for an incompatible endpoint.
|
||||
|
||||
For small-context local models, configure both consolidation limits. The input
|
||||
target accounts for the complete rendered prompt, including bounded slot and
|
||||
current-page context plus the structured-output schema; the output limit is
|
||||
sent to the provider. Their sum must fit the model context window, with extra
|
||||
headroom because provider tokenizers differ:
|
||||
|
||||
```toml
|
||||
[consolidation]
|
||||
max_input_tokens = 6500
|
||||
max_output_tokens = 1000
|
||||
```
|
||||
|
||||
The equivalent environment variables are
|
||||
`AI_MEMORY_CONSOLIDATION__MAX_INPUT_TOKENS` and
|
||||
`AI_MEMORY_CONSOLIDATION__MAX_OUTPUT_TOKENS`. Provider failures during an
|
||||
automatic PreCompact/PostCompaction checkpoint fall back to the deterministic
|
||||
rule-based page; admission, storage, and scope errors still fail closed. The
|
||||
validated minimums are 6,000 input and 1,000 output tokens.
|
||||
|
||||
Every chat provider bounds each completion request at 300 seconds
|
||||
(`AI_MEMORY_LLM_TIMEOUT_SECS` to override, or `llm_timeout_secs = 900` in
|
||||
config.toml; the quick openai-oauth token refresh keeps the default ceiling).
|
||||
The default tolerates a local engine cold-loading a large model; slow hosted
|
||||
gateways whose long completions exceed the ceiling fail every request with
|
||||
`http: error sending request`, so raise the value there instead of watching
|
||||
consolidation exhaust its retries.
|
||||
|
||||
Reranking is optional and off by default. With an LLM provider configured,
|
||||
`AI_MEMORY_RERANKER=llm` makes project and explicit-scope `memory_query`
|
||||
calls over-fetch from the hybrid stage, fuse scopes, and make at most one LLM
|
||||
call to reorder the best candidates. This can promote a relevant page that
|
||||
RRF ranked below the requested cut, at the cost of LLM latency and usage. The
|
||||
request sends the query plus at most 30 bounded page titles and search snippets
|
||||
to the configured provider; all values are JSON-encoded and treated as
|
||||
untrusted data. A timeout, provider error, or incomplete/invalid score set
|
||||
preserves the normal order. `global=true` and supplemental global-preference
|
||||
hits keep their existing non-RRF ranking. Concurrent provider calls are capped
|
||||
at four; saturated queries keep their local ranking without waiting.
|
||||
|
||||
Embeddings are optional and separate from the LLM provider. Set
|
||||
`AI_MEMORY_EMBEDDING_PROVIDER=openai`, `voyage`, `google`/`gemini`, or
|
||||
`openai-compat` when you want vector retrieval in addition to FTS5 + entity +
|
||||
graph-neighbor retrieval. `openai-compat` targets self-hosted engines
|
||||
(Ollama, LM Studio, vLLM): it needs no API key and requires explicit
|
||||
`AI_MEMORY_EMBEDDING_BASE_URL`, `AI_MEMORY_EMBEDDING_MODEL`, and
|
||||
`AI_MEMORY_EMBEDDING_DIM`. The optional `EMBEDDING_API_KEY` credentials the
|
||||
embedder alone and is checked before `OPENAI_API_KEY` and `LLM_API_KEY`, so
|
||||
embeddings can run on a different provider than the LLM. Both the FTS-only and
|
||||
hybrid paths apply the same bounded page-authority adjustment after candidate
|
||||
generation; embeddings improve relevance recall but do not decide which source
|
||||
is canonical.
|
||||
|
||||
See [`docs/install.md#llm-provider-tiers`](docs/install.md#llm-provider-tiers)
|
||||
for env vars and Ollama/OpenRouter/Atlas Cloud/OrcaRouter examples, and
|
||||
[`docs/llm-provider-comparison.md`](docs/llm-provider-comparison.md)
|
||||
for the empirical model comparison.
|
||||
@@ -0,0 +1,113 @@
|
||||
# Security
|
||||
|
||||
> The full security model. Moved verbatim from the README front page; the README keeps a summary.
|
||||
|
||||
Loopback-only (`127.0.0.1:49374`) with no auth is the default because
|
||||
it is safe for a single-user laptop: no process outside the machine can
|
||||
reach the server.
|
||||
|
||||
Unauthenticated non-loopback HTTP now fails closed. Set
|
||||
`AI_MEMORY_AUTH_TOKEN` or bind loopback; `--allow-insecure-no-auth` is an
|
||||
intentional, dangerous exception for plain HTTP only. Authentication does not
|
||||
encrypt bearer tokens: for LAN or remote access, use the ready
|
||||
[Caddy](docker/compose.tls.caddy.yml) or
|
||||
[Cloudflare Tunnel](docker/compose.tls.cloudflared.yml) templates described in
|
||||
the [HTTPS reverse-proxy guide](docs/https-via-proxy.md).
|
||||
|
||||
Enable bearer auth when the server is exposed beyond loopback, when
|
||||
untrusted local processes share the machine, or when the data dir holds
|
||||
sensitive project history:
|
||||
|
||||
```bash
|
||||
TOKEN=$(ai-memory generate-auth-token)
|
||||
|
||||
docker run -d --name ai-memory \
|
||||
--restart unless-stopped \
|
||||
-p 0.0.0.0:49374:49374 \
|
||||
-v ai-memory-data:/data \
|
||||
-e AI_MEMORY_AUTH_TOKEN="$TOKEN" \
|
||||
-e AI_MEMORY_ALLOWED_HOSTS="<server-ip>,localhost,127.0.0.1" \
|
||||
akitaonrails/ai-memory:latest
|
||||
|
||||
ai-memory install-mcp --client claude-code --apply \
|
||||
--server-url "http://<server-ip>:49374/mcp" --auth-token "$TOKEN"
|
||||
ai-memory install-hooks --agent claude-code --apply \
|
||||
--server-url "http://<server-ip>:49374" --auth-token "$TOKEN"
|
||||
```
|
||||
|
||||
Bearer auth protects `/mcp`, `/hook`, `/handoff`, `/workstream/*`, and
|
||||
machine calls to `/admin/*` and `/api/v1/*`. Humans sign in at
|
||||
`POST /auth/login`; the console uses an `HttpOnly` session cookie plus CSRF,
|
||||
not a Bearer in `localStorage`. Custom SPA HTML at `/web` is public static.
|
||||
When human auth listens beyond loopback,
|
||||
`AI_MEMORY_AUTH__SECURE_COOKIE=true` is required and signals that a trusted
|
||||
HTTPS reverse proxy owns the browser-facing edge. It makes the session cookie
|
||||
HTTPS-only. Close or redirect direct HTTP access to that hostname.
|
||||
Non-loopback binds should also set `AI_MEMORY_ALLOWED_HOSTS` to guard against
|
||||
DNS rebinding.
|
||||
|
||||
Busy shared hook servers can also set `AI_MEMORY_HOOK_RATE_PER_SEC` (tokens per
|
||||
second per actor/session source) and optionally `AI_MEMORY_HOOK_RATE_BURST` to
|
||||
bound one runaway session without blocking unrelated hook sources. Unset or `0`
|
||||
rate leaves the limiter disabled.
|
||||
|
||||
For shared servers where each developer should authenticate their own hook
|
||||
writes, native Claude Code hooks can use a stored OIDC device token instead of
|
||||
embedding a shared static token:
|
||||
|
||||
```bash
|
||||
ai-memory auth login oidc-device \
|
||||
--issuer "https://issuer.example.com/realms/team" \
|
||||
--client-id "ai-memory-cli"
|
||||
|
||||
ai-memory install-hooks --agent claude-code --apply \
|
||||
--server-url "http://<server-ip>:49374"
|
||||
```
|
||||
|
||||
OIDC hook auth requires the native `ai-memory hook ...` command path. The Docker
|
||||
wrapper keeps shell-script hooks by default; set up OIDC from a native release
|
||||
binary or source install. Thin-client HTTP commands such as `ai-memory status`
|
||||
and `ai-memory search` also use the stored OIDC access token when no static
|
||||
`AI_MEMORY_AUTH_TOKEN` / `[auth].bearer_token` is configured; the static bearer
|
||||
still wins when present. This is for OIDC-aware gateways/bridges; native
|
||||
ai-memory server auth still accepts static root bearer / DB-user tokens, and
|
||||
`/admin/*` remains root-only unless a gateway translates accepted OIDC auth into
|
||||
upstream auth that ai-memory accepts.
|
||||
|
||||
OIDC/Keycloak session ids are login-provider sessions, not ai-memory agent
|
||||
sessions. Shared servers that rely on `[auto_scope]` session isolation still
|
||||
need explicit `workspace` + `project` / `scopes`, or a bridge that forwards the
|
||||
real lifecycle-hook session id on MCP requests.
|
||||
|
||||
**Want HTTPS?** ai-memory deliberately does not terminate TLS itself —
|
||||
the right answer is a battle-tested reverse proxy in front of it.
|
||||
[`docs/https-via-proxy.md`](docs/https-via-proxy.md) is the deployment
|
||||
guide, with copy-paste docker compose templates in
|
||||
[`docker/compose.tls.caddy.yml`](docker/compose.tls.caddy.yml) (Caddy
|
||||
with Let's Encrypt or internal CA) and
|
||||
[`docker/compose.tls.cloudflared.yml`](docker/compose.tls.cloudflared.yml)
|
||||
(Cloudflare Tunnel — no open ports). Both are recommended once you
|
||||
turn on multi-user or bind beyond loopback. The Quick Start happy
|
||||
path of single-user on loopback doesn't need TLS — that case is
|
||||
called out explicitly in the guide so you don't add ceremony where
|
||||
it doesn't earn its keep.
|
||||
|
||||
**Multi-user attribution (v0.8, optional) plus human login.** When more
|
||||
than one human shares a server, ai-memory attributes each write to a
|
||||
named user. Humans sign in with username/password; agents and CLIs use
|
||||
`Authorization: Bearer` (`AI_MEMORY_AUTH_TOKEN` for root automation, or
|
||||
an `aim_` key from `ai-memory api-key add`). Data stays single-tenant —
|
||||
there is no per-page RBAC. A
|
||||
`[auth].token_pepper` is required for DB-user authentication, but creating the
|
||||
first user row is what immediately switches every `/admin/*` endpoint to
|
||||
root-only, including status/search/read-page and user-management routes.
|
||||
`ai-memory init` generates a pepper for new installs without changing
|
||||
single-user behavior until a user is added. An SSO gateway can instead use a
|
||||
dedicated `[auth].actor_proxy_bearer_token` and trusted `X-Memory-Actor-*`
|
||||
headers; its credential is deliberately separate from the root bearer so a
|
||||
missing identity cannot become root. See
|
||||
[`docs/users.md`](docs/users.md) for the full walkthrough and the
|
||||
four-rung auth ladder.
|
||||
|
||||
See [`docs/deploy.md`](docs/deploy.md) for the full homelab pattern
|
||||
with bearer auth, host allowlisting, and TLS/reverse-proxy options.
|
||||
@@ -0,0 +1,39 @@
|
||||
# Support Matrix
|
||||
|
||||
> The complete platform and agent matrix, with integration notes.
|
||||
> Moved verbatim from the README front page, which keeps a compact view.
|
||||
|
||||
|
||||
| Area | Status | Notes |
|
||||
|---|---|---|
|
||||
| Linux | Supported | Primary Docker/server target and CI platform. Published Docker images support `linux/amd64` and `linux/arm64`. Native Arch/AUR packages include system and user systemd units. |
|
||||
| macOS | Supported | Workspace tests run in CI; tagged releases publish native `ai-memory-macos-aarch64.tar.gz` and `ai-memory-macos-x86_64.tar.gz` binaries. The native binary is the recommended path on Apple Silicon. See [`docs/macos.md`](docs/macos.md). |
|
||||
| Windows via WSL2 | Supported | Use the Linux install path inside WSL2 when the agent runs there. |
|
||||
| Native Windows | Experimental | Tagged releases publish `ai-memory-windows-x86_64.zip` with `ai-memory.exe`; Docker Desktop wrapper and source builds are also available. Local supported profiles default to host-native hook commands; Claude Code may use its Windows exec form, while other agents use native single command strings matching their hook schema. PowerShell/Git Bash scripts are compatibility fallbacks. See [`docs/windows.md`](docs/windows.md). |
|
||||
| Claude Code | Supported | MCP config + lifecycle hooks; native commands enforce capture exclusions. `install-mcp --session-aware` optionally enables per-session auto-scope isolation through a local stdio bridge. Optionally captures the assistant's final turn on `Stop` when installed with `--capture-assistant` and the server enables `capture_assistant` (double opt-in, off by default). |
|
||||
| Codex | Supported | MCP config + lifecycle hooks; native commands enforce capture exclusions. No automatic true session-end hook, so run `ai-memory finalize-session` when you need a final summary/handoff. |
|
||||
| Command Code | Supported | MCP config (`~/.commandcode/mcp.json`) + its four stable lifecycle-hook events (`~/.commandcode/settings.json`); native commands enforce capture exclusions and `SessionStart` injects handoffs. `Stop` is only a turn boundary, so use `ai-memory finalize-session --agent command-code` after the final turn. `ai-memory run command-code` adds exact v3 native-session resume and visible-event import; experimental unsandboxed Mods remain excluded. |
|
||||
| Devin CLI | Supported | MCP config + lifecycle hooks. Hooks use Devin's `PostCompaction` event, inject handoffs via `hookSpecificOutput.additionalContext`, and omit subagent events because Devin does not expose them. |
|
||||
| OpenCode | Supported | Remote MCP config + generated TypeScript plugin; generated plugin enforces capture exclusions. |
|
||||
| Cursor | Supported | MCP config + lifecycle hooks. |
|
||||
| Gemini CLI | Supported | MCP config + lifecycle hooks. |
|
||||
| Oh My Pi / OMP | Supported | Use `--client omp` / `--agent omp` (or `oh-my-pi`) for native `.omp` MCP config + TypeScript extension; generated extension enforces capture exclusions. |
|
||||
| Pi | Supported | Generated `~/.pi/agent/extensions/ai-memory-pi.ts` extension provides lifecycle capture and an HTTP MCP bridge; generated extension enforces capture exclusions. |
|
||||
| Crush | Managed-only | `ai-memory run crush` resumes its project-local session database and supplies portable context through a temporary supported global-context file; no lifecycle-hook installer is provided. |
|
||||
| Managed workstreams | Opt-in | `ai-memory run` provides transparent cross-harness continuity for Claude Code, Codex, OpenCode, Pi, Crush, Kimi Code, Command Code, both incompatible Kiro CLI engines, OMP, Grok Build CLI, and Antigravity CLI. Direct launches remain unchanged. See [`docs/managed-workstreams.md`](docs/managed-workstreams.md). |
|
||||
| Claude Desktop | MCP-only | Uses `mcp-remote`; no lifecycle hooks. |
|
||||
| OpenClaw | Supported | MCP config + native plugin lifecycle hooks; generated plugin enforces capture exclusions. |
|
||||
| Antigravity CLI | Supported | MCP config (`serverUrl`) + lifecycle hooks (`agy` alias). Only `PreInvocation` with `invocationNum = 0` maps to SessionStart; later model calls cannot consume a next-session handoff. No automatic true session-end hook, so run `ai-memory finalize-session --agent antigravity-cli` after the final turn when you need a summary, handoff, and opt-in SessionEnd consolidation. `ai-memory run antigravity` (aliases `antigravity-cli`, `agy`) adds managed workstream resume via `--conversation`; conversation text is not decoded, so the ledger for this harness comes from hook capture. |
|
||||
| Grok Build CLI | Supported | MCP config (`install-mcp --client grok` → `$GROK_HOME/config.toml`, default `~/.grok/config.toml`) + lifecycle hooks (`install-hooks --agent grok` → `$GROK_HOME/hooks/ai-memory.json`, default `~/.grok/hooks/ai-memory.json`, Grok-specific hook bundle). Capture works; no hook handoff injection — Grok ignores `SessionStart` stdout, so recover handoffs via MCP `memory_handoff_accept`. `ai-memory run grok` adds managed workstream resume with the context packet delivered natively through `--rules`. Skills root: `.grok/skills` / `$GROK_HOME/skills` (default `~/.grok/skills`). |
|
||||
| Swival CLI | MCP-only | `install-mcp --client swival --apply` merges a native HTTP entry into the project-root `.swival/mcp.json`, preserving sibling servers. Lifecycle and managed-workstream support are not claimed because Swival's callback contract does not expose a stable session identifier. |
|
||||
| Zero | Supported | `install-mcp --client zero` (native HTTP + bearer in `~/.config/zero/config.json`) + lifecycle hooks via `install-hooks --agent zero --apply` (exec-form native commands in `~/.config/zero/hooks.json`, JSON payload on stdin, no shell). Capture works incl. specialist (subagent) events; no handoff injection — Zero discards `sessionStart` stdout, so recover handoffs via MCP `memory_handoff_accept`. |
|
||||
| ZCode | MCP-only | `install-mcp --client zcode --apply` merges a native HTTP entry (strict schema: `type`/`url`/`headers` only) into the `mcp.servers` map of `~/.zcode/cli/config.json`, preserving sibling servers. Lifecycle hooks are not yet available — see #512. |
|
||||
| Kimi Code | Supported | MCP config (`url` entry in `~/.kimi-code/mcp.json`) + lifecycle hooks (`[[hooks]]` in `~/.kimi-code/config.toml`, 10 events including subagent start/stop and `PostToolUseFailure` for tool-failure capture); both paths honor `$KIMI_CODE_HOME`. Handoffs inject via `UserPromptSubmit` stdout (Kimi Code discards `SessionStart` hook stdout); `ai-memory run kimi` adds managed workstream resume. |
|
||||
| Kiro CLI | Supported | MCP config uses `install-mcp --client kiro-cli` (alias `kiro`) and Kiro's Bedrock-compatible schema flavor. `install-hooks --agent kiro-cli` merges v2 hooks into existing agent configs; the explicit `--agent kiro-cli-v3` target writes the incompatible standalone v3 registration. Both preserve unrelated entries, honor `$KIRO_HOME`, enforce capture exclusions, and inject pending handoffs at session start. Kiro has no true SessionEnd hook; use `ai-memory finalize-session --agent kiro-cli`, with `--session-id <uuid>` for concurrent sessions. `ai-memory run kiro` manages v2; add `--v3`, `--mode`, or `--agent-engine v3` for version-safe v3 resume. |
|
||||
| Pool | Hooks-only | Poolside Agent CLI (`pool`). Lifecycle-hook capture via `install-hooks --agent pool` (alias `poolside`): Pool reads project-scoped hooks from the repo-root `.poolside/settings.yaml`, so ai-memory stages the scripts and prints a ready-to-paste `hooks:` snippet rather than writing project-local files; native commands enforce capture exclusions. Five Claude-shaped events (`SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `Stop`) and no true session-end — `Stop` is a turn boundary, so run `ai-memory finalize-session --agent pool` after the final turn. `SessionStart` stdout injection is not demonstrated, so capture works but handoff injection does not — recover handoffs via MCP `memory_handoff_accept`. No first-party `install-mcp` client and no managed workstream (`ai-memory run pool`) are claimed: Pool's native session-store contract is not demonstrated (see [`docs/managed-harness-contributions.md`](docs/managed-harness-contributions.md)). Verified against Poolside CLI v1.0.16. |
|
||||
| ZCode | Hooks-only | ZCode (z.ai, `zcode`, alias `zai`). Lifecycle-hook capture via `install-hooks --agent zcode --apply`: exec-form native commands (`type: "process"`, no shell) merged into the root `hooks` block of `~/.zcode/cli/config.json` around any third-party hooks; native commands enforce capture exclusions. Six documented triggers (`SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PostToolUseFailure`, `Stop`); `PostToolUseFailure` fires instead of `PostToolUse` when a tool throws and lands on the same capture channel with the error preserved. `PermissionRequest` is deliberately not installed — its hook chain races the interactive permission client, so passive capture of that event is unreliable by design. No true session-end — `Stop` is a per-turn boundary, so run `ai-memory finalize-session --agent zcode` after the final turn. Unlike Pool and Zero, `SessionStart` stdout injection works (`hookSpecificOutput.additionalContext`, verified live against the embedded engine v0.16.5), so the prior session's handoff is delivered automatically. No first-party `install-mcp` client and no managed workstream are claimed yet. |
|
||||
| VS Code Copilot | MCP-only | `.vscode/mcp.json` for Copilot agent mode; no lifecycle hooks (Copilot does not expose them yet). |
|
||||
| Zed | MCP-only | Native remote MCP under `context_servers` in Zed's user `settings.json`; no lifecycle hooks or managed-workstream support. |
|
||||
| Hermes Agent | Community | Core hook ingestion recognizes `agent=hermes` and Hermes' documented shell-hook `tool_name` / `tool_input` payload for concrete session attribution, tool-family titles, and capture exclusions. A community-maintained [`ai-memory-hermes-plugin`](https://github.com/MrLuciano/ai-memory-hermes-plugin) is available, but no first-party installer is shipped; review its compatibility matrix, install/uninstall scripts, and secret handling before using it. Hermes ignores session-start hook stdout, so recover handoffs through MCP. |
|
||||
| LLM/auth providers | Supported | Anthropic, OpenAI, OpenAI OAuth/Codex, GitHub Copilot, Gemini, OpenCode Zen/Go, OpenAI-compatible endpoints, and generic OIDC device auth for native hooks. |
|
||||
| Embedding providers | Supported | OpenAI, Voyage, Google Gemini, and keyless OpenAI-compatible endpoints such as Ollama, LM Studio, and vLLM. |
|
||||
@@ -0,0 +1,227 @@
|
||||
# Use Cases
|
||||
|
||||
> Moved verbatim from the README front page; the scenarios and walkthroughs are unchanged.
|
||||
|
||||
- **"Quit Claude Code and continue the same work in Codex."** Use the optional
|
||||
managed launcher when you want native session resume plus the portable visible
|
||||
history, not only a summary handoff:
|
||||
|
||||
```bash
|
||||
cd /path/to/project
|
||||
ai-memory run claude
|
||||
|
||||
# Quit Claude Code, then continue the same workstream in Codex.
|
||||
ai-memory run codex --yolo
|
||||
|
||||
# Continue in Command Code, preserving its own exact native session.
|
||||
ai-memory run command-code
|
||||
|
||||
# Later, omit the name to resume the newest usable managed session here.
|
||||
ai-memory run
|
||||
|
||||
# Start a new Codex session in the same workstream, keeping portable history.
|
||||
ai-memory run --fresh codex
|
||||
|
||||
# Kiro defaults to v2; select its incompatible v3 engine explicitly once.
|
||||
ai-memory run kiro --v3
|
||||
|
||||
# List the workstreams that can be selected from this checkout.
|
||||
ai-memory workstreams
|
||||
|
||||
# Fix a name you regret; the ledger and the current selection stay put.
|
||||
ai-memory rename-workstream --from typo-nmae --to refactor-db
|
||||
|
||||
# Pick a managed workstream from any linked local checkout, then resume it.
|
||||
ai-memory resume
|
||||
|
||||
# List open cross-agent handoffs, oldest first, with the id
|
||||
# `memory_handoff_cancel` needs to clear a stale one.
|
||||
ai-memory handoffs
|
||||
```
|
||||
|
||||
- **"Pick the project instead of remembering where it lives."** Start from a
|
||||
directory containing your checkouts and choose the checkout before the
|
||||
managed harness:
|
||||
|
||||
```bash
|
||||
ai-memory show
|
||||
|
||||
# Machine-readable discovery without launching anything.
|
||||
ai-memory show --json
|
||||
```
|
||||
|
||||
Each successful `ai-memory run` saves a client-local checkout link keyed by
|
||||
the configured server plus workspace/project. `show` joins those links with
|
||||
the server's public activity and page-count metadata. A fast, bounded depth-1
|
||||
scan of the current directory also finds new checkouts carrying a project
|
||||
marker (`.git`, `Cargo.toml`, `package.json`, `go.mod`, `pyproject.toml`, and
|
||||
friends), while skipping dependency and build directories. The server never
|
||||
exposes a checkout path, so two client machines can safely use different
|
||||
local paths for the same project on a remote homeserver.
|
||||
|
||||
The list always leads with **`+ New project`**: type a name and ai-memory
|
||||
validates a portable directory name, stages the new checkout privately, pins
|
||||
its workspace and project in `.ai-memory.toml`, and installs the routing block
|
||||
and managed Agent Skills for the chosen agent. The final directory appears
|
||||
only after every setup step succeeds, then `show` launches from it.
|
||||
|
||||
The harness menu only offers agents actually installed on the host, using the
|
||||
same `PATH` lookup `run` enforces at launch.
|
||||
|
||||
`--no-scan` uses only saved links; `--workspace` filters both sources;
|
||||
`--yolo`, `--fresh`, and trailing native arguments are forwarded unchanged.
|
||||
Non-terminal use must pass `--json`; JSON mode is discovery-only and never
|
||||
launches a harness.
|
||||
|
||||
The first explicit run can offer an existing session from this exact checkout
|
||||
or start a new one. Switching harnesses starts or resumes the native session
|
||||
linked to the shared workstream, so an obsolete local session cannot replace
|
||||
newer cross-harness history. After a normal quit, the next launch waits
|
||||
briefly if the previous launcher is still finalizing; handled failures release
|
||||
the workstream immediately. If a linked native transcript was deleted,
|
||||
ai-memory detects the orphan before launch and starts fresh; `--fresh` forces
|
||||
that recovery for one harness. Managed mode currently covers Claude Code,
|
||||
Codex, OpenCode, Pi, Crush, Kimi Code, Command Code, Kiro CLI v2/v3, OMP,
|
||||
Grok Build CLI, and Antigravity CLI; direct harness launches remain unchanged. See
|
||||
[Managed cross-harness workstreams](docs/managed-workstreams.md).
|
||||
- **"Just put me back where I was."** From any directory, with no name to
|
||||
type and no list to read:
|
||||
|
||||
```bash
|
||||
ai-memory continue
|
||||
```
|
||||
|
||||
It picks the checkout whose managed launch is most recent, revalidates the
|
||||
path and its resolved scope, then continues there exactly as bare
|
||||
`ai-memory run` would. A link whose directory moved, was replaced, now
|
||||
resolves to a different project, or has a corrupt ordering timestamp is
|
||||
reported on stderr and skipped, so a resume never quietly lands in the wrong
|
||||
project. `--workspace` narrows the search; `--yolo` and `--fresh` are
|
||||
forwarded.
|
||||
- **"Let me choose which workstream to resume."** From any directory, select
|
||||
one of the recent workstreams from your valid client-local managed checkouts:
|
||||
|
||||
```bash
|
||||
ai-memory resume
|
||||
ai-memory resume --workspace work
|
||||
```
|
||||
|
||||
The picker shows the workstream name, project scope, activity, and linked
|
||||
harnesses. Use Up/Down (or `j`/`k`) to choose a workstream and Left/Right to
|
||||
cycle its launch harness. Every row starts at `auto`, preserving normal
|
||||
session discovery; the other choices are supported harnesses currently found
|
||||
in `PATH`. Enter launches the displayed combination. The checkout is
|
||||
revalidated first, while the server continues to receive fingerprints rather
|
||||
than a local path. Use `ai-memory workstreams` when you are already in a
|
||||
checkout and only want the read-only list.
|
||||
- **"Quit at 4 PM, pick up at 9 AM in a different agent."** The
|
||||
classic. SessionStart hook in the next supported hook client prepends a
|
||||
typed handoff with open questions, next steps, and a session summary. Grok
|
||||
captures lifecycle events but ignores SessionStart stdout, so ask it to call
|
||||
`memory_handoff_accept` when resuming from a handoff. Zero has the same
|
||||
no-stdout behavior and also must call `memory_handoff_accept`.
|
||||
- **"What did we decide about X six weeks ago?"** Use `memory_query X` from
|
||||
the agent for FTS5 fused with entity matches and linked-page expansion (plus
|
||||
vector similarity when an embedder is configured). For a quick terminal-only
|
||||
FTS5 lookup, use `ai-memory search X`; that admin command does not run the
|
||||
hybrid streams. Pages are
|
||||
LLM-consolidated, so the hit is a coherent decision page, not a raw
|
||||
chat log. Pass `explain: true` to see why each hit ranked where it
|
||||
did in project or explicit-scope retrieval. Cross-project
|
||||
`global: true` search uses its separate FTS-only ranker and reports
|
||||
that active stream without per-hit RRF details.
|
||||
- **"Remember this permanently."** When something is worth keeping
|
||||
beyond auto-captured session logs - a decision, a convention, a
|
||||
gotcha - tell the agent "save a permanent note that we standardised
|
||||
on Postgres for X" or "annotate this as a project rule" and it calls
|
||||
`memory_write_page` to write a durable, git-versioned wiki page. From
|
||||
a terminal it's `ai-memory write-page --path decisions/0007-db.md
|
||||
--body $'# Standardised on Postgres\n\n...' --pinned`. `--pinned`
|
||||
exempts it from the decay sweep; the H1 on the first line of
|
||||
`--body` becomes the page title (omit `--title` — it's still
|
||||
accepted, but LLM callers trip over JSON-escaping their way through
|
||||
it, see issue #67). Unlike a handoff (single-use) or an
|
||||
auto-synthesised session page (rewritten on consolidation), a
|
||||
write-page note is yours: it shows up in `memory_query`, renders in
|
||||
`/web`, and stays until you change it.
|
||||
- **"That page you found is out of date."** The agent calls
|
||||
`memory_feedback` with the page's path and a signal: `helpful` /
|
||||
`not_helpful` tune how strongly retention keeps a sweep-eligible episodic
|
||||
page (they move its salience, which scales the decay formula's time term),
|
||||
while `stale` / `wrong` floor the salience *and* make any current page
|
||||
show up as a `feedback_flagged` finding in the next `memory_lint` report.
|
||||
Feedback never deletes anything — it lowers confidence and flags for review —
|
||||
and it attaches to the version current when feedback is recorded, so a
|
||||
later rewrite clears the flag. Retrieved page text is untrusted and never
|
||||
authorizes feedback by itself.
|
||||
- **"Remember this, but only until the sprint ends."** Pass
|
||||
`expires_at` to `memory_write_page` (RFC3339 or `YYYY-MM-DD` = end of
|
||||
that day, UTC) — or put `expires_at:` in a page's frontmatter by
|
||||
hand. Past the TTL the page disappears from search/recent/briefing
|
||||
(pass `include_expired: true` to `memory_query` to still see it) and
|
||||
the next forget sweep hard-deletes the file and its rows. A TTL beats
|
||||
a pin; `memory_lint` warns about pinned+expiring combos.
|
||||
- **"This new project has months of history before ai-memory."**
|
||||
`cd /path/to/my-project && ai-memory bootstrap` collects
|
||||
`git log`, README, `docs/`, module headers, project rules and
|
||||
one-shot-summarises them into seed wiki pages. Future sessions
|
||||
build on top.
|
||||
- **"What durable lesson did that session teach?"**
|
||||
When an LLM provider is configured, ai-memory runs a background
|
||||
auto-improvement scheduler for newly completed sessions in every project. It
|
||||
records proposed wiki edits in the pending-writes audit trail, then approves
|
||||
them immediately through the normal wiki write path by default. Scheduler ticks
|
||||
are non-overlapping: if reviewing all projects takes longer than the interval,
|
||||
the next tick is delayed until the current one finishes. Scheduling and
|
||||
approval are separate: set `[auto_improve.scheduler] enabled = false` to stop
|
||||
automatic review, or set `[auto_improve] require_approval = true` to keep both
|
||||
scheduled and manual proposals pending for human review. `ai-memory
|
||||
auto-improve --session-id <uuid>` and MCP `memory_auto_improve` remain
|
||||
available for manual catch-up or targeted reruns. When its `session_id` is
|
||||
omitted, the MCP tool selects the newest completed session without a
|
||||
persisted auto-improvement run, so repeated calls advance past short
|
||||
preflight-skipped sessions; an explicit ID reruns that session. `ai-memory
|
||||
auto-improve-report --workspace <w> --project <p>` returns a read-only
|
||||
telemetry report for recent auto-improvement outcomes without staging or
|
||||
creating proposals; add `--stage` to create one pending report page for
|
||||
audit/approval. On deployments that distinguish operators, pending learning
|
||||
proposals are isolated by qualified operator identity, so one person's
|
||||
proposal for a page does not block another's; unattributed and single-user
|
||||
deployments retain the shared pending queue. See
|
||||
[`docs/auto-improve-eval-gates.md`](docs/auto-improve-eval-gates.md) for
|
||||
example executable eval scorers.
|
||||
|
||||
Existing installs do not need per-project migration. The scheduler initializes
|
||||
a per-project first-run watermark so historical sessions are not reviewed
|
||||
automatically on upgrade, then records per-session claims so failed scheduled
|
||||
reviews do not retry forever; use manual auto-improve for old sessions or
|
||||
failed scheduled sessions you want to catch up. Older configs may still contain
|
||||
an `[auto_improve] mode = ...` line; current ai-memory ignores that legacy key,
|
||||
so you can remove it when convenient.
|
||||
- **"What housekeeping should I consider?"**
|
||||
`ai-memory curator` runs a no-LLM, rule-based maintenance report over cold
|
||||
episodic pages, stale slots, duplicate exact normalized titles, and dangling
|
||||
cross-project links. It is report-only unless `--stage` is passed; staging
|
||||
queues one report page for approval and still performs no maintenance actions
|
||||
itself. Shared servers can opt into `[decay] breadth_weight` to give pages
|
||||
reinforced by several identified operators a retention bonus; the default
|
||||
`0.0` leaves existing retention scores unchanged.
|
||||
- **"Run one ai-memory for the whole household."** Stand the server
|
||||
up on a homelab box at `0.0.0.0:49374` with a bearer token; every
|
||||
laptop/desktop talks to it. Per-cwd routing keeps each project's
|
||||
pages cleanly separated; the `/web` UI is reachable from a
|
||||
browser anywhere on the LAN.
|
||||
- **"Audit what landed before sharing with a teammate."** Browse
|
||||
the wiki at `http://<server>:49374/web` - sign in with username and
|
||||
password when human auth is on. Per-project tree view,
|
||||
rendered markdown, supersession chain visible per page.
|
||||
- **"Undo one bad page edit without rolling back the whole server."**
|
||||
`ai-memory checkpoints` shows recent wiki commits, then
|
||||
`ai-memory restore-page --path notes/foo.md --from <rev>` restores that one
|
||||
markdown file and reindexes it into SQLite. Full `backup` / `restore` is
|
||||
still the answer for DB-only state such as sessions, observations, handoffs,
|
||||
users, audit rows, and embeddings.
|
||||
- **"Drop an experiment, keep the rest."**
|
||||
`ai-memory purge-project --project experimental --confirm`.
|
||||
Atomic: that project's DB rows cascade away, its wiki subdir gets
|
||||
`rm -rf`'d, every sibling project is untouched by construction.
|
||||
Reference in New Issue
Block a user