130 KiB
Installation cookbook
The README quick-start covers the happy paths (Docker + Claude Code, Arch AUR, Fedora RPM, macOS menu bar app). This page covers everything else:
- Server on a different machine (homelab, LAN box, remote server)
- Configuring the CLI URL and auth
- Fedora native package (RPM) (systemd system service or user service)
- Arch Linux native packages (AUR) (systemd system service or user service)
- macOS menu bar app
(self-contained
.app+ LaunchAgent) - Configuring other agent CLIs (Codex, Command Code, Devin CLI, OpenCode, OMP, Pi, Cursor, Claude Desktop, Gemini CLI, Antigravity CLI, Grok Build CLI, Zero, ZCode, Kimi Code, Kiro CLI, Pool, OpenClaw, VS Code Copilot, Zed)
- Installing hooks without docker (curl-based installer)
- Running ai-memory without docker (mise, building from source)
- Managed cross-harness workstreams
(
ai-memory run, transparent native resume, and argument forwarding) - LLM provider tiers + self-hosted Ollama
- Common subcommands
- Managed routing snippets and Agent Skills
- Human password bootstrap and recovery
- Operating without auth (local-only)
- Keeping ai-memory up to date
Shorthand. Most snippets use
$TOKENandhomelab:49374. If you're following along verbatim:export TOKEN=$(docker run --rm akitaonrails/ai-memory:latest generate-auth-token)and replace
homelabwithlocalhostif the server runs on the same machine as the agent CLI.
The Docker image is published for linux/amd64 and linux/arm64; Apple
Silicon Macs and ARM64 Linux hosts should not need --platform linux/amd64.
Podman. The
bin/ai-memorywrapper automatically uses rootless Podman when Docker is not installed. It also works through thepodman-dockerdockershim; setAI_MEMORY_DOCKER=podmanto force Podman when both engines are installed. The default image name is fully qualified for non-interactive Podman short-name resolution. See SELinux-enforcing hosts for how it detects the engine's rootless and SELinux state.
Server on a different machine
When the ai-memory server runs on a LAN box (homelab, headless server) and you use Claude Code / Codex / etc. on a laptop:
Server side (the homelab host)
docker run -d --name ai-memory \
--restart unless-stopped \
-p 0.0.0.0:49374:49374 \
-v ai-memory-data:/data \
-e AI_MEMORY_AUTH_TOKEN="$TOKEN" \
-e AI_MEMORY_ALLOWED_HOSTS="<server-ip>,localhost,127.0.0.1" \
-e AI_MEMORY_LLM_PROVIDER=anthropic \
-e ANTHROPIC_API_KEY=sk-ant-... \
akitaonrails/ai-memory:latest
See Security in the README for why
AI_MEMORY_AUTH_TOKEN and AI_MEMORY_ALLOWED_HOSTS are both required for
normal non-loopback binds. Bearer auth does not encrypt traffic: use the ready
Caddy or
Cloudflare Tunnel templates from the
HTTPS reverse-proxy guide for LAN or remote access.
When the proxy serves /web over HTTPS, also set
AI_MEMORY_AUTH__SECURE_COOKIE=true in the server environment and close or
redirect direct HTTP access to that hostname. Do not set it for direct HTTP:
browsers then correctly withhold the session cookie.
Client side (the laptop)
export AI_MEMORY_SERVER_URL="http://<server-ip>:49374"
export AI_MEMORY_AUTH_TOKEN="$TOKEN"
ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply
--session-aware is an optional Claude Code MCP mode:
ai-memory install-mcp --client claude-code --session-aware --apply
It replaces the static HTTP MCP entry with a local ai-memory stdio bridge that
still connects to the configured remote server and bearer token, while
forwarding Claude's lifecycle session id. Pair it with
[auto_scope] mode = "per_session" when the same operator runs concurrent
Claude Code sessions in different projects. The default static HTTP
registration remains appropriate for one active project at a time.
If CLAUDE_CONFIG_DIR is set, the claude-code installers match Claude Code's
own config resolution: install-mcp writes the MCP registration to
$CLAUDE_CONFIG_DIR/.claude.json (instead of ~/.claude.json),
install-hooks / setup-agent target $CLAUDE_CONFIG_DIR/settings.json
(instead of ~/.claude/settings.json), and install-skills --scope global
uses $CLAUDE_CONFIG_DIR/skills (instead of ~/.claude/skills). uninstall
sweeps the active relocated paths alongside the home defaults. It cannot
discover an older arbitrary CLAUDE_CONFIG_DIR that is no longer set. The
Docker wrapper forwards the variable for config roots under its existing
$HOME bind mount; use the native binary when the relocated root is outside
$HOME.
The CLI commands (bootstrap, status, search, lint, auto-improve,
curator, pending-writes, etc.) inherit the two env vars automatically. So do
install-mcp, install-hooks, and
setup-agent: with AI_MEMORY_SERVER_URL set, install-mcp derives the
/mcp endpoint and install-hooks uses the bare server origin.
After upgrading ai-memory, refresh the managed routing package in existing
projects so Claude Code/OpenCode/Codex/Gemini pick up new tool guidance and
proactive retrieval rules. From an agent, ask "refresh the ai-memory routing in
this project"; from the terminal, run ai-memory install-instructions (or pass
--target AGENTS.md for non-Claude prompt files). The update is idempotent:
legacy long snippets between <!-- ai-memory:start --> /
<!-- ai-memory:end --> are replaced in place with the slim snippet, and
managed Agent Skills are installed or updated alongside it.
If you install into AGENTS.md and the project is also used from Claude Code,
make CLAUDE.md import it with a bare @AGENTS.md first line. Claude Code
loads CLAUDE.md and does not read AGENTS.md, so without that import the
installed block is absent from context at session start. See
Claude Code memory.
Configuring the CLI URL and auth
The ai-memory binary is a thin HTTP client. It never opens the wiki
or SQLite directly; state-touching commands go through the running
server, which is the sole writer.
Configuration is two optional environment variables:
| Variable | Default | When to set it |
|---|---|---|
AI_MEMORY_SERVER_URL |
http://127.0.0.1:49374 |
When the server runs somewhere other than the same machine, such as http://192.168.0.90:49374. |
AI_MEMORY_AUTH_TOKEN |
unset | When the server has bearer auth enabled. |
For a single-laptop loopback server, set neither variable. For a remote or homelab server, put both in your shell rc or direnv file:
export AI_MEMORY_SERVER_URL="http://192.168.0.90:49374"
export AI_MEMORY_AUTH_TOKEN="<token>"
Explicit --server-url and --auth-token flags on install-mcp,
install-hooks, and setup-agent override the environment. That is
useful when you are generating config for a client that talks to a
different server than your default CLI target.
If you run install-mcp --apply first and later run install-hooks --apply
without env vars or flags, hooks reuse the existing ai-memory MCP entry for
that agent when possible. This keeps remote MCP config and lifecycle capture
pointed at the same server instead of falling back to loopback.
All installer --apply modes preserve symlinked configuration files: the
atomic update is written to the symlink target, including a missing final
target, while the timestamped backup stays next to the user-facing config
path. This keeps stow, chezmoi, and similar dotfile-managed installs linked.
init, serve, and generate-auth-token do not need these env vars because
they either create local files or start the server itself.
Default project resolution (--project-strategy)
By default each session files memory under basename(cwd). Because an agent
shell keeps its working directory between tool calls, a single
mkdir sub && cd sub reparents the rest of the session into a phantom project
named sub. To make every session for an install resolve its project from the
git repo root instead — collapsing subdirectories and worktrees — bake the
strategy into the hooks:
ai-memory install-hooks --apply --agent claude-code --project-strategy repo-root
--project-strategy accepts basename (the new-install default; bakes nothing)
or repo-root. Omitting it during a later --apply preserves the strategy
already baked into that agent's ai-memory hooks, including during the wrapper's
automatic post-upgrade refresh. Pass basename explicitly to remove an
existing repo-root default. This works for every agent and delivery path. A
per-repo .ai-memory.toml marker's own project_strategy / project still
take precedence — see
the marker-file reference.
Fedora (RPM)
Download the x86_64 or aarch64 RPM from the
latest GitHub release
and install it with sudo dnf install ./ai-memory-*.rpm. The package includes
the binary, hook sources, system and user systemd units, and system service
configuration. Follow the user-level service or
system service steps below to initialize and start it.
Arch Linux native packages (AUR)
Use the native packages when you want /usr/bin/ai-memory plus systemd units
instead of the Docker wrapper. The package installs the binary and hook sources
once; each user still stages their agent hook scripts into their own home dir
with install-hooks --apply.
Package choice
yay -S ai-memory-bin # prebuilt Linux x86_64/aarch64 binary, fastest install
yay -S ai-memory # builds from source, works on x86_64 and aarch64
Both packages install the same runtime layout:
| Path | Purpose |
|---|---|
/usr/bin/ai-memory |
Native CLI/server binary. |
/usr/share/ai-memory/hooks/ |
Packaged hook source bundle used by install-hooks. |
/usr/lib/systemd/system/ai-memory.service |
System-wide service unit. |
/usr/lib/systemd/user/ai-memory.service |
Per-user service unit. |
/usr/lib/sysusers.d/ai-memory.conf |
Creates the ai-memory system user. |
/usr/lib/tmpfiles.d/ai-memory.conf |
Creates /var/lib/ai-memory for the system service. |
/etc/ai-memory/config.toml |
System-service config file, tracked as a pacman backup file. |
/etc/ai-memory/env |
System-service environment/secrets file, tracked as a pacman backup file. |
The binary itself does not guess between system and user mode. The unit file chooses explicitly:
| Mode | Data dir | Config | Env/secrets | Requires sudo? |
|---|---|---|---|---|
| User service | ~/.local/share/ai-memory |
~/.config/ai-memory/config.toml |
~/.config/ai-memory/env |
No |
| System service | /var/lib/ai-memory |
/etc/ai-memory/config.toml |
/etc/ai-memory/env |
Yes |
Do not run both services on the same bind address. They can coexist on disk, but
only one can listen on 127.0.0.1:49374 unless you change bind in one config.
User-level service
Use this on a single-user workstation. It needs no sudo after package install and keeps all state in your home directory.
mkdir -p ~/.config/ai-memory ~/.local/share/ai-memory
ai-memory \
--data-dir ~/.local/share/ai-memory \
--config ~/.config/ai-memory/config.toml \
init
Edit provider/auth settings if you want LLM consolidation or bearer auth:
$EDITOR ~/.config/ai-memory/config.toml
$EDITOR ~/.config/ai-memory/env
For a loopback-only local service, bearer auth is optional. If you want one:
TOKEN=$(ai-memory generate-auth-token)
printf 'AI_MEMORY_AUTH_TOKEN=%s\n' "$TOKEN" >> ~/.config/ai-memory/env
Start and inspect the service:
systemctl --user daemon-reload
systemctl --user enable --now ai-memory.service
systemctl --user status ai-memory.service
journalctl --user -u ai-memory.service -f
If the service should keep running after you log out:
loginctl enable-linger "$USER"
Verify the HTTP server:
curl http://127.0.0.1:49374/mcp
# Expect a JSON-RPC error, which means the server is reachable.
System-level service
Use this for a shared workstation, LAN box, or homelab-style host where the server should run independently of any logged-in user.
Make sure the package-created user and state directory exist, then initialize the data layout as that service user:
sudo systemd-sysusers /usr/lib/sysusers.d/ai-memory.conf
sudo systemd-tmpfiles --create /usr/lib/tmpfiles.d/ai-memory.conf
sudo -u ai-memory ai-memory \
--data-dir /var/lib/ai-memory \
--config /etc/ai-memory/config.toml \
init
Edit system config and secrets:
sudoedit /etc/ai-memory/config.toml
sudoedit /etc/ai-memory/env
The package installs /etc/ai-memory/env as root-readable only because it may
hold API keys. Keep that file out of backups or logs that other users can read.
For LAN exposure, set a non-loopback bind and allowed hosts in
/etc/ai-memory/config.toml, and set a bearer token in /etc/ai-memory/env:
bind = "0.0.0.0:49374"
allowed_hosts = ["homelab", "192.168.0.90", "localhost", "127.0.0.1"]
TOKEN=$(ai-memory generate-auth-token)
printf 'AI_MEMORY_AUTH_TOKEN=%s\n' "$TOKEN" | sudo tee -a /etc/ai-memory/env
Start and inspect the service:
sudo systemctl daemon-reload
sudo systemctl enable --now ai-memory.service
sudo systemctl status ai-memory.service
journalctl -u ai-memory.service -f
Verify from the host:
curl -sI http://127.0.0.1:49374/handoff
# 401 Unauthorized when AI_MEMORY_AUTH_TOKEN is set.
LLM provider login with native services
You do not need a paid platform API key. ai-memory's LLM features (consolidation, lint, auto-improve) are opt-in, and when you enable them you can authenticate with a subscription you already pay for instead of a metered API key: a Claude Pro/Max plan via
anthropic-oauth(claude setup-token), a ChatGPT Plus/Pro/Codex plan viaopenai-oauth(ai-memory auth login openai-oauth), or a GitHub Copilot plan viacopilot(ai-memory auth login copilot). Seedocs/llm-providers.mdfor the full table. And you can skip an LLM entirely: the default zero-LLM path still captures, searches (FTS), and writes rule-based summaries with no provider at all —docs/local-embeddings.mdmakes vector search keyless too.
API-key providers go in the relevant env file:
# User service
printf 'AI_MEMORY_LLM_PROVIDER=anthropic\nANTHROPIC_API_KEY=sk-ant-...\n' >> ~/.config/ai-memory/env
systemctl --user restart ai-memory.service
# System service
sudoedit /etc/ai-memory/env
sudo systemctl restart ai-memory.service
OAuth-style providers write tokens into the selected data dir. Run the login
with the same --data-dir and --config pair as the service:
# User service
ai-memory \
--data-dir ~/.local/share/ai-memory \
--config ~/.config/ai-memory/config.toml \
auth login openai-oauth
# System service
sudo -u ai-memory ai-memory \
--data-dir /var/lib/ai-memory \
--config /etc/ai-memory/config.toml \
auth login openai-oauth
Use auth login copilot the same way for GitHub Copilot. For per-developer
native hook auth against an OIDC issuer, run auth login oidc-device in the
developer's selected data dir instead:
ai-memory auth login oidc-device \
--issuer "https://issuer.example.com/realms/team" \
--client-id "ai-memory-cli"
The stored OIDC access token is also used by thin-client HTTP commands
(status, search, read-page, write-page, backup, embed, and
similar) when no static AI_MEMORY_AUTH_TOKEN / [auth].bearer_token is
configured. Static bearer auth still has precedence. This is for external
OIDC-aware gateways/bridges; native ai-memory server auth still uses static root
bearer / DB-user tokens, and /admin/* remains root-only unless a gateway
translates accepted OIDC auth into upstream auth that ai-memory accepts.
OIDC/Keycloak sid claims describe the login provider's session, not the
coding-agent session ai-memory uses for [auto_scope] isolation. Gateways may
propagate the authenticated user/client/agent headers, but
X-Memory-Actor-Session-Id should only contain a real lifecycle-hook session id
from a session-aware bridge.
Restart the service after changing provider settings:
systemctl --user restart ai-memory.service # user mode
sudo systemctl restart ai-memory.service # system mode
Wire agent CLIs after native install
For a local loopback server with no bearer token:
ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply
For concurrent Claude Code sessions, set [auto_scope] mode = "per_session" in
the server config and add --session-aware to the install-mcp command. This
works for local and LAN servers; the generated stdio bridge keeps using
AI_MEMORY_SERVER_URL / --server-url and AI_MEMORY_AUTH_TOKEN.
For a bearer-protected local or LAN server, export the endpoint first. The MCP
URL includes /mcp; the hook URL is the bare origin.
export AI_MEMORY_SERVER_URL="http://127.0.0.1:49374"
export AI_MEMORY_AUTH_TOKEN="$TOKEN"
ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply
install-hooks finds packaged hook sources under /usr/share/ai-memory/hooks,
then stages runnable copies under ~/.local/share/ai-memory/hooks/<agent>/ so
the agent can execute files owned by your user. Re-run install-hooks --apply
after package upgrades to refresh those staged copies.
Hook latency expectations
Hook-capable agents launch one short-lived hook process per lifecycle event. A
successful tool call normally fires both a pre-tool and a post-tool event, so
per-invocation overhead is paid twice. As an order-of-magnitude reference, one
independent v1.29.0 evaluation on native macOS aarch64 measured about 145 ms per
posix-native invocation (about 290 ms per completed tool call) and about 172
ms per legacy .sh invocation.
Those figures are one host's measurements, not a benchmark or performance guarantee. Process startup, filesystem and security software, the selected data directory, authentication, and host load can all change the result. The native hook fast path skips full config loading and tracing, and normally spools locally instead of waiting for the server, so keep the installer's native default where supported. Latency-sensitive deployments should measure their own agent host before opting into additional captured events or choosing a script fallback.
Capture-policy capability and refresh
[capture] ignore_paths is enforced only by native ai-memory hook commands
and generated OpenCode/OMP/Pi/OpenClaw integrations. Local installers select
native commands where supported; legacy .sh/.ps1 hooks and remote-only or
Docker script bundles do not enforce it. Re-run install-hooks --agent <agent> --apply or refresh/reinstall generated plugins after upgrading; installer
capability output reflects the selected integration. See the canonical
capture exclusions reference.
Lifecycle observation bodies are bounded separately from the 10 MiB HTTP
request limit. User prompts and post-compaction summaries retain up to 16 KiB;
notifications and tool excerpts retain up to 2 KB. Native ai-memory hook
commands truncate those fields UTF-8-safely before they enter the local spool
or wire. The server repeats the event-specific caps for every integration,
including script and generated clients, then applies a 16 KiB backstop after
sanitization before any observation reaches SQLite or FTS. Native hook commands
invoke the installed binary directly, so upgrading that binary is enough to
receive the client-side cap.
Disable Claude Code prompt capture. Regulated or mixed-trust machines can
leave the rest of Claude Code's lifecycle capture enabled while preventing
UserPromptSubmit text from entering the local spool or wire:
ai-memory install-hooks --agent claude-code --no-capture-prompts --apply
The installer removes only ai-memory's prompt hook and preserves third-party
hooks registered under the same event. A later bare install-hooks --apply
(including an upgrade refresh) inherits the disabled state. Re-enable prompt
capture explicitly with --capture-prompts. These options are Claude Code-only:
some other agents use their prompt hook to inject handoff context, so removing
it would break continuity. Disabling prompt capture reduces session summaries
and recall quality because user intent is no longer part of the observation
stream; tool and session-boundary capture continues unchanged.
Capture only repositories that opt in. The controls above narrow what is
captured; this one narrows where. By default a repository with no
.ai-memory.toml marker is still captured, so a machine that works across many
checkouts captures every new one automatically — forgetting a marker means
capturing more, not less. Allowlist mode inverts that:
ai-memory install-hooks --apply --capture-mode allowlist
A repository without a marker then emits no lifecycle event at all — not a
trimmed one. The event is dropped in the hook process before it can reach the
local spool or the wire, so nothing is written to disk for a repository that
never opted in. Opting a repository in is just placing a .ai-memory.toml
marker in it, which is the same file that already configures routing and
ignore_paths.
It is enforced by native ai-memory hook commands only — the same
boundary that already applies to [capture] ignore_paths, and for the same
reason: the gate runs inside the hook binary, immediately before it spools.
That is what a normal install-hooks --apply writes on Linux, macOS and
Windows, so the usual install is covered. It is the script installs that are
not: the bundled shell/PowerShell hooks POST to the server directly and never
execute the binary, so nothing reads the mode. In practice that means the
legacy posix/windows platform override (AI_MEMORY_HOOK_PLATFORM), the
Docker host wrapper, and setup-agent snippets, which emit script commands by
design. install-hooks --apply prints the mode and warns when the install it
is writing cannot enforce it.
Within that boundary the mode is not per-agent: it is stored once in the data
directory and every native hook command reads it, whichever agent invoked it.
That also means a later bare
install-hooks --apply — including the auto-refresh inside ai-memory upgrade
— leaves it alone by construction rather than by re-detecting it. Every
--apply prints the mode in force. Return to the default with
--capture-mode denylist.
Verify it on any repository without changing anything:
printf '{"cwd":"%s"}' "$PWD" | ai-memory hook --event user-prompt-submit \
--agent claude-code --server-url "$AI_MEMORY_SERVER_URL" --check-capture
--check-capture inspects policy without spooling, draining, or contacting the
server, and needs a JSON payload naming the directory to test. In the output,
"admits_capture": false means that repository captures nothing;
"marker_present" shows whether a marker was found by the upward walk.
Note the trade: recall is lost for every repository you do not mark, and a repository you intended to capture stays silent until you add its marker.
Some agent harnesses attach the assistant's final turn to their Stop event —
Claude Code sends it as a raw last_assistant_message. By default that text is
never persisted: the native hook binary strips the raw field before it can reach
the local spool or the wire, and the server strips it defensively on arrival.
Opt-in capture (#196). You can opt in to storing a sanitized, 2 KB-capped excerpt of the assistant's final turn as the Stop body. It is a double opt-in — enable the server first, then the client:
-
Server: set
capture_assistant = truein the live<data_dir>/config.toml(or the service's configured TOML file), or setAI_MEMORY_CAPTURE_ASSISTANT=true, then restartai-memory serve. -
Client: re-install the Claude Code (or Codex) hooks with the flag:
ai-memory install-hooks --agent claude-code --capture-assistant --apply # Codex is supported too — its Stop payload carries last_assistant_message: ai-memory install-hooks --agent codex --capture-assistant --apply
The client sanitizes (built-in patterns) and truncates the excerpt before it
touches the spool or wire; the server re-scrubs with its [sanitize] patterns
before storing. If either side is off — or the marker is malformed — the Stop
stays empty. Re-running install-hooks without --capture-assistant removes
the flag (idempotent). --capture-assistant is Claude Code, Codex and
OpenCode 2 (--agent opencode2) on a native hook platform only; on any other agent or the script fallback the
installer refuses it rather than enabling something that cannot take effect. Assistant text is
privacy-sensitive — read the SECURITY.md notes on what it can contain and where
it flows (consolidation/reviewer prompts, and out to a cloud LLM provider if one
is configured) before enabling it.
Upgrading the binary is sufficient for native Claude Code installs, and pending
spooled events drain with the raw field stripped as well. Installs that run the
.sh/.ps1 script fallback (the Docker script bundle or an explicit
AI_MEMORY_HOOK_PLATFORM=posix) cannot sanitize the assistant text, so a Stop
payload still carrying the raw field is dropped whole by the script rather than
POSTed verbatim. The Docker wrapper deliberately keeps script commands because a
binary path inside its helper container is not valid on the host; running
install-hooks through that wrapper refreshes the scripts but does not convert
them. To capture assistant text safely, install a native ai-memory client on the
agent host, then use that native executable to run
install-hooks --agent claude-code --apply. Even if the script fallback is
retained, the server still strips any raw field on receipt before persistence.
Native ai-memory hook --event ... commands spool events locally. The POSIX
shell bundle spools too, but only on failure: it POSTs first and writes the
event to the same <data_dir>/hook-spool/ contract when the server is
unreachable or answers 5xx, then flushes the backlog behind the next delivery
that succeeds. A 4xx is a permanent rejection and is not retried. (The
PowerShell bundle still drops an undelivered event.) Session start
does a short bounded cleanup drain before fetching a handoff; cancellation-prone
boundary events (stop, pre-compact, and session-end) start a detached
hook-drain helper so delivery does not depend on one shutdown hook surviving.
The POSIX bundle assigns one idempotency key before its initial POST and keeps
that key if the event enters the spool. A server that processed an event but
lost the response will not duplicate its observation or completed session-end
effects; if processing stopped after the observation commit, the retry re-runs
downstream work. SessionEnd atomically
commits its end watermark with its automatic handoff; a retry that finds that
transaction complete finishes any interrupted wiki commit, durable provider
enqueue, and ingest-key completion without adding a second handoff. Those
incomplete effects remain at-least-once until the server marks the event
complete.
On Unix, the helper uses a trusted setsid launcher when available and falls
back to a separate process group otherwise; Windows uses detached/breakaway
process flags. The spool is capped, so a permanently undrained backlog is
eventually pruned rather than unbounded, but old undelivered events can be lost.
The built-in timings stay short on agent-facing paths, but high-latency or
large-backlog instances can raise them with whole-minute runtime env vars in the
agent's environment; no install-hooks rerun is needed:
| Env var | Built-in default | Max override | What it caps |
|---|---|---|---|
AI_MEMORY_HOOK_DRAIN_TIMEOUT_MINUTES |
3 seconds | 60 minutes | each event POST during a drain |
AI_MEMORY_HOOK_HANDOFF_TIMEOUT_MINUTES |
3 seconds | 60 minutes | the synchronous session-start handoff GET |
AI_MEMORY_HOOK_START_BUDGET_MINUTES |
3 seconds | 60 minutes | total time session-start may spend waiting for the drain lock and cleanup draining |
AI_MEMORY_HOOK_BACKGROUND_DRAIN_BUDGET_MINUTES |
5 minutes | 60 minutes | total time the detached hook-drain helper may spend after a background-drain boundary |
AI_MEMORY_HOOK_INCREMENTAL_THRESHOLD |
32 events | positive integer | spool backlog size that triggers a 250 ms post-tool-use catch-up drain |
Timing values must be positive whole minutes. Missing, empty, non-numeric, or zero values fall back to the built-in defaults; values above 60 are clamped. The incremental threshold is a positive event count; invalid values fall back to 32.
Server-side hook ingest also has an optional per-source limiter for shared or
remote installs that need protection from one runaway agent session. Set
AI_MEMORY_HOOK_RATE_PER_SEC on the server to the token refill rate per
actor/session source; 0 or unset disables the limiter. Set
AI_MEMORY_HOOK_RATE_BURST to override the burst size (defaults to the refill
rate, minimum one token when enabled). The limiter is bounded in both key count
and key bytes, and /hook/batch drains can skip over-budget sources while still
accepting later unrelated sources.
Native service operations
# User service
systemctl --user restart ai-memory.service
systemctl --user stop ai-memory.service
journalctl --user -u ai-memory.service -n 100
# System service
sudo systemctl restart ai-memory.service
sudo systemctl stop ai-memory.service
journalctl -u ai-memory.service -n 100
Backups still use the same CLI, just point it at the service data dir:
# User service
ai-memory --data-dir ~/.local/share/ai-memory backup --to ~/ai-memory-backup.tar.gz
# System service
sudo -u ai-memory ai-memory --data-dir /var/lib/ai-memory backup --to /var/lib/ai-memory/backup.tar.gz
Package removal does not delete data. Stop the service and remove state only when you intentionally want to erase memory:
systemctl --user disable --now ai-memory.service
sudo systemctl disable --now ai-memory.service
# Optional destructive cleanup:
rm -rf ~/.local/share/ai-memory ~/.config/ai-memory
sudo rm -rf /var/lib/ai-memory /etc/ai-memory
Maintainer integration test
The normal CI runs scripts/check-native-packaging.sh, a host-safe regression
check that uses a temporary alternate root for systemd-analyze,
systemd-sysusers, and systemd-tmpfiles. It verifies unit syntax, expected
paths, sysusers output, tmpfiles rules, env-file mode, and AUR shell syntax
without writing to host /usr, /etc, /var, or touching real services.
The repo also includes a manual Arch integration harness that is intentionally kept out of routine CI because it creates a disposable distrobox, installs packages, starts real systemd services, and can take several minutes:
scripts/test-native-arch-systemd-distrobox.sh
It verifies the AUR metadata shape, builds the current working tree, installs
the native layout into the disposable Arch container, starts the system service
with systemctl, starts the user-profile command under transient systemd
supervision, and checks that packaged hook sources under
/usr/share/ai-memory/hooks can be staged by install-hooks.
The destructive part of that script refuses to run unless it detects a container/distrobox environment.
Useful knobs:
AI_MEMORY_NATIVE_TEST_BOX=ai-memory-native-test scripts/test-native-arch-systemd-distrobox.sh
AI_MEMORY_NATIVE_TEST_KEEP_BOX=1 scripts/test-native-arch-systemd-distrobox.sh
AI_MEMORY_NATIVE_TEST_IMAGE=quay.io/toolbx/arch-toolbox:latest scripts/test-native-arch-systemd-distrobox.sh
macOS menu bar app
On a Mac, the self-contained menu bar app is the GUI install: it bundles the
native ai-memory binary and hooks/ tree, governs the existing LaunchAgent
(com.github.akitaonrails.ai-memory), and opens /web, ai-memory status,
config.toml, the data directory, and logs. It does not replace those tools
with a second dashboard.
Needs a Rust toolchain and Xcode / Swift 6 (the same as a source build):
git clone https://github.com/akitaonrails/ai-memory
cd ai-memory
./companions/ai-memory-macos/build.sh
open "companions/ai-memory-macos/dist/AI Memory.app"
Drag AI Memory.app to /Applications, then Install & Start Server
from the menu extra (no Dock icon). When the status item is green, wire an
agent with the bundled binary so install-hooks finds the sibling hooks/
tree:
BIN="/Applications/AI Memory.app/Contents/Resources/runtime/ai-memory"
"$BIN" install-mcp --client claude-code --apply
"$BIN" install-hooks --agent claude-code --apply
Durable memory stays in ~/Library/Application Support/ai-memory. Replacing
the .app is an update and does not rewrite that tree. Prebuilt tarball,
source-build, Docker-wrapper, and hand-installed launchd paths remain in
docs/macos.md. Companion source:
companions/ai-memory-macos.
Configuring other agent CLIs
install-mcp --server-urlaccepts either the bare server origin or the full MCP endpoint and appends a missing/mcpexactly once.install-hooks --server-urltakes the bare server origin (e.g.http://homelab:49374) — hook scripts append/hook,/handoff, etc. themselves.
Each agent CLI needs two things:
- MCP registration - so the agent can call
memory_query,memory_recent,memory_handoff_accept. - Lifecycle hooks - so the server auto-captures session events. Without this, the agent can still query memory but capture becomes manual.
Claude Desktop, VS Code Copilot, and Zed are MCP-only today. The
hook-capable clients in the README Support Matrix,
including Pi and Zero, have lifecycle capture paths through install-hooks.
Hook install pattern. Local supported profiles default to host-native commands. Claude Code may use its supported Windows exec form (
command= realai-memory.exe,args= argv tokens forhook --event ...); other agents use native single command strings according to their hook schema. PowerShell/Git Bash script bundles are compatibility fallbacks and do not enforce capture-policy v1. Remote-only/Docker script installs still use the two-step path: (1)docker cpbundled scripts to your home dir, (2)docker run --rm install-hooksrenders the config snippet. OpenClaw, OpenCode, OMP, and Pi are different: they use generated TypeScript plugin/extension files, so no shell-script extraction is needed for those clients.
OpenAI Codex
# MCP snippet (merge into $CODEX_HOME/config.toml, default ~/.codex/config.toml):
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client codex \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
# Hooks — extract scripts + render config:
docker cp ai-memory:/usr/local/share/ai-memory/hooks ~/.ai-memory/
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent codex \
--hooks-dir ~/.ai-memory/hooks \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
Native Codex tool hooks use top-level tool_name, tool_input, tool_response,
and tool_use_id fields (verified against CLI 0.154.0). ai-memory records the
tool family and call ID on PreToolUse and PostToolUse; recognized tools such
as Bash and apply_patch also retain a sanitized response excerpt on
PostToolUse, capped at 2 KB including metadata. Structured JSON responses are
flattened using the same bounded excerpt path. Inputs are not copied into
observations, and unknown tools (including unrecognized MCP names) retain only
metadata. PostToolUse alone does not prove success, so Codex outcomes remain
unknown.
Capture exclusions still run before native spooling. Codex's apply_patch
passes patch text in tool_input.command, which does not provide direct file
paths to the capture policy. With active ignore_paths, those events retain
only metadata; ai-memory does not parse patch or shell text to infer paths.
Tool events are delivered at the normal 32-event catch-up threshold or a
lifecycle drain boundary, so a small active turn may still have queued events.
Codex CLI 0.145.0 and later expose a native SessionEnd hook. Stop is captured
as a turn boundary and leaves the session open; a native SessionEnd triggers
the final summary, handoff, and auto-improvement eligibility. See the
Codex hook lifecycle for when
Codex ends a session. For older clients or a missed session-end delivery, run:
ai-memory finalize-session
# add --all to close every matching open Codex session in this workspace/project
# or target one exact concurrent session (mutually exclusive with --all)
ai-memory finalize-session --session-id <uuid>
Antigravity CLI also lacks a true session-end event. Its Stop hook marks the
end of one execution loop, so ai-memory intentionally records it without
closing the conversation. Its PreInvocation hook likewise runs before every
model call; ai-memory treats only the documented invocationNum = 0 call as
SessionStart. Later invocations return an empty hook result without capturing
another start or fetching the single-use handoff, so a handoff created while
the current conversation winds down remains available to the next session.
After the final turn, finalize the latest matching Antigravity session
explicitly:
ai-memory finalize-session --agent antigravity-cli
# add --all only to close every matching open Antigravity session in this scope
# or add --session-id <uuid> to close one exact concurrent session
# if the conversation continued after a first finalize, re-close it to cover
# the new observations (a re-run with nothing new is a harmless no-op)
ai-memory finalize-session --agent antigravity-cli --reopen --session-id <uuid>
Devin CLI
Devin uses ~/.devin/config.json for MCP servers and ~/.devin/hooks.v1.json
for lifecycle hooks by default. If you prefer one combined Devin config file,
pass --config-file ~/.devin/config.json to install-hooks; ai-memory then
merges the hook entries under that file's hooks key.
ai-memory install-mcp --client devin --apply \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
ai-memory install-hooks --agent devin --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
ai-memory install-skills --agent devin
Devin's hook vocabulary is close to Claude Code's, with two important differences:
- Devin emits
PostCompactionafter compaction and includes asummaryfield; ai-memory records it aspost-compaction. - Devin does not expose subagent start/stop hooks, so ai-memory cannot capture nested subagent boundaries for Devin.
The SessionStart hook injects pending handoffs through Devin's
hookSpecificOutput.additionalContext. Real Devin SessionStart and
PostToolUse payloads may omit session_id and cwd; ai-memory now infers cwd
from DEVIN_PROJECT_DIR or the hook process working directory when the payload
omits it, and mints/reuses a per-host session id from hook state when necessary,
so those events are still captured. A payload-provided value always wins.
Kimi Code
Kimi Code keeps MCP servers in ~/.kimi-code/mcp.json and lifecycle hooks in
~/.kimi-code/config.toml; both move together when $KIMI_CODE_HOME is set.
The CLI also accepts --agent kimi as an alias. install-mcp writes the
server URL with a ?flavor=moonshot query because the Moonshot API rejects
root-level anyOf/oneOf/allOf in tool parameter schemas ("moonshot
flavored json schema") — the ai-memory server answers flavored requests with
flat schemas, and all other clients keep the upstream shape.
ai-memory install-mcp --client kimi-code --apply \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
ai-memory install-hooks --agent kimi-code --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
install-hooks merges [[hooks]] entries into config.toml, preserving the
provider/model settings the same file holds. Entries cover 10 events —
SessionStart, SessionEnd, UserPromptSubmit, PreToolUse, PostToolUse,
PostToolUseFailure (Kimi Code reports tool failures separately from
successful calls; it reuses the post-tool-use handler), Stop,
SubagentStart, SubagentStop, and PreCompact — and default to
native ai-memory hook --event … --agent kimi-code commands on local installs
(local spool plus batched delivery, capture-policy v1 enforced); the staged
script bundle under ~/.local/share/ai-memory/hooks/kimi-code/ is the
compatibility fallback (POSTs to /hook, spooling a failed delivery for a
later drain, without capture-policy v1 enforcement). A pending handoff
is injected at UserPromptSubmit through the hook's stdout, which Kimi Code
appends to the model context as a user message before the turn; Kimi Code
fires SessionStart but discards that hook's stdout, so hooks installed by
an older release consumed handoffs without delivering them. Existing native
hook commands invoke the current ai-memory binary and pick up the corrected
delivery behavior on upgrade. Re-run
ai-memory install-hooks --agent kimi-code --apply only for a
script-fallback installation so its staged scripts are refreshed.
Kimi Code hook entries accept only event, matcher, command, and
timeout; extra fields make the whole config.toml fail to load, so prefer
install-hooks --apply over hand edits.
Command Code
Command Code keeps user-scope MCP and hook configuration in separate JSON
files under ~/.commandcode/. Install both integrations with:
ai-memory install-mcp --client command-code --apply \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
ai-memory install-hooks --agent command-code --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
The aliases commandcode, cmdc, and cmd are accepted. install-mcp
merges a native HTTP entry into ~/.commandcode/mcp.json; install-hooks
merges only Command Code's four stable events (SessionStart, PreToolUse,
PostToolUse, and Stop) into ~/.commandcode/settings.json, preserving
other settings and hook handlers. The hook definitions deliberately omit
matcher: Command Code documents omission as "all tools", while any matcher
on SessionStart or Stop prevents that lifecycle hook from firing.
Local installs use the native ai-memory hook command, so Command Code's
native session_id and cwd are attributed directly;
recognized shell_command, read_file, write_file, and edit_file
payloads pass through the same bounded capture-exclusion policy as other
native integrations. A pending handoff is injected through
hookSpecificOutput.additionalContext at SessionStart.
Command Code's stable Stop event ends a turn, not a session. Finalize the
open session after the last turn when you need immediate consolidation and a
handoff:
ai-memory finalize-session --agent command-code
ai-memory finalize-session --agent command-code --session-id <uuid>
ai-memory does not install Command Code Mods. Mods run arbitrary unsandboxed code and are not needed for the stable hook or managed-session paths.
Managed sessions are opt-in:
ai-memory run command-code
ai-memory run command-code --yolo --model <model-id>
The aliases commandcode, cmdc, and cmd select the same adapter. The
default executable is command-code on Unix and cmdc on native Windows.
Fresh sessions keep Command Code's native UUID; returning sessions use exact
--session <uuid>. The read-only adapter accepts only the observed v3 header,
requires its UUID filename and canonical cwd to match the checkout, excludes
checkpoint/prompt sidecars, hidden reasoning, images, and provider metadata,
and preserves parentId on visible events for branch provenance. An unknown
future transcript version fails closed until its schema is audited. Direct
cmd, cmdc, or command-code launches remain unchanged.
Kiro CLI
Kiro CLI has one MCP surface and two incompatible lifecycle-hook formats.
ai-memory supports both through explicit installer targets: kiro-cli remains
the v2 target, while kiro-cli-v3 selects the standalone v3 registration. The
global MCP file is $KIRO_HOME/settings/mcp.json, defaulting to
~/.kiro/settings/mcp.json; pass --config-file .kiro/settings/mcp.json for a
project-scoped entry.
ai-memory install-mcp --client kiro-cli --apply \
--server-url "https://memory.example/mcp" \
--auth-token "$TOKEN"
The kiro alias is equivalent. The installed URL includes
?flavor=bedrock so Kiro's Bedrock backend receives schemas without
root-level anyOf, oneOf, or allOf; nested schemas and runtime validation
remain intact. Kiro requires HTTPS for non-loopback remote servers, so the CLI
rejects a plain-HTTP homelab URL before changing the file. Configure a reverse
proxy as described in HTTPS via reverse proxy.
Install v2 hooks with the kiro-cli agent value. When --server-url is omitted,
install-hooks can infer the hook origin and bearer token from the managed MCP
entry above.
# Default v2 engine: merge hooks into every existing global agent config.
ai-memory install-hooks --agent kiro-cli --apply
# A project-local v2 agent overrides a same-named global agent. Update the
# selected local config explicitly instead of assuming the global copy runs.
ai-memory install-hooks --agent kiro-cli --apply \
--config-file .kiro/agents/<agent-name>.json
The v2 engine stores camelCase hooks inside agent JSON files. ai-memory updates
existing $KIRO_HOME/agents/*.json files only; it will not fabricate an agent
that Kiro never selects. Create and select an agent first when that directory
is empty. Kiro gives project-local agents precedence over global agents,
so use --config-file when the active definition lives under
.kiro/agents/. All target files are parsed before any one is changed, and
unrelated agent fields, third-party hooks, and each agent's existing
--project-strategy remain intact.
The install registers spawn, user-prompt, pre-tool, post-tool, and stop capture,
remains fail-open when ai-memory is unavailable, and delivers a pending handoff
through successful agentSpawn stdout. Verified v2 tool payloads enforce
[capture] ignore_paths; an unrecognized payload shape is stored as bounded
metadata rather than exposing file content.
Install v3 hooks with the explicit kiro-cli-v3 target. This distinction is
intentional: kiro and kiro-cli continue to mean v2 so an upgrade cannot
silently rewrite an existing installation into an incompatible format. The
standalone registration was acceptance-tested with an interactive Kiro CLI
2.16.2 --v3 session.
# Global v3 registration under $KIRO_HOME/hooks (default ~/.kiro/hooks).
ai-memory install-hooks --agent kiro-cli-v3 --apply
# Project-local v3 registration.
ai-memory install-hooks --agent kiro-cli-v3 --apply \
--config-file .kiro/hooks/ai-memory.json
The v3 installer writes the documented standalone version: "v1" schema with
PascalCase triggers. It preserves third-party entries in a shared file,
refuses an unsupported schema version or a third-party collision with an
ai-memory-reserved hook name, and bounds capture-only commands to one second.
SessionStart gets five seconds so ai-memory's bounded handoff fetch can finish.
Both engines use the same sanitized hook-ingress boundary: documented and live
tool_name/tool_input file operations honor [capture] ignore_paths, while
unknown file-tool payload shapes degrade to metadata-only capture.
Kiro v2's stop event ends a turn, not the session. After the final turn, close
the matching session explicitly; use the exact id when several Kiro sessions
are open in the same project:
ai-memory finalize-session --agent kiro-cli
ai-memory finalize-session --agent kiro-cli --session-id <uuid>
ai-memory uninstall --only hooks --apply --yes removes only exact ai-memory
entries from global v2 agents, the current project's .kiro/agents directory,
and ai-memory's global/current-project v3 registration. A purely generated v3
file is deleted; third-party entries in a shared file remain. ai-memory run kiro (alias kiro-cli) manages the default v2 engine and honors $KIRO_HOME;
add --v3, --mode, or --agent-engine v3 for version-safe v3 resume. Once
linked, a later plain Kiro launch recovers the stored engine transparently, and
bare ai-memory run considers checkout-local sessions from both incompatible
stores. See
managed workstreams.
Pool (Poolside Agent CLI)
Pool reads lifecycle hooks from a project-scoped .poolside/settings.yaml at
the root of each repository it runs in — there is no user-global hook file for
ai-memory to merge. install-hooks --agent pool (alias poolside) therefore
stages the hook scripts to the stable user-global location and prints a
ready-to-paste hooks: snippet; ai-memory deliberately does not write files
inside your repositories.
# Stage the scripts and print the snippet to paste into
# <repo>/.poolside/settings.yaml:
ai-memory install-hooks --agent pool --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
The snippet wires Pool's five documented events — SessionStart,
UserPromptSubmit, PreToolUse, PostToolUse, and Stop (Claude-shaped
names, snake_case JSON payload on stdin, verified against Poolside CLI
v1.0.16). Local installs use the native ai-memory hook command, so Pool's
documented tool_name/tool_input file operations honor [capture] ignore_paths and unknown file-tool payload shapes degrade to metadata-only
capture.
Pool has no true session-end event; Stop is a turn boundary. After the final
turn, close the session explicitly; use the exact id when several Pool
sessions are open in the same project:
ai-memory finalize-session --agent pool
ai-memory finalize-session --agent pool --session-id <uuid>
Pool tolerates hook stdout, but model-visible context injection from
SessionStart stdout is not demonstrated, so the session-start hook captures
only and never fetches the (single-use) handoff — recover a prior session's
handoff via the MCP memory_handoff_accept tool. No first-party
install-mcp client and no managed workstream (ai-memory run pool) are
claimed: Pool's native session-store contract is not demonstrated, per
managed-harness contributions.
ZCode (z.ai)
ZCode wires lifecycle hooks in the root hooks block of
~/.zcode/cli/config.json — the same file that holds its other CLI
registration. install-hooks --agent zcode (alias zai) merges ai-memory's
entries into that block around any third-party hooks you already have, and is
idempotent: re-running strips only ai-memory's own entries (marked by
statusMessage: "ai-memory capture") and rewrites them in place.
ai-memory install-hooks --agent zcode --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
# Preview the exact hooks block without writing anything:
ai-memory install-hooks --agent zcode \
--server-url "http://homelab:49374"
Entries are exec-form {"type": "process", "command": <ai-memory>, "args": […]} — ZCode spawns them with the event JSON on stdin and no shell, which the
native ai-memory hook command reads directly, so the local spool, bearer
auth, and [capture] ignore_paths exclusions all apply. Every emitted key is
from ZCode's documented hook schema (type, command, args, enabled,
timeoutMs, statusMessage); ZCode drops entries carrying undocumented keys,
so none are emitted. Six documented triggers are wired: SessionStart,
UserPromptSubmit, PreToolUse, PostToolUse, PostToolUseFailure, and
Stop (verified live against the embedded engine v0.16.5).
PostToolUseFailure fires instead of PostToolUse when a tool throws and
is forwarded onto the same capture channel with the error preserved as the
observation outcome. PermissionRequest is deliberately not installed: its
hook chain races the interactive permission client, and a fast client decision
aborts the losing hook, so passive capture of that event is unreliable by
design.
ZCode injects SessionStart stdout as model context
(hookSpecificOutput.additionalContext), so unlike Pool, Zero, and Grok the
prior session's handoff is delivered automatically at session start. There is
no true session-end event — Stop fires at the end of every turn — so close
finished sessions explicitly; use the exact id when several ZCode sessions are
open in the same project:
ai-memory finalize-session --agent zcode
ai-memory finalize-session --agent zcode --session-id <uuid>
No first-party install-mcp client and no managed workstream
(ai-memory run zcode) are claimed yet.
Hermes Agent (Nous Research)
Hermes declares lifecycle hooks in the hooks: block of
~/.hermes/config.yaml. install-hooks --agent hermes (alias hermes-agent)
prints the block to paste there:
ai-memory install-hooks --agent hermes \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
Hermes splits each configured command with shlex.split and runs it with
no shell, passing the event JSON on stdin — so the generated entry invokes
the native ai-memory hook command directly (the same exec-form shape Zero and
ZCode use) and there is no script bundle to stage. There is also no --apply
that writes the file: ~/.hermes/config.yaml is YAML you also edit, and Hermes
gates user hooks behind its own acceptance prompt (hooks_auto_accept). Paste
the block and let Hermes accept it.
Two events are wired: pre_tool_call → pre-tool-use and post_tool_call →
post-tool-use. Their payload carries tool_name / tool_input, which is what
gives Hermes sessions tool observations, tool-family titles, and
[capture] ignore_paths exclusion enforcement.
Session lifecycle is deliberately not wired here: automatic recall, prompt
capture, session-end and the automatic handoff for Hermes belong to the memory
provider, the community-maintained
ai-memory-hermes-plugin.
A second, hook-driven session-end would close the same session twice. Hermes
ignores session-start hook stdout, so a pending handoff is recovered through
MCP: memory_handoff_list then memory_handoff_accept.
No first-party install-mcp client and no managed workstream
(ai-memory run hermes) are claimed yet.
OpenCode
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client opencode \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
# Plugin — write to ~/.config/opencode/plugins/ai-memory.ts.
# If you have the local wrapper installed, prefer `--apply`:
ai-memory install-hooks --agent opencode --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
# Docker-only preview path; redirect only if you want to write the file yourself:
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent opencode \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
Restart OpenCode after installing or changing the plugin; plugins are loaded at startup.
OpenCode 2 (beta)
The generated plugin targets the OpenCode 2.0.10+ event API (checked against
2.0.14). OpenCode 2 runs one long-lived service behind every CLI, so closing a
terminal is not a session end: each completed root turn instead writes a
deterministic checkpoint (no LLM call) of sessions/<id>.md and refreshes the
session's automatic handoff, keeping one open baton per live session. The next
session claims the latest checkpoint of a session that has captured nothing for
ten minutes. Nothing tells a closed terminal from a parallel session still at
work in the same directory, so this is a heuristic: a session in use keeps its
baton, one session's turn never retires another live session's baton, and a
claim retires only older batons of quiet sessions. A tool or model call that
runs longer than ten minutes without a captured event looks quiet. Startup context is claimed once per root
session and retained on every later model request; child sessions never claim
it or publish a baton.
Assistant text stays opt-in, as for Claude Code and Codex: set
capture_assistant = true on the server and install with
ai-memory install-hooks --agent opencode2 --capture-assistant --apply. The
plugin hands the last completed text to the native hook, which sanitizes and
caps it before it reaches the spool or the wire; the excerpt then rides in the
next session's automatic handoff. A bare re-apply preserves the opt-in.
For a commented opencode.jsonc, preview install-mcp --client opencode2 and
merge the entry into the existing mcp.servers object by hand: the apply path
writes strict JSON.
The 2.0 beta installs side by side as opencode2 and shares v1's config
dir and session store, but its MCP schema and plugin API changed. Wire it
with the opencode2 client/agent names:
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client opencode2 \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
# Plugin — write to ~/.config/opencode/plugins/ai-memory-opencode2.ts.
# If you have the local wrapper installed, prefer `--apply`:
ai-memory install-hooks --agent opencode2 --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
V2 nests servers under mcp.servers (no enabled field), so the v1
(mcp) and v2 (mcp.servers) entries coexist in the one
~/.config/opencode/opencode.json(c) file — the beta explicitly supports
this mixed nesting, so keep both entries and do not "convert" the file by
removing the v1 one
(see Migrate from V1). ai-memory run opencode2
resumes the same native sessions as ai-memory run opencode through the
opencode2 binary. Both plugins share the one auto-loaded dir while the
beta is side-by-side; a host may warn about its sibling's file (the two
plugin APIs are incompatible) — that warning is benign, and uninstall
removes each file only on its own ownership markers.
Back up
~/.local/share/opencodebefore running the beta against your real data. The two binaries share the oneopencode.dbfile and the beta has migrated its schema in place before, leaving stableopencode1.x broken (upstream #42260). ai-memory only ever opens that database read-only; the migration risk comes from launchingopencode2itself, not from this integration.The beta's background service defaults to port 49374 — ai-memory's own default. Running both on defaults crash-loops the opencode2 service (
Managed service port 49374 ... is already in use). Move one side:opencode2 service set port <free-port>, or start ai-memory with--bind 127.0.0.1:<free-port>(and matching--server-urlinstalls).
On a Gemini/Vertex model, serve Gemini-safe schemas. OpenCode forwards MCP
tool schemas to the configured provider verbatim, and Google's Schema
(Vertex/Gemini functionDeclaration.parameters) accepts only a single type per
field. schemars renders every optional tool argument as a nullable union
("type": ["integer", "null"]), so the provider rewrites it into any_of with
description still beside it and Vertex fails the whole session at
tools/list:
Unable to submit request because `ai-memory_memory_auto_improve` functionDeclaration
`parameters.max_proposals` schema specified other fields alongside any_of.
When using any_of, it must be the only field set.
Either set gemini_safe_schemas = true in config.toml
(AI_MEMORY_GEMINI_SAFE_SCHEMAS=true) on the server, which collapses those
unions to "type": "integer" plus nullable: true for every client, or append
?flavor=gemini to just this client's MCP URL. Runtime validation is unchanged
either way. Gemini CLI and Antigravity CLI normalize schemas client-side and
need neither.
Oh My Pi / OMP
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client omp \
--server-url "http://homelab:49374/mcp" \
--auth-token "$TOKEN"
# Extension — write to ~/.omp/agent/extensions/ai-memory-omp.ts.
# If you have the local wrapper installed, prefer `--apply`:
ai-memory install-hooks --agent omp --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
Restart OMP after installing or changing the extension; extensions are
loaded at startup. The ai-memory CLI accepts --client omp (or
--client oh-my-pi) for MCP and --agent omp (or --agent oh-my-pi)
for hooks; both target OMP's native .omp integration surface.
Pi
Pi does not read a native mcp.json. ai-memory supports Pi through one
generated TypeScript extension at ~/.pi/agent/extensions/ai-memory-pi.ts; the
same file captures lifecycle events and bridges ai-memory's HTTP MCP tools into
Pi with pi.registerTool. When PI_CODING_AGENT_DIR is set (it relocates
Pi's whole ~/.pi/agent home), the extension is written to
$PI_CODING_AGENT_DIR/extensions/ai-memory-pi.ts instead.
The Pi and OMP extensions use distinct filenames (ai-memory-pi.ts and
ai-memory-omp.ts) so installing one never overwrites the other. They are
not interchangeable — only Pi's bridges MCP tools.
OMP profiles
omp --profile <name> relocates OMP's agent home to
~/.omp/profiles/<name>/agent. Point the installers at the same profile so
the extension and MCP entry are installed in that directory:
ai-memory install-hooks --agent omp --profile work --apply
# Or pass OMP_PROFILE to each installer:
OMP_PROFILE=work ai-memory install-hooks --agent omp --apply
OMP_PROFILE=work ai-memory install-mcp --client omp --apply
--profile takes precedence over OMP_PROFILE. The legacy PI_PROFILE is
used only when OMP_PROFILE is unset. Names are trimmed and validated using
OMP's rules. An empty environment value, a whitespace-only name or default
selects the default profile; an explicit --profile "" is rejected.
A named profile ignores PI_CODING_AGENT_DIR. The default profile uses it
unless it points to the named profile directory exported by a parent OMP
process. PI_CONFIG_DIR changes the .omp root relative to your home using
OMP's path-joining rules; it does not replace an explicit agent directory.
The extension and mcp.json stay in the agent directory even when sessions
move to XDG storage. See native adapter behavior
for the session paths.
uninstall --profile <name> removes the profile's integration files. It also
checks the default profile and the legacy ~/.omp locations, including when
PI_CONFIG_DIR selects a different root.
ai-memory install-hooks --agent pi --apply \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
# `install-mcp --client pi` prints this guidance instead of writing mcp.json:
ai-memory install-mcp --client pi --server-url "http://homelab:49374/mcp"
Restart Pi after installing or changing the extension. OMP / Oh My Pi remains
separate and continues to use .omp paths.
Bind mounts vs docker cp
The setup-agent subcommand does the extract + render in one shot
using a bind mount:
docker run --rm -v "$HOME/.ai-memory:/host" \
akitaonrails/ai-memory:latest \
setup-agent --agent claude-code --to /host/hooks \
--host-prefix "$HOME/.ai-memory/hooks" \
--server-url "http://homelab:49374" --auth-token "$TOKEN"
This works cleanly when the container user's UID matches the host
user's UID (e.g. the homelab where both are 1000). It fails on
rootless Docker and on hosts with userns-remap enabled - the
container can't write to a host directory that belongs to a UID
outside the user-namespace mapping.
The docker cp pattern recommended above sidesteps all of that
because docker cp is mediated by the docker daemon and outputs
files owned by the user running the command. Prefer it as the
default; reach for setup-agent only when your docker setup is
known not to remap UIDs.
Other MCP clients
See docs/mcp-install.md for the per-client MCP
config file path and snippet, or one-shot it via:
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client cursor --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent cursor --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client claude-desktop --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client gemini-cli --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent gemini-cli --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client antigravity-cli --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent antigravity-cli --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client grok --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent grok --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client openclaw --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent openclaw --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client kiro-cli --auth-token "$TOKEN" \
--server-url "https://memory.example/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent kiro-cli --auth-token "$TOKEN" \
--server-url "https://memory.example"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client command-code --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent command-code --auth-token "$TOKEN" \
--server-url "http://homelab:49374"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client vscode-copilot --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client zed --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
docker run --rm akitaonrails/ai-memory:latest \
install-mcp --client zcode --auth-token "$TOKEN" \
--server-url "http://homelab:49374/mcp"
Cursor, Gemini CLI, Antigravity CLI, Grok Build CLI, Kiro CLI, Command Code, and OpenClaw support both
install-mcp and install-hooks. Grok's install-mcp --client grok writes
$GROK_HOME/config.toml (default ~/.grok/config.toml); its hooks live under
$GROK_HOME/hooks (default ~/.grok/hooks). install-hooks --agent grok
captures lifecycle events.
Grok ignores SessionStart stdout and discards an allowing UserPromptSubmit,
so those hooks do not accept the handoff. The first PostToolUse prints
hookSpecificOutput.additionalContext (pending handoff, plus an opted-in
[briefing]). The model sees it after that tool result, not before the
first prompt. A session with no tool call leaves the handoff for
memory_handoff_accept. Claude Desktop, VS Code Copilot, Zed,
and ZCode
are MCP-only here, so you'll need to nudge the model to call
memory_query / memory_handoff_accept itself.
For clients with install-hooks support, the capture path handles
handoff injection at session start or the client's closest equivalent, except
for Zero's no-stdout SessionStart behavior. Grok delivers on the first
PostToolUse instead (Antigravity CLI uses PreInvocation).
Installing hooks without docker
If you only need to use ai-memory from a machine (i.e. that machine doesn't run the server), download and verify the release installer. The installer then downloads and verifies the release's hook archive before writing any scripts:
installer_base=https://github.com/akitaonrails/ai-memory/releases/latest/download/ai-memory-install-hooks
installer_tmp="$(mktemp -d)"
trap 'rm -rf "$installer_tmp"' EXIT
curl -fsSL "$installer_base" -o "$installer_tmp/ai-memory-install-hooks"
curl -fsSL "$installer_base.sha256" -o "$installer_tmp/ai-memory-install-hooks.sha256"
expected="$(awk 'NR == 1 { print $1 }' "$installer_tmp/ai-memory-install-hooks.sha256")"
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$installer_tmp/ai-memory-install-hooks" | awk '{ print $1 }')"
else
actual="$(shasum -a 256 "$installer_tmp/ai-memory-install-hooks" | awk '{ print $1 }')"
fi
[ -n "$expected" ] && [ "$actual" = "$expected" ] || { echo "installer checksum mismatch" >&2; exit 1; }
chmod +x "$installer_tmp/ai-memory-install-hooks"
"$installer_tmp/ai-memory-install-hooks" --agent claude-code
rm -rf "$installer_tmp"
trap - EXIT
# Then render the JSON config (still wants `ai-memory` somewhere —
# either via docker as a one-shot, or installed locally):
docker run --rm akitaonrails/ai-memory:latest \
install-hooks --agent claude-code \
--hooks-dir "$HOME/.ai-memory/hooks" \
--server-url "http://homelab:49374" \
--auth-token "$TOKEN"
The curl script installer supports
--agent claude-code|codex|cursor|gemini-cli|antigravity-cli|grok|opencode|opencode2|openclaw|omp|oh-my-pi|pi
and --to <dir>; --help prints the full flag list. OpenCode,
OpenClaw, OMP / Oh My Pi, and Pi do not need script extraction because
install-hooks generates TypeScript plugin/extension files for them
instead. For Pi, the generated extension also provides the MCP bridge.
The generated TypeScript integrations survive an unreachable server the
same way the native hooks do: a delivery that fails at the network level
(or gets a 5xx) is written to <data_dir>/hook-spool/ in the exact
format ai-memory hook-drain reads, and the plugin drains that backlog
itself once the server is reachable again — so a day of laptop work off
the server's network is captured, not silently dropped. Each spooled
entry carries an idempotency key, so the plugin's own drain and a
manual hook-drain can race without double-ingesting. Note the spool
lands in the agent host's data dir (AI_MEMORY_DATA_DIR or the
platform default), which is where a native ai-memory install looks.
This path is friction-free when:
- You have curl + bash but not docker
- You don't need to run a local ai-memory server (you're a client of a homelab/remote ai-memory)
Hook command paths across a container boundary
install-hooks --apply stages the hook scripts into the data dir and
writes their absolute paths into the agent's config. When the CLI runs
inside a container but the agent runs on the host, those staged paths
would be container paths the host can't see. Set
AI_MEMORY_HOOKS_HOST_ROOT to the host directory that the staged
hooks/ tree is mounted from and the rendered config uses
<host-root>/<agent>/… command paths instead. The bundled docker
wrappers (bin/ai-memory, bin/ai-memory.ps1) forward this variable
automatically; you only set it by hand for custom container setups.
Running ai-memory without docker
Most users should stick to the docker wrapper from the Quick start. Arch
Linux users have the AUR packages. For any
other host, mise installs a tagged release binary directly from GitHub —
no Rust toolchain needed:
mise use -g github:akitaonrails/ai-memory
This uses mise's GitHub backend,
which downloads the release archive matching your OS/arch
(ai-memory-linux-x86_64.tar.gz, ai-memory-macos-aarch64.tar.gz, etc.),
verifies its checksum against the published .sha256 sidecar, then extracts
it and puts ai-memory on PATH. (mise can also check GitHub artifact
attestation and SLSA provenance where a project publishes them; ai-memory's
release workflow does not emit either today, so only the checksum applies.)
No dedicated mise plugin or
registry entry is required — the backend works against any repo whose
release assets follow this naming convention. By default mise also holds back
the very newest release for a short safety window
(minimum_release_age),
so a fresh tag may resolve to the previous version for a day or so; pin an
exact tag with mise use -g github:akitaonrails/ai-memory@1.30.0 to bypass
that.
cargo install ai-memory is not available: the crate name is already taken
by an unrelated project on crates.io, as is ai-memory-core (the workspace's
foundational internal crate). Publishing would require renaming at least
that crate for the registry — a naming decision the project hasn't made yet.
Build from source only when hacking on ai-memory itself or running on a
platform none of the above covers. On macOS, tagged releases also publish
native ai-memory-macos-aarch64.tar.gz and ai-memory-macos-x86_64.tar.gz
archives when you only need the client CLI.
git clone https://github.com/akitaonrails/ai-memory ~/.ai-memory
cd ~/.ai-memory
cargo build --release --workspace
./target/release/ai-memory init # one-time
./target/release/ai-memory serve --transport http \
--bind 127.0.0.1:49374 # MCP + hook HTTP server
Data dir defaults to ~/.local/share/ai-memory on Linux,
~/Library/Application Support/ai-memory on macOS, and the platform
local-data directory on Windows, typically
%LOCALAPPDATA%\ai-memory. Override with AI_MEMORY_DATA_DIR=/path.
To require bearer-token auth, set AI_MEMORY_AUTH_TOKEN in the
server's environment.
Optional serve flags
The serve subcommand also accepts:
| Flag | Env var | What it does |
|---|---|---|
--enable-web |
AI_MEMORY_ENABLE_WEB=true |
Mount the read-only web browser + /api/v1 JSON API. |
--base-path /wiki |
AI_MEMORY_BASE_PATH |
Host the entire HTTP surface (/mcp, /hook, /admin/*, /api/v1, /web) under a configurable subpath — useful behind a reverse proxy sharing a hostname. . and .. segments are rejected; unsafe chars cause a fallback to root with a warning. See docs/https-via-proxy.md. |
--web-slug /web |
AI_MEMORY_WEB_SLUG |
Where the web UI mounts within the base-path. Default /web; set to / to mount the UI at the base-path root. |
--web-ui-dir <path> |
AI_MEMORY_WEB_UI_DIR |
Serve a custom SPA from <path> instead of the built-in browser. ai-memory injects <base href> and <meta name="ai-memory-base-path"> so the SPA can build relative URLs and API calls under the configured prefix. |
--cors-allow-origin <origin> |
AI_MEMORY_CORS_ALLOW_ORIGINS (CSV) |
Allow listed origins to call /api/v1. Layer is scoped only to that route — /mcp, /hook, /admin, and /web remain origin-locked. |
| (config only) | AI_MEMORY_HOOK_RATE_PER_SEC, AI_MEMORY_HOOK_RATE_BURST |
Optional per-actor/session hook ingest token bucket. Unset/0 rate disables it; burst defaults to the rate (minimum one token when enabled). |
On macOS, see docs/macos.md; use the archive matching your
architecture: aarch64 for Apple Silicon, x86_64 for Intel. The
menu bar app is the self-contained GUI path (bundles
the binary, starts the LaunchAgent, opens /web and status). On Windows, see
docs/windows.md.
The short version: run the install commands from the same environment that
launches the agent. WSL2-launched agents need WSL paths and POSIX .sh hooks.
Native Windows agents can use the tagged ai-memory-windows-x86_64.zip, the
Docker Desktop wrapper, or a source build. Native Claude Code uses Claude exec
form with a real ai-memory.exe by default; the Windows Docker wrapper renders
other native Windows script-hook agents through encoded PowerShell .ps1
fallback commands.
When run from source, install-hooks finds the bundled scripts in
the repo's hooks/ automatically. Extracted release archives also
auto-discover the sibling hooks/ bundle beside the ai-memory binary:
./target/release/ai-memory install-hooks --agent claude-code --auth-token "$TOKEN"
(No need for setup-agent in this case - the scripts already live
at the right host path.)
LLM provider tiers
ai-memory works in three intensity tiers:
| Tier | What you get | Env vars | Cost |
|---|---|---|---|
| Zero-LLM (default) | FTS5 + manually declared entity + graph search, rule-based session summaries, auto-handoffs from prompt + tool-call history | (none) | $0 |
| + LLM consolidation | LLM rewrites session pages as coherent narratives; PreCompact checkpoints; LLM-driven contradiction lint | AI_MEMORY_LLM_PROVIDER=anthropic + ANTHROPIC_API_KEY |
~$0.01–0.05 / session |
| + Anthropic via subscription | Same LLM features using a Claude Pro/Max subscription instead of an API key | AI_MEMORY_LLM_PROVIDER=anthropic-oauth + ANTHROPIC_OAUTH_TOKEN |
Uses your Claude subscription |
| + ChatGPT/Codex OAuth | Same LLM features using a ChatGPT Pro/Plus login instead of an OpenAI Platform key | AI_MEMORY_LLM_PROVIDER=openai-oauth + ai-memory auth login openai-oauth |
Uses your ChatGPT subscription |
| + Codex credential reuse | Same LLM features using the Codex CLI-owned login without copying or owning its refresh token | AI_MEMORY_LLM_PROVIDER=codex + an authenticated Codex CLI |
Uses your ChatGPT subscription |
| + GitHub Copilot | Same LLM features using a GitHub Copilot subscription | AI_MEMORY_LLM_PROVIDER=copilot + ai-memory auth login copilot or COPILOT_GITHUB_TOKEN |
Uses your Copilot subscription |
| + LLM reranking | At most one relevance pass over up to 30 bounded project/scopes search candidates; normal order is preserved on invalid, failed, timed-out, or concurrency-saturated responses | AI_MEMORY_RERANKER=llm + any configured LLM provider |
One LLM call per eligible query, at most four concurrently |
| + Hybrid retrieval | Adds vector cosine similarity to FTS5 + entity + graph RRF. Better recall on paraphrased queries | AI_MEMORY_EMBEDDING_PROVIDER=openai + OPENAI_API_KEY (or EMBEDDING_API_KEY) |
~$0.0001 / page on backfill |
Recommended models (chosen as defaults)
If you set only the provider, ai-memory picks a sensible default:
| Setting | Default | Why |
|---|---|---|
AI_MEMORY_LLM_PROVIDER=anthropic |
claude-haiku-4-5 |
Recommended default. Best balance of speed, restraint, and classification quality. Not a reasoning model. Consistently classifies durable project rules as kind: rule. |
AI_MEMORY_LLM_PROVIDER=anthropic-oauth |
claude-sonnet-4-6 |
Anthropic via Claude subscription. Run claude setup-token once; set ANTHROPIC_OAUTH_TOKEN (or CLAUDE_CODE_OAUTH_TOKEN). No ANTHROPIC_API_KEY needed. Same /v1/messages endpoint, Bearer token auth. |
AI_MEMORY_LLM_PROVIDER=openai |
gpt-5.4-mini |
Cheaper + faster alternative. Same parse reliability; mild over-classification on thin sessions. |
AI_MEMORY_LLM_PROVIDER=openai-oauth |
gpt-5.5 |
ChatGPT/Codex backend. Run ai-memory auth login openai-oauth once; ai-memory stores the refresh token in <data_dir>/auth.json and refreshes access tokens automatically. Optional AI_MEMORY_LLM_REASONING_EFFORT (none/minimal/low/medium/high/xhigh/max/ultra/persistent) is mapped to each provider's native reasoning field; omit it to keep the model default. |
AI_MEMORY_LLM_PROVIDER=codex |
gpt-5.6-luna |
Reuses only access_token and account_id from Codex's auth.json; token renewal is delegated to codex app-server --stdio. |
AI_MEMORY_LLM_PROVIDER=copilot |
gpt-5.5 |
GitHub Copilot Chat backend. ai-memory stores a GitHub user token in <data_dir>/auth.json, exchanges it for a short-lived Copilot API token, and refreshes before expiry. |
AI_MEMORY_LLM_PROVIDER=gemini |
gemini-3.5-flash |
Google's hosted option with a generous free tier. ai-memory disables Gemini 3.5 Flash's default dynamic thinking so hidden thought tokens do not truncate strict JSON. Set GEMINI_API_KEY (or GOOGLE_API_KEY). |
AI_MEMORY_LLM_PROVIDER=opencode |
claude-sonnet-4-6 |
OpenCode cloud API. Defaults to the Go endpoint, opencode.ai/zen/go/v1 — a cost-optimised model subset. GPT-5.6 Luna uses Go's Responses endpoint; other models use Chat Completions. For Zen's full catalogue, set AI_MEMORY_LLM_BASE_URL=https://opencode.ai/zen/v1 plus an AI_MEMORY_LLM_MODEL from it; the default model id is Go's. Requests identify ai-memory by version and reuse one session header across related attempts. Both endpoints take OPENCODE_API_KEY (key from opencode.ai/auth). Alias: opencode-zen — historical, and it selects Go like the others; the endpoint is chosen by the base URL, not the alias. |
AI_MEMORY_EMBEDDING_PROVIDER=openai |
text-embedding-3-small (1536-dim) |
5× cheaper than -3-large with marginal recall loss. |
AI_MEMORY_EMBEDDING_PROVIDER=openai + AI_MEMORY_EMBEDDING_BASE_URL=https://openrouter.ai/api/v1 |
openai/text-embedding-3-small via OpenRouter |
Uses EMBEDDING_API_KEY, else reuses LLM_API_KEY or OPENAI_API_KEY, with the OpenAI-compatible embedding client. |
AI_MEMORY_EMBEDDING_PROVIDER=openai + AI_MEMORY_EMBEDDING_BASE_URL=https://api.orcarouter.ai/v1 |
openai/text-embedding-3-small via OrcaRouter |
Uses EMBEDDING_API_KEY, else reuses LLM_API_KEY, with the OpenAI-compatible embedding client. |
AI_MEMORY_EMBEDDING_PROVIDER=voyage |
voyage-3 (1024-dim) |
Voyage's current general-purpose recommendation. |
AI_MEMORY_EMBEDDING_PROVIDER=google / gemini |
gemini-embedding-001 (768-dim) |
Google-hosted embeddings via embedContent. Set GEMINI_API_KEY (or GOOGLE_API_KEY). |
AI_MEMORY_EMBEDDING_PROVIDER=openai-compat |
no default — set model, dim, and base URL explicitly | Self-hosted engines (Ollama, LM Studio, vLLM). Keyless by default; EMBEDDING_API_KEY, else LLM_API_KEY, is sent as a bearer token when present (gateways). Example: AI_MEMORY_EMBEDDING_BASE_URL=http://localhost:11434/v1, AI_MEMORY_EMBEDDING_MODEL=nomic-embed-text, AI_MEMORY_EMBEDDING_DIM=768. Switching an existing openai+base-URL setup to openai-compat changes the stored {provider, model, dim} triple — run ai-memory embed --force to re-embed. Asymmetric models need AI_MEMORY_EMBEDDING_QUERY_PREFIX / AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX, e.g. nvidia/Nemotron-3-Embed-1B-BF16 wants query: / passage: — see docs/llm-providers.md for that and for Qwen3-Embedding/instruction-tuned E5, which need a different (query-only) format. |
AI_MEMORY_EMBEDDING_PROVIDER=copilot |
text-embedding-3-small (1536-dim) |
Reuses the copilot LLM provider's OAuth login (ai-memory auth login copilot, COPILOT_GITHUB_TOKEN, or GITHUB_COPILOT_API_TOKEN) — no separate API key. Calls Copilot's /embeddings endpoint following the OpenAI-compatible contract Copilot documents for chat; that endpoint's exact shape is not covered by a live test against Copilot here, so treat it as needing a real-Copilot smoke test. |
What we don't recommend: reasoning-mode models (Claude with extended thinking, GPT-o3, Gemini "thinking" variants) — they burn token budget on internal reasoning and hang or emit empty responses with the strict-JSON consolidation prompt. Turn reasoning off if you must use one.
Embedding credentials (EMBEDDING_API_KEY)
EMBEDDING_API_KEY is optional and credentials the embedding role alone. Set
it when the embedder should authenticate against a different provider than the
LLM — the openai and openai-compat embedders otherwise borrow the chat
model's OPENAI_API_KEY or LLM_API_KEY, which is then sent to whatever
AI_MEMORY_EMBEDDING_BASE_URL points at and rejected with a 401.
Precedence for AI_MEMORY_EMBEDDING_PROVIDER=openai:
EMBEDDING_API_KEYOPENAI_API_KEYLLM_API_KEY, only whenAI_MEMORY_EMBEDDING_BASE_URLis set
openai-compat checks EMBEDDING_API_KEY, then LLM_API_KEY, and stays
keyless when neither is set. voyage and google/gemini are unaffected:
they read their own VOYAGE_API_KEY and GEMINI_API_KEY/GOOGLE_API_KEY and
never borrow another role's key. With EMBEDDING_API_KEY unset, key
resolution is exactly what it was before the variable existed.
# Chat on api.openai.com — `openai` is the provider that sends
# `max_completion_tokens`, which gpt-5 / o-series models require and
# `openai-compat` never emits — with embeddings on another endpoint.
export AI_MEMORY_LLM_PROVIDER=openai
export OPENAI_API_KEY=sk-...
export AI_MEMORY_EMBEDDING_PROVIDER=openai
export AI_MEMORY_EMBEDDING_BASE_URL=https://openrouter.ai/api/v1
export EMBEDDING_API_KEY=sk-or-v1-...
Anthropic via Claude subscription (OAuth)
Warning
Unofficial and against Anthropic's usage policies — use at your own risk. Anthropic provides no public OAuth API for the Claude Pro/Max subscription; this reuses the
claude setup-tokencredential against/v1/messages, which is not a supported or sanctioned integration. Anthropic's terms reserve subscription (Claude Code) access for interactive use, and using it as an automated API backend may breach those terms and could get your account rate-limited, flagged, or banned. The header recipe is also undocumented and can change without notice. If you want a supported path, use theanthropicprovider with a real Platform API key. We ship this purely as an opt-in convenience and make no guarantees about it.
anthropic-oauth is for Claude Pro/Max subscribers who want to use their
existing subscription instead of an Anthropic Platform API key. It hits the
same /v1/messages endpoint as the anthropic provider — only the auth
headers differ (Bearer token + anthropic-beta: oauth-2025-04-20).
# Obtain a token once using the Claude Code CLI:
claude setup-token
# Then export it (the CLI may also write CLAUDE_CODE_OAUTH_TOKEN automatically):
export ANTHROPIC_OAUTH_TOKEN=<paste token here>
export AI_MEMORY_LLM_PROVIDER=anthropic-oauth
ai-memory serve
For Docker, pass the token as an env var:
docker run -d --name ai-memory \
-p 127.0.0.1:49374:49374 \
-v ai-memory-data:/data \
-e AI_MEMORY_LLM_PROVIDER=anthropic-oauth \
-e ANTHROPIC_OAUTH_TOKEN=<token> \
akitaonrails/ai-memory:latest
Both ANTHROPIC_OAUTH_TOKEN and CLAUDE_CODE_OAUTH_TOKEN are accepted;
ai-memory checks ANTHROPIC_OAUTH_TOKEN first. When either variable is exported
on the host, the POSIX and PowerShell Docker wrappers forward its name to
short-lived helper commands such as llm-test; the token value is inherited by
Docker rather than placed in the wrapper's command line. The long-lived server
container still needs the provider and token variables in its own environment,
as in the example above.
For both Anthropic providers, ai-memory omits temperature for Claude
4.7 and later models and Claude Mythos Preview because those models reject
non-default sampling parameters. llm-test deliberately starts with the same
representative 0.2 value as bootstrap and consolidation, then exercises the
provider's compatibility normalization before sending the request.
Tip
Pick a small, fast model. ai-memory's LLM work — session consolidation, lint, and explore — is summarisation/extraction, not hard reasoning, so a Haiku-class model is plenty: faster, cheaper, and far easier on subscription rate limits than Sonnet/Opus. Set e.g.
AI_MEMORY_LLM_MODEL=claude-haiku-4-5. Save the high-effort thinking models for your actual coding agent.
OpenAI OAuth / Codex
openai-oauth is for ChatGPT Pro/Plus/Codex accounts. It does not use
OPENAI_API_KEY and it does not call api.openai.com; requests go to the
ChatGPT/Codex Responses backend with a refreshable OAuth token.
For the Docker quick start wrapper, this writes into the same named volume the
server mounts at /data:
ai-memory auth login openai-oauth
docker run -d --name ai-memory \
-p 127.0.0.1:49374:49374 \
-v ai-memory-data:/data \
-e AI_MEMORY_LLM_PROVIDER=openai-oauth \
akitaonrails/ai-memory:latest
For a remote Docker host, run the login on that host against the same container or data volume:
docker exec -it ai-memory ai-memory auth login openai-oauth
Use ai-memory auth status to check whether a token is present and
ai-memory auth logout openai-oauth to remove it.
Tip
Leave the model at the provider default (
gpt-5.5). The Codex/ChatGPT backend behindopenai-oauthonly accepts a small server-defined set of model ids and rejects others — includinggpt-5-mini— with a deterministic 400, so do not setAI_MEMORY_LLM_MODELfor this backend. Consolidation / lint / explore are summarisation tasks, so if the default reasoning is too heavy setAI_MEMORY_LLM_REASONING_EFFORT=noneorlowinstead, so hidden thought tokens do not eat the JSON budget. Reserve high-effort reasoning for your coding agent.
Codex credential reuse
The independent codex provider reads $CODEX_HOME/auth.json, falling back to
the platform home's .codex/auth.json. It materializes only
tokens.access_token and tokens.account_id, reloads them before every call,
and never copies or writes the file. On the first 401, it asks
codex app-server --stdio to refresh the Codex-owned credential and retries
the Responses request once.
export AI_MEMORY_LLM_PROVIDER=codex
export AI_MEMORY_LLM_MODEL=gpt-5.6-luna
export AI_MEMORY_LLM_REASONING_EFFORT=medium
ai-memory llm-test --provider codex --model gpt-5.6-luna --prompt "Reply with OK"
ai-memory llm-test --provider codex --model gpt-5.6-luna --structured --prompt "Return a short answer"
AI_MEMORY_CODEX_EXECUTABLE optionally selects another Codex binary. File
storage is supported; auto is supported when it resolves to the same
auth.json. Keyring-only and ephemeral storage are not supported. Docker is
outside the automatic setup path: both the executable and credentials must be
available inside the same container/environment.
GitHub Copilot
copilot uses a GitHub user token, then exchanges it for a short-lived Copilot
API token through https://api.github.com/copilot_internal/v2/token. The raw
GitHub token is never sent to api.githubcopilot.com.
For the Docker quick start wrapper:
ai-memory auth login copilot
docker run -d --name ai-memory \
-p 127.0.0.1:49374:49374 \
-v ai-memory-data:/data \
-e AI_MEMORY_LLM_PROVIDER=copilot \
akitaonrails/ai-memory:latest
For a remote Docker host, run the login against the same data volume:
docker exec -it ai-memory ai-memory auth login copilot
Non-interactive deploys can set COPILOT_GITHUB_TOKEN instead. ai-memory also
accepts GH_TOKEN and GITHUB_TOKEN when running natively; prefer the explicit
COPILOT_GITHUB_TOKEN in Docker so you do not pass a broad token by accident.
Advanced users with a pre-minted Copilot API token can set
GITHUB_COPILOT_API_TOKEN and optionally COPILOT_API_URL.
auth login copilot defaults to GitHub Copilot's public device-flow client id.
Pass --client-id or set AI_MEMORY_COPILOT_CLIENT_ID if you operate your own
OAuth app.
OpenAI-compatible providers (Ollama / vLLM / LM Studio / hosted APIs)
docker run -d --name ai-memory \
-p 49374:49374 \
-v ai-memory-data:/data \
-e AI_MEMORY_AUTH_TOKEN="$TOKEN" \
-e AI_MEMORY_LLM_PROVIDER=openai-compat \
-e AI_MEMORY_LLM_BASE_URL=http://host.docker.internal:11434/v1 \
-e AI_MEMORY_LLM_MODEL=qwen2.5-coder:14b \
akitaonrails/ai-memory:latest
There is no safe default model for openai-compat; the env var is
required. For OpenRouter (Kimi, DeepSeek, etc.):
-e AI_MEMORY_LLM_PROVIDER=openai-compat
-e AI_MEMORY_LLM_BASE_URL=https://openrouter.ai/api/v1
-e AI_MEMORY_LLM_MODEL=moonshotai/kimi-k2.6
-e LLM_API_KEY=sk-or-v1-...
AI_MEMORY_LLM_REASONING_EFFORT is honoured on this path: OpenRouter hosts
send reasoning: { effort, exclude: true }; https://api.x.ai (Grok)
sends Chat Completions reasoning_effort (clamped to low/medium/high/
xhigh, because Grok cannot disable reasoning); other compat endpoints send
OpenAI-style reasoning_effort. Anthropic and Anthropic-OAuth map the same
key to output_config.effort and adaptive/disabled thinking on models that
accept those fields (Haiku 4.5 omits them so the default model does not 400;
Fable 5 / Mythos 5 / Mythos Preview omit thinking: disabled because those
models reject it). ultra and persistent clamp to max on OpenAI-style
hosts. Gemini and Copilot ignore the key.
Atlas Cloud uses the same provider; no Atlas-specific ai-memory provider is needed. Pass its API key through the generic compatibility credential:
-e AI_MEMORY_LLM_PROVIDER=openai-compat
-e AI_MEMORY_LLM_BASE_URL=https://api.atlascloud.ai/v1
-e AI_MEMORY_LLM_MODEL=qwen/qwen3.5-flash
-e LLM_API_KEY="$ATLASCLOUD_API_KEY"
Replace the model with another current Atlas model id when needed. ai-memory does not select a default for hosted compatibility endpoints.
OrcaRouter uses the same provider; no OrcaRouter-specific ai-memory provider is needed. Pass its API key through the generic compatibility credential:
-e AI_MEMORY_LLM_PROVIDER=openai-compat
-e AI_MEMORY_LLM_BASE_URL=https://api.orcarouter.ai/v1
-e AI_MEMORY_LLM_MODEL=openai/gpt-4o
-e LLM_API_KEY=sk-orca-...
Replace the model with another current OrcaRouter model id (same
provider/model format as OpenRouter, e.g. anthropic/claude-sonnet-4.6 or
deepseek/deepseek-v4-flash) when needed.
Cheaper Inference is an OpenAI-compatible LLM gateway that advertises models below each lab's list price. It uses the same provider; no dedicated ai-memory provider is needed. Pass its API key through the generic compatibility credential:
-e AI_MEMORY_LLM_PROVIDER=openai-compat
-e AI_MEMORY_LLM_BASE_URL=https://api.cheaperinference.com/v1
-e AI_MEMORY_LLM_MODEL=gpt-5.4-mini
-e LLM_API_KEY=ci_live_...
Model ids are bare, without a provider/ prefix. Replace the model with
another current Cheaper Inference model id (e.g. claude-haiku-4.5 or
deepseek-v4-flash) when needed. Cheaper Inference serves chat models only
and has no embeddings endpoint; configure embeddings separately.
API Route also uses the existing
openai-compat provider. Supply an API Route key and a model ID from its
current catalog:
-e AI_MEMORY_LLM_PROVIDER=openai-compat
-e AI_MEMORY_LLM_BASE_URL=https://global.api-route.com/v1
-e AI_MEMORY_LLM_MODEL=gpt-5.5
-e LLM_API_KEY="$API_ROUTE_API_KEY"
Replace gpt-5.5 with the exact model ID you intend to use. As with other
hosted compatibility endpoints, no dedicated ai-memory provider is required.
Configure embeddings separately if your chosen API Route model does not
provide an OpenAI-compatible embeddings endpoint.
OpenAI-compatible structured calls use the operation's JSON Schema by default:
-e AI_MEMORY_LLM_COMPAT_STRICT=true
Modern Ollama, vLLM, LM Studio, llama.cpp, and gateway endpoints honour this
OpenAI-style response_format=json_schema request. ai-memory retries with its
tolerant parser when an endpoint explicitly rejects the structured-output field
or returns a malformed response shape. For an incompatible endpoint, opt out:
-e AI_MEMORY_LLM_COMPAT_STRICT=false
Hosted gateways that stream long completions past the default 300-second
per-request ceiling fail with http: error sending request; raise the
ceiling to match the gateway's worst-case generation time:
-e AI_MEMORY_LLM_TIMEOUT_SECS=900
Send a caller-identifying header to a gateway that requires one
Every chat request already carries User-Agent: ai-memory/<version>, so a
gateway can tell what is calling it. Some also require a header of their own
for request correlation, and reject or throttle traffic without it. Declare
those once — they are sent on every chat request, whatever the provider:
-e AI_MEMORY_LLM_HEADERS=x-opencode-session=prod-01,x-opencode-client=ai-memory
Entries are Name=Value or Name: Value, comma separated. A header value
cannot contain a comma through the env var; use llm_headers = [...] in
config.toml when one must. Headers ai-memory sets itself (authorization,
content-type, x-api-key, x-goog-api-key, anthropic-version,
anthropic-beta, openai-beta, host, content-length) are refused at
startup rather than duplicated onto the request. An entry for user-agent
overrides the default. Values are never logged.
The opencode provider needs no configuration for this: OpenCode asks
callers for x-opencode-session, and that provider already sends one id per
logical operation — stable across retries and the structured-output
fallback, so one consolidation pass reads as one operation in OpenCode's
metrics. Supply the header through AI_MEMORY_LLM_HEADERS to override that,
for instance to tell several ai-memory instances apart under one account.
Match the consolidation budget to a local model's context window
Consolidation defaults to an approximate 100k-token input target plus a 32k
output limit, sized for a 200k-context provider. A local model with a smaller
window can reject the whole request (exceed_context_size_error from
llama.cpp, HTTP 400 from most gateways). Lower both limits so their sum fits
the real context window, with additional headroom for tokenizer variance:
# e.g. a model loaded with an 8k context window
-e AI_MEMORY_CONSOLIDATION__MAX_INPUT_TOKENS=6500
-e AI_MEMORY_CONSOLIDATION__MAX_OUTPUT_TOKENS=1000
The double underscore separates the [consolidation] section from each key.
The input target accounts for the rendered observations, current page body,
system prompt, page conventions, bounded slot snapshots, structured-output
schema, and provider-envelope reserve. Tokenizers differ, so this is a
conservative estimate rather than an exact provider token count. An automatic
checkpoint provider failure degrades to a rule-based page rather than losing
the checkpoint, but right-sized limits are what allow LLM consolidation to
succeed. Startup rejects input targets below 6,000 and output limits below
1,000 because the batch schema and a useful response cannot fit reliably below
those floors.
Common subcommands
This is the operational shortlist. Run ai-memory --help for the authoritative
full command tree.
Two ways to invoke a subcommand against the docker deploy:
# A) Against the running container (stateful: status, search, backup,
# checkpoints, restore-page, audit-contamination, forget-sweep, lint, embed).
docker exec ai-memory ai-memory status --json
docker exec ai-memory ai-memory search "karpathy"
docker exec ai-memory ai-memory backup --to /data/snapshot.tar.gz
# B) One-shot, no running container needed for pure-stdout helpers
# (generate-auth-token, completions, install-mcp, install-hooks, setup-agent,
# llm-test).
# Auth login is stateful: use docker exec against the running container or
# the wrapper so it writes into the same data volume as the server.
docker run --rm akitaonrails/ai-memory:latest generate-auth-token
docker run --rm akitaonrails/ai-memory:latest completions zsh
docker run --rm akitaonrails/ai-memory:latest install-mcp --client cursor
docker run --rm akitaonrails/ai-memory:latest --help # full subcommand tree
| Subcommand | Pattern | What it does |
|---|---|---|
serve |
docker compose up -d (already done) |
Run the HTTP MCP server |
run [harness] [args...] |
host wrapper or native binary | Opt into one managed cross-harness workstream; omit the harness to resume the newest usable local session, or name Claude Code, Codex, OpenCode, Pi, Crush, Kimi Code, Command Code, Kiro CLI v2/v3, OMP, Grok Build CLI, or Antigravity CLI explicitly; exact --yolo and --fresh flags are wrapper-owned and other native arguments pass through |
show [--json] |
host wrapper or native binary | Choose a client-local checkout and installed managed harness, or return structured discovery data without launching; remote servers never provide checkout paths |
continue [--workspace NAME] |
host wrapper or native binary | From any directory, revalidate and resume the newest client-local managed checkout; accepts --yolo and --fresh but no harness-native arguments |
resume [--workspace NAME] [--limit N] |
host wrapper or native binary | Interactively choose a recent managed workstream from valid client-local checkouts; Up/Down selects the workstream and Left/Right cycles auto plus installed harnesses before launch; accepts --yolo and --fresh |
workstreams [--workspace NAME] [--project NAME] [--limit N] [--json] |
host wrapper or native binary | List recent workstreams selectable from the current checkout, including current selection and linked harnesses, without exposing paths or native session ids |
rename-workstream (--from NAME | --workstream-id ID) --to NAME [--workspace NAME] [--project NAME] [--json] |
host wrapper or native binary | Retitle one workstream selectable from the current checkout; metadata only, so the stable id, the ledger, the listing order, and the current selection are unaffected |
workstream-search [query] |
managed child or thin HTTP client | Search the complete visible managed-workstream ledger; the managed child receives its workstream id automatically |
status |
docker exec |
Counts, paths, derived-index diagnostics, and passive LLM/embedding provider health |
search "<query>" |
docker exec |
Wiki FTS5 search + bounded source authority; use MCP memory_query for entity/graph/vector RRF |
write-page |
docker exec |
Manual page write (atomic + indexed) |
backup --to / restore --from |
docker exec |
Snapshot or restore the data dir |
checkpoints / restore-page |
docker exec |
List wiki git checkpoints or restore one markdown page and reindex it |
audit-contamination |
docker exec |
Read-only structural audit for likely cross-project contamination |
forget-sweep / lint / embed |
docker exec |
Manual maintenance; sweep + lint also run on the server schedule by default |
commit -m "…" |
docker exec |
Stage + commit the wiki tree |
reset --confirm |
docker exec |
Wipe data (refuses while siblings alive) |
generate-auth-token |
docker run --rm |
Print a random hex bearer token |
user add-human / list / reset-password / disable / enable / patch |
docker exec or native binary |
Human identities and temporary passwords; never issues an API key |
user add / expire / revive / rotate-token |
docker exec or native binary |
Deprecated 1.x compatibility-token lifecycle backed by legacy-user-token |
api-key add / list / rotate / revoke |
docker exec or native binary |
Native aim_ machine credentials |
auth login openai-oauth |
same data volume as the server | Store a ChatGPT/Codex OAuth refresh token for the optional openai-oauth LLM provider |
auth login copilot |
same data volume as the server | Store a GitHub token for the optional copilot LLM provider |
auth login oidc-device |
same developer data dir as native hooks and thin-client CLI commands | Store a per-developer OIDC device token for native hook authentication and HTTP CLI fallback auth |
install-mcp --client |
docker run --rm |
MCP-config snippet per client |
install-hooks --agent |
docker run --rm |
Hook-config snippet for an existing hooks dir |
setup-agent --agent --to --host-prefix |
docker run --rm -v |
Extract bundled scripts + print config (one-shot) |
install-instructions [--target] [--print] [--no-skills] |
same host environment used for the agent prompt files | Install or update the slim CLAUDE.md / AGENTS.md routing block and, by default, the managed ai-memory Agent Skills |
install-skills [--scope] [--agent] |
same host environment used for the agent skill dirs | Install or update only the managed ai-memory Agent Skills |
uninstall --apply |
same host environment used for install | Remove only ai-memory-owned hooks, MCP entries, instruction blocks, managed skill files, and generated plugin files after content/marker validation. Removing hooks or MCP also deletes the ai-memory run auto-wire sentinels under <data_dir>/autowire-state/, so the next managed launch wires again. Use --mcp-url for custom MCP endpoints and --mcp-name only to narrow removal. |
llm-test --provider … |
docker run --rm -e … |
Smoke-test an LLM provider |
completions <shell> |
docker run --rm or native binary |
Print a bash/zsh/fish/PowerShell/elvish completion script; see shell-completions.md |
Managed routing snippets and Agent Skills
ai-memory's routing install is agent-facing prompt packaging. It does not add a
runtime skill router, and SKILL.md files are not durable memory pages. The
wiki remains the durable source of truth.
ai-memory install-instructions now writes two managed prompt artifacts by
default:
- A slim instruction block in
CLAUDE.md,AGENTS.md, or the file passed with--target. The block is bounded by<!-- ai-memory:start -->and<!-- ai-memory:end -->delimiters that appear alone on their own lines. - Managed ai-memory Agent Skills containing the detailed tool-routing guidance.
Re-running the command is safe. If a project still has the old long ai-memory
block between line-anchored markers, the refresh replaces that block in place
with the slim snippet, leaves unrelated instructions before and after it alone,
and writes a timestamped .bak-* backup before changing an existing file.
Managed skill files contain an ai-memory ownership marker; same-name user skills
without that marker are preserved unless you explicitly force replacement.
Their embedded payloads use LF line endings on every release platform, so CLI
installs and memory_install_self_routing return the same bytes on Windows,
Linux, and macOS. This does not rewrite line endings in user-authored files.
install-instructions --print previews only the instruction snippet; run
install-skills --print when you want to preview the managed skill payloads.
install-instructions flags for skills:
| Flag | Meaning |
|---|---|
--no-skills |
Refresh only the markered instruction block. |
--skills-scope <scope> |
Choose project-local or user-global skill roots. Values: project, global. Defaults to project. |
--skills-agent <agent> |
Choose .claude/skills, .agents/skills, .devin/skills, .grok/skills, or both Claude/Agents roots. Values: claude-code, agents, devin, grok, both. By default, CLAUDE.md targets imply claude-code, AGENTS.md targets imply agents, and both instruction files imply both. |
--skills-target-dir <dir> |
Write managed skill directories below an explicit root instead of inferring from scope and agent. |
--skills-force |
Replace unmanaged same-name skills during install-instructions; without it, they are left untouched and the command exits with an actionable error. |
Use install-skills when the instruction block is already right and only the
Agent Skill files need a refresh:
ai-memory install-skills
ai-memory install-skills --scope global --agent agents
ai-memory install-skills --scope global --agent devin
ai-memory install-skills --scope global --agent grok
ai-memory install-skills --agent both --print
ai-memory install-skills --target-dir .custom/skills --force
install-skills flags:
| Flag | Meaning |
|---|---|
--scope <scope> |
Install into this project or the current user's global skill roots. Values: project, global. Defaults to project. |
--agent <agent> |
Install into Claude Code's skill root, the cross-agent skill root, Devin's skill root, Grok's skill root, or both Claude/Agents roots. Values: claude-code, agents, devin, grok, both. Defaults to claude-code. |
--target-dir <dir> |
Write managed skill directories below an explicit root; --scope and --agent are ignored. |
--print |
Print target paths and SKILL.md contents without writing files. |
--force |
Replace unmanaged same-name skills; without it, user-authored same-name skills are preserved. |
Default skill target roots:
| Scope | --agent claude-code |
--agent agents |
--agent devin |
--agent grok |
|---|---|---|---|---|
project |
.claude/skills |
.agents/skills |
.devin/skills |
.grok/skills |
global |
~/.claude/skills |
~/.agents/skills |
Windows: %APPDATA%\devin\skills; non-Windows: ~/.devin/skills |
$GROK_HOME/skills (default ~/.grok/skills) |
Each managed skill is written as <root>/<skill-name>/SKILL.md.
ai-memory uninstall --only skills --apply removes managed skill files only
from the default project/global roots shown above, after validating the
ai-memory ownership marker. If you installed with --target-dir or
--skills-target-dir, clean up that custom root manually.
Data dir inside the container is /data (mounted via the compose
volume). Outside docker, override with AI_MEMORY_DATA_DIR=/path.
Scheduled maintenance is configured in [maintenance] in config.toml.
By default, rule-based lint and forget sweep run daily outside hook
latency across every existing workspace/project. Embedding backfill is
supported but defaults to off because it can call a paid provider; if you
enable embedding_backfill_interval_secs after configuring an embedder,
each scheduled tick backfills every existing workspace/project and may
increase provider usage accordingly.
Forget sweep and rule-based lint persist their last successful completion. On restart, a job that is not due waits only its remaining interval; a never-run or overdue job runs once after a bounded startup delay. Failed runs are not recorded as successful and retry after that bounded delay. Embedding backfill remains opt-in and keeps its interval-only behavior (no startup catch-up).
reconcile_tombstones_deleted_pages (default false, also under
[maintenance]) is a separate, experimental opt-in: it lets the watcher's
own 30s reconcile pass (not the scheduled jobs above) tombstone an
OKF-imported content page (session summary pages are excluded — see
docs/okf.md) whose file has disappeared from disk, after it has been
missing on two consecutive passes and survived a circuit breaker that
refuses to act when more than max(3, 50%) of a scope's candidate pages
look missing at once, or when a walk finds nothing at all. The tombstone is
soft (is_latest = 0 + superseded_at, the same shape decay eviction uses)
and is picked up by the same aged-tombstone hard-delete sweep — it is not
exempt from it. What actually protects it: a reconcile tombstone is never
itself destroyed while its chain has no successor; if the file returns, the
new version re-links to the tombstoned chain instead of starting fresh, so
nothing is orphaned for that sweep to destroy. It never runs the blocking
admission gate (nothing can refuse it), but does fire-and-forget any
non-blocking observer/mirror webhook. With it off (the default), reconcile's
behavior is unchanged: a deleted file still requires ai-memory delete-page.
See
docs/okf.md for the full design.
Bootstrap mid-project
When you adopt ai-memory in a project that's already been around for
a while, the wiki starts empty. ai-memory bootstrap ingests the
project's existing history into seed pages so the first session has
warm context.
cd /path/to/project
ai-memory bootstrap
If you installed the Docker wrapper from the quick start and started the
server on 127.0.0.1:49374, the wrapper automatically reaches that host
loopback server from its short-lived helper container. Set
AI_MEMORY_SERVER_URL=http://<server>:49374 only when the server is
remote or uses a custom host/port.
What gets ingested by default:
| Source | Priority (dropped first when over budget) |
|---|---|
CLAUDE.md / AGENTS.md (project rules) |
never dropped |
README.md at the repo root |
very-late |
docs/**/*.md |
late |
| Substantive git commits (body >120 chars OR conventional-commit prefix) | mid |
Module-level //! doc-comments in **/*.rs |
first to drop |
Flags:
--repo-path <PATH> (default: git rev-parse --show-toplevel)
--workspace <NAME> (default: the nearest `.ai-memory.toml` marker's
`workspace`, else "default")
--project <NAME> (default: the marker's `project` when pinned,
else derived from cwd — main repo root's
basename via `git rev-parse --show-toplevel`,
or basename(cwd) when no repo is found.
"scratch" only as a defensive fallback for
hook events with no usable cwd.)
--max-input-tokens N (default: 150000; total source budget after prune)
--chunk-input-tokens N (default: 24000; per LLM call, each asking for up
to 16K output tokens; 0 = single call asking for
up to 64K, so input + 64K must fit the context)
--since "30 days ago" (git log filter; supports "N days/months/years ago" + YYYY-MM-DD)
--exclude-git (skip commit history)
--exclude-readme (skip README)
--exclude-docs (skip docs/**/*.md)
--exclude-code (skip Rust module headers)
--dry-run (collect + estimate but don't call LLM or write)
--force (re-bootstrap, overwrites the prior manifest)
Budgets and the estimate. Bootstrap estimates tokens as bytes ÷ 4, which
undercounts non-English text and code, so it fills only 80% of
--max-input-tokens and of each --chunk-input-tokens by that estimate. The
dry run still reports tokens by the same estimate; the sources it keeps and the
chunks it plans reflect the margin.
Cost. With Kimi 2.6 via OpenRouter ($0.73/$3.49 per M):
- 50k input tokens cap → ~$0.04 worst case input
- 1-2k generated tokens → ~$0.007 output
- Total: well under $0.20 per run.
Idempotency. The first run produces a per-project bootstrap.md
manifest (at <wiki>/<workspace>/<project>/bootstrap.md) listing every
page generated + a one-paragraph rationale. Re-running without --force
errors out. Delete the manifest (and the generated pages) if you want a
clean re-bootstrap.
Dry-run first. Always worth doing before the real call to see which sources would actually be sent + how many tokens that represents. Output is JSON to stdout.
ai-memory bootstrap --dry-run
{
"sources_collected": 117,
"sources_sent": 22,
"sources_dropped": 95,
"estimated_input_tokens": 48760,
"pages_written": [],
"rationale": "(dry-run; LLM not invoked)",
"dry_run": true,
"llm_chunks": 1
}
Large repos (e.g. years of git history) are pruned client-side before
POST, then processed in sequential LLM chunks so provider context limits
are not exceeded. The CLI logs llm_chunks in dry-run and the final
outcome.
Caveat: LLM-fabricated detail. A bootstrap run can produce
plausible-but-wrong pages (the LLM doesn't know your project, it's
inferring from git history). The wiki is git-versioned precisely so
this is recoverable: review what landed, docker exec ai-memory git -C /data/wiki diff HEAD~1, and revert if it's off.
Logs and read-only sandboxes
The CLI and server write daily-rolling logs to <data_dir>/logs/
(~/.local/share/ai-memory/logs/ by default). When that location is not
writable — sandboxes like ai-jail
mount $HOME read-only or as throwaway tmpfs — ai-memory degrades instead
of failing: it falls back to the OS temp dir, then to stderr-only logging,
printing the exact path that failed at each step. Commands keep working
either way. To keep durable file logs (and durable hook spooling) inside a
sandbox, map the data dir read-write, e.g. ai-jail --rw-map ~/.local/share/ai-memory ….
Human password bootstrap and recovery
Human console login is username/password, not a Bearer pasted into the
browser. Machine APIs keep using Authorization: Bearer (AI_MEMORY_AUTH_TOKEN
or a native aim_ key). Before human auth activates, deprecated GET-only
browser compatibility may exchange the root bearer through HTTP Basic for an
HttpOnly ai_memory_auth cookie; activation disables that path immediately.
The active classes stay isolated: a password only issues a session; a session
never authenticates /mcp/hooks/workstreams; recovery never issues a session;
an API key never logs into /auth/login.
First-time root (greenfield)
Set a one-shot password, then start serve. The engine creates the configured
root_username (default root) with must_change_password=true and marks
bootstrap complete. Later restarts ignore the value even if it stays in the
environment — unset it so the plaintext is not kept in process env.
# At least 12 characters. Must not equal the root bearer, actor-proxy bearer,
# or recovery token, and must not use an `ams_`/`aim_`/`amk_` prefix.
export AI_MEMORY_AUTH__INITIAL_ROOT_PASSWORD='choose-a-long-password'
docker compose up -d # or: ai-memory serve
unset AI_MEMORY_AUTH__INITIAL_ROOT_PASSWORD
The first login must change that password (POST /auth/password with CSRF).
If the password collides with an existing native API-credential hash, startup
fails closed and bootstrap is not marked complete.
Break-glass recovery
If the last human root is lost, set a high-entropy recovery secret (at least
32 characters) and POST /auth/recovery with a new password. Success returns
204, expires the legacy ai_memory_auth cookie, revokes that user's sessions,
and does not set ai_memory_session. Sign in again with the new password.
export AI_MEMORY_AUTH__RECOVERY_TOKEN='at-least-32-characters-of-entropy-here'
# restart serve so the process picks up the new value
curl -sS -D - -o /dev/null -X POST http://127.0.0.1:49374/auth/recovery \
-H 'content-type: application/json' \
-d '{"recovery_token":"'"$AI_MEMORY_AUTH__RECOVERY_TOKEN"'","new_password":"brand-new-pass!!","new_password_confirmation":"brand-new-pass!!"}'
Rotate by changing the env var and restarting; the previous token stops working immediately. Public failures (wrong token, recovery unset, password policy) share one 401 body. The recovery token is not a Bearer, cookie, or login password.
When human mode is on (bootstrap completed, any password hash, or
initial/recovery secrets configured), serve refuses to start unless a
recoverable root exists (role=root with a password and not disabled) or
recovery is configured. Explicit machine-only remote deploys with
AI_MEMORY_AUTH_TOKEN and no human secrets remain valid. Loopback with
neither Bearer nor human auth stays anonymous.
CIDRs in AI_MEMORY_AUTH__TRUSTED_PROXY_CIDRS may supply X-Forwarded-For
for login rate limits; untrusted peers' XFF is ignored. Behind HTTPS, set
AI_MEMORY_AUTH__SECURE_COOKIE=true as noted above.
Operating without auth
For local-only / single-machine deploys you can skip the bearer token:
docker run -d --name ai-memory \
-p 127.0.0.1:49374:49374 \
-v ai-memory-data:/data \
akitaonrails/ai-memory:latest
Notice the bind: 127.0.0.1:49374, not 0.0.0.0:49374. This is the
critical pairing - no bearer token AND loopback only is the only
safe combination. The server refuses an unauthenticated LAN bind before it
accepts requests. --allow-insecure-no-auth can override that refusal only
for an intentional dangerous plain-HTTP deployment; prefer
AI_MEMORY_AUTH_TOKEN or loopback instead.
In a container that refusal becomes a warning, because the container must
bind 0.0.0.0 internally for -p to work at all and cannot see which host
address you published to. The -p above is therefore doing the real work: it
is what keeps this container loopback-only. If you change it to publish on a
LAN address, set AI_MEMORY_AUTH_TOKEN as well.
Then wire up the agent CLI. Both commands default to no auth and
http://127.0.0.1:49374 - no extra flags needed for the local case:
ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply
The installed Docker wrapper runs CLI commands inside a short-lived
helper container. For local loopback servers, it automatically bridges
that helper back to the host's 127.0.0.1:49374, so ai-memory status,
ai-memory search, and ai-memory bootstrap work with the same default
URL as the generated agent config.
SELinux-enforcing hosts
On SELinux-enforcing Linux systems such as Fedora, RHEL, and openSUSE, normal
home-directory labels can prevent the helper container from reaching agent
config even when its UID and GID match the host user. The wrapper checks both
the host enforcement mode and the engine's advertised security options. For the
short-lived helper commands that touch host files (install-*, setup-agent,
uninstall, backup, restore, and bootstrap), it adds --security-opt label=disable; thin-client commands remain confined when they use the named
data volume and implicit configuration. An explicit --config path or a valid
host-backed AI_MEMORY_DATA_DIR also activates the host-file treatment. This
relaxes SELinux label confinement only for that trusted helper invocation. It
does not modify the long-lived ai-memory server, which uses an engine-managed
named volume.
bootstrap is in that list even though it only reads host files: an
unmapped UID and a confined label block reads just as hard, and the failure is
misleading — it degrades silently to no .git found at /work; bootstrapping from README/docs/rules only before dying with Permission denied (os error 13).
The two engines report these facts under different keys. Docker answers
docker info --format '{{.SecurityOptions}}'; podman has no such field and
fails that template, so the wrapper falls back to podman's
{{.Host.Security.Rootless}} and {{.Host.Security.SELinuxEnabled}} when the
Docker probe comes back empty. Rootless engines additionally need -u 0:0,
because only container UID 0 maps back to the invoking host user — on rootless
podman with SELinux enforcing, both adjustments are required and neither alone
lets the write land.
Do not add :z or :Z to the wrapper's whole $HOME bind. Docker's
bind-mount documentation
warns that relabeling system directories such as /home can make the host
inoperable. Docker documents label=disable in the
docker run security options.
ai-memory run, ai-memory show, ai-memory continue, ai-memory resume,
ai-memory workstreams, and ai-memory rename-workstream are the exceptions:
the current wrapper intercepts them and starts a cached checksum-verified native
client on the host, where local checkouts, harness executables, and session
stores exist. It preserves an explicit remote AI_MEMORY_SERVER_URL. If one of
these commands logs data_dir=/data, cannot find a checkout, or cannot find
codex, claude, or another host executable, refresh the stale wrapper with
ai-memory upgrade on that client machine.
Docker compose alternative
If you prefer compose, clone the repo and run:
docker compose -f docker/docker-compose.yml up -d
The bundled compose file already has restart: unless-stopped, a
healthcheck, and the named volume wired up. Agent setup is the same as
the regular Docker path.
Keeping ai-memory up to date
Docker wrapper
The wrapper checks Docker Hub at most once every 24 hours and prints a one-line warning when a newer image is available. Upgrade with:
ai-memory upgrade
The command downloads the wrapper and its SHA-256 checksum from the latest
GitHub Release, refuses an unverified update, pulls the latest Docker
image, re-stages hook scripts under
~/.local/share/ai-memory/hooks/<agent>/ for configured agents, and
prints how to restart the server container so the new binary is used.
Re-running install-hooks --apply remains idempotent: ai-memory
replaces only the hook entries it owns and leaves unrelated hooks alone.
When a Compose file is found, the wrapper first verifies that its project owns
the running ai-memory container. A standalone container is never handed to an
unrelated Compose project just because its file occupies a conventional path;
the wrapper instead writes the inspected standalone recreation script for
review, preserving the existing /data mount and other runtime options.
The macOS menu bar app is not covered by ai-memory upgrade. Rebuild with
./companions/ai-memory-macos/build.sh (or replace /Applications/AI Memory.app
with a newer staged bundle). Wiki, SQLite, config, and models stay in
~/Library/Application Support/ai-memory.
Set AI_MEMORY_NO_VERSION_CHECK=1 to silence the daily check. To pin wrapper
self-upgrades to a fork or tagged release, set AI_MEMORY_WRAPPER_URL=<url>;
the wrapper requires <url>.sha256 unless
AI_MEMORY_WRAPPER_SHA256_URL=<checksum-url> is also set.
Native release binary (Linux / macOS / Windows x86_64)
When PATH points at a GitHub-release ai-memory binary under a writable
user prefix (for example ~/.local/bin, or %LOCALAPPDATA%\ai-memory on
Windows), the same command upgrades the binary itself:
ai-memory upgrade
# optional: pin a tag, or force a re-download of the current tag
ai-memory upgrade --version v2.3.2
ai-memory upgrade --force
The native path downloads the matching release archive
(ai-memory-<os>-<arch>.tar.gz on Unix, ai-memory-windows-x86_64.zip on
Windows) and its .sha256 sidecar from GitHub Releases, verifies the
checksum, replaces the on-disk binary (and a sibling hooks/ directory when
present), then re-stages hooks for agents already under the data-dir hooks
tree. Windows uses rename-aside (.exe → .old, then promote .new) because
a running image cannot be overwritten in place. It refuses Homebrew/AUR//usr
installs (use the package manager), unwritable prefixes (for example Program
Files — download the zip manually), and in-container binaries (upgrade the
host wrapper/image instead). Windows aarch64 has no release asset yet. For
mirrors or hermetic tests, set AI_MEMORY_RELEASE_BASE_URL (or
release_base_url in config.toml) to a Releases-compatible base that serves
{base}/latest/tag and {base}/download/<tag>/<asset> (+ .sha256). That
override is a trust boundary: archive and checksum are fetched from the same
base, so the .sha256 only proves the base served a consistent pair, not that
the binary is genuine — point it only at origins you control. Prefer an
https:// base; a plain-http:// base has no transit protection, so an
on-path attacker can substitute both the archive and its matching checksum.
Each response body is capped at 128 MiB.
Shared notes
When the upgraded server starts, it applies SQLite schema migrations and
pending wiki-structure migrations automatically. No manual database
reset or wiki rewrite is required for normal upgrades. Migrations are
forward-only: after a newer version has applied its schema, an older binary
will refuse to open that data dir (it fails closed rather than risk
corruption), so take a ai-memory backup before upgrading if you might need
to roll back to the previous version.
Upgrading to 2.3.0
2.3.0 is a normal forward upgrade (the only new migration, V64, just adds the
cross-project agent_messages table — nothing existing is altered or removed).
Two capture/UX conveniences are on by default; both are additive and
non-destructive, but worth knowing about for your first session after upgrading:
- First
ai-memory run <harness>auto-installs that harness's hooks + MCP if they were not already wired (idempotent, one-time per harness; it preserves your existing hook config, including a--capture-assistantopt-in). Disable withai-memory run --no-autowireorAI_MEMORY_RUN_AUTOWIRE=false. - The first session in a brand-new (empty) project imports that project's
existing local harness history once, so installing ai-memory mid-project is
not amnesiac. It only ever runs on an empty project (never touches one that
already has captured memory) and is hard-capped. Disable with
AI_MEMORY_BACKFILL_ON_START=false; run it by hand withai-memory backfill.
If the server runs on another host, ai-memory upgrade refreshes only
the local client (wrapper/image or native binary) and local hook scripts.
Redeploy the remote server separately with bin/deploy or
docker compose pull && docker compose up -d in that deploy directory.
Inside ai-jail or another bwrap sandbox, the wrapper is usable from the
sandbox, but run install-* commands outside the sandbox because they
write to ~/.local/share/ai-memory/hooks/.
See also
docs/macos.md- macOS install paths: menu bar app, native release tarball, source build, Docker wrapper, and launchddocs/deploy.md- homelab deploy walkthrough (bin/deploy, cloudflared TLS, env-file management)docs/usage.md- handoffs, proactive querying, web UI, slim routing snippet + managed Agent Skills, migration from other memory tools, and raw-wiki inspectiondocs/mcp-install.md- per-client MCP config reference for every client in the README Support Matrixdocs/ARCHITECTURE.md- what's actually running inside ai-memory