ai-memory's reqwest client is built with rustls-tls-native-roots, so it trusts
the OS store and honors SSL_CERT_FILE/SSL_CERT_DIR. Document how to give the
container a combined CA bundle (public roots + the interception root) so
outbound LLM/embedding calls stop failing with UnknownIssuer behind a
corporate MITM appliance or inspecting antivirus. Cross-referenced from
deploy.md's provider-failures troubleshooting bullet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The docs already tell users that pushing the wiki to a remote git repository
is a supported backup pattern: docs/deploy.md#backups says 'markdown - back
up with rsync or git push to a remote', docs/design-decisions.md says 'No
remote/cloud sync (use git remote on the wiki dir)', and both
docs/airgapped-install.md and SECURITY.md defer the channel security to the
operator ('Remote sync security'). But the recipe itself is nowhere. A
homelab or laptop install has no built-in remote-push cadence, and building
one from scratch takes 100+ lines of bash plus systemd units plus a
gitignore that catches derived state and secrets.
Fill the gap with a docs-only addition:
- New docs/backup.md walks through what to include (wiki/, raw/, config.toml
minus secrets) vs exclude (db/ derived from wiki via reindex, models/
redownloadable, logs/, .serve.lock), the rsync + git push + tarball flow,
the systemd --user timer schedule, restore, and the SECURITY.md-aligned
posture (private repo, least-privilege push credential, secret exclusion,
encryption at rest is out of scope for ai-memory).
- A worked example under docs/examples/backup/ following the precedent set
by docs/examples/jev-reranker-adapter/ and docs/examples/auto-improve-eval/:
the snapshot script (configurable through six env vars), a systemd --user
.service oneshot, a daily .timer, a .gitignore for the mirror repo, and
a README with the install-and-enable steps plus non-systemd equivalents
for macOS launchd, Windows Task Scheduler via WSL, and Docker sidecar.
- One-line pointers from docs/deploy.md#backups (right after the tarball
routine) and docs/airgapped-install.md (extending the existing 'git
remote sync' bullet).
- A row in the README docs table between lifecycle-ops.md and
llm-providers.md, positioned as a companion to lifecycle-ops.md.
The on-box "ai-memory backup --to TARBALL" command is unchanged. This is
docs-only; no core CLI subcommand for remote push is proposed here (the
issue leaves that decision to the maintainer).
Refs #950.
Co-Authored-By: Claude <noreply@anthropic.com>
ai-memory serve leaked one file descriptor per dead hook/MCP peer.
A client whose connection dies without sending FIN (laptop sleep, a
VPN/Tailscale flap, an abrupt kill) leaves the accepted socket
ESTABLISHED forever, since the OS keepalive default is off. Over ~2-3
days of normal churn that exhausts the 1024-fd default and breaks the
healthcheck -- an unauthenticated availability/DoS. The rmcp
session-table half of this leak was already fixed in 2.4.0 by the
rmcp 2.x bump; this closes the remaining half-open-socket half.
Accepted sockets now get TCP keepalive via socket2, wired through
axum::serve::ListenerExt::tap_io (a hand-rolled axum::serve::Listener
newtype was tried first, but into_make_service_with_connect_info's
Connected<IncomingStream<'_, L>> bound is only implemented by axum for
its own TcpListener and, generically, for TapIo<L, F> -- never for an
arbitrary third-party L, and the orphan rule blocks implementing it
ourselves since neither Connected, SocketAddr, nor IncomingStream is
local to this crate). tap_io keeps the real peer SocketAddr flowing to
ConnectInfo while still touching every accepted stream.
New tcp_keepalive_secs config field (default 60s; AI_MEMORY_TCP_KEEPALIVE_SECS
env override; 0 disables keepalive entirely), read once through the
existing Config::load path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The stdio transport listened for no signal at all, and the HTTP transport
listened for SIGINT only, so SIGTERM — what `docker stop`, `docker compose
down` and `systemctl stop` send — reached no handler on either.
What that cost depended on whether the server was PID 1. In the image it
is: the ENTRYPOINT is exec form with no init shim, and for PID 1 the kernel
discards a signal whose handler is not installed, so the signal was not
merely unhandled, it was invisible. `docker stop` sat out its whole grace
period and ended in SIGKILL, `docker kill` was the only way out, and Ctrl-C
on stdio did nothing at all.
Everywhere else — under the native systemd unit, or a plain `ai-memory
serve` in a terminal — the process is not PID 1, so the same signal fell
through to the kernel's default disposition and killed the process
instantly instead. That stop was fast but unclean: no drain, and the
durable SessionEnd consolidation worker cut off mid-flight. For those
operators this change makes stopping slower, up to the five-second bound,
and correct.
Handling the signal was still not enough on stdio. The MCP transport reads
stdin from a tokio blocking thread, a blocking read cannot be cancelled,
and dropping the runtime waits for in-flight blocking work forever — so a
handled Ctrl-C left the process parked on that read. The runtime is now
built by hand and shut down without waiting for blocking work nothing is
reading the result of any more.
Both transports install their listeners before the transport starts, so a
signal arriving during a slow boot (migrations, the pre-migration archive)
is handled rather than lost, and every wait on the shutdown path — the
connection drain and the consolidation worker's join — is bounded at five
seconds, so a stateful or SSE client holding a connection open cannot stall
the exit.
The default destination was the user's home - which inside a container
is the ephemeral container layer, destroyed on the next compose
recreation: the safety archive would silently vanish on the first
redeploy while the receipt still pointed at it. Containers are now
detected (AI_MEMORY_IN_CONTAINER from the official image, /.dockerenv,
/run/.containerenv) and the archive defaults to <data_dir>/backups/ on
the persistent volume. AI_MEMORY_BACKUP_DIR still wins when set.
A destination inside the data dir (the container default, or an
override) is excluded from the walk - self-inclusion would tar the
half-written archive - and the data dir root itself is refused.
MIGRATION-2.0.md gains the server/container section (docker cp
retrieval, volume persistence) and deploy.md points at it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
* feat(scope): make multi-session and multi-user work without configuration
One operator running several harnesses, and several operators sharing a
server, are core capabilities. They worked only if you found and set
`[auto_scope] mode`, which defaulted to `single` - a process-wide
last-write-wins slot - and which neither docs/users.md nor docs/deploy.md
mentioned at all.
That default was not only a stale-read problem. `resolve_write_args`
falls back to the same pointer, so an unscoped `memory_write_page` from
one session could land in whichever project another session had most
recently published.
The default is now `per_actor`: the pointer is keyed by the caller's
identity and session. The reason this is safe as a default is the
fallback rule added here. A keyed miss means one of two very different
things, and they must not be treated alike:
* nothing has EVER been keyed on this process - an MCP-only install
with no lifecycle hooks. There is no better information anywhere, so
the shared slot still answers, exactly as before.
* something HAS been keyed, but not this coordinate - a genuine
mismatch. Answering from the shared slot is what routed a request
into whichever project published last, so it fails closed instead.
Without that distinction, flipping the default would have broken every
hookless MCP-only install. A control proves it is load-bearing: removing
the fallback fails `a_hookless_install_still_reads_the_shared_slot`.
Tests are integration-level on purpose. Unit tests here exercise one
session at a time, which is the exact shape that cannot see a
collaboration or concurrency defect:
* a page written in a project is readable by another operator - pages
are shared, `author_id` is attribution and never a filter;
* a divergent concurrent write supersedes rather than destroys;
* an identical rewrite is idempotent;
* a second accept cannot steal a claimed handoff;
* an owned baton stays with its owner while pages stay shared;
* parallel harnesses and separate operators keep distinct pointers, a
static MCP client resolves through the identity-only slot, and an
actorless caller still reads the shared slot.
Controls verified: making page reads owner-filtered fails the
collaboration test, and removing the handoff state guards fails the
steal test. That second control also corrected my understanding - the
claim is protected by TWO independent `state = 'open'` guards, and
removing only the UPDATE's leaves the metadata lookup still catching it.
AGENTS.md gains invariant 16 so a later change has to argue against this
rather than assume it.
Documented as a behaviour change with a per-setup upgrade table in
docs/auto-scope.md, plus the team guidance docs/users.md and
docs/deploy.md were missing entirely. `mode = "single"` restores the old
behaviour exactly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
* test(scope): pin the same-project-two-machines case
The scenario named first in the request, and the only one covered by
reasoning rather than evidence: identity derives from the checkout's
NAME, never its absolute path, so `/home/alice/work/proj-b` and
`/Users/alice/src/proj-b` are one project and each machine reads what the
other wrote.
Lives in ai-memory-consolidate because that crate owns the name
derivation and also depends on the store, so the property is asserted end
to end instead of in two halves that never meet.
Control: making the derived name path-dependent fails it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
* test(store): measure writer capacity instead of assuming it
All writes funnel through one writer actor, which raises a fair question:
when does that become the ceiling? Measured rather than guessed.
writers= 1 42/s mean 23.894ms
writers= 8 295/s mean 3.392ms
writers= 32 698/s mean 1.432ms
writers=128 700/s mean 1.428ms
Saturation is ~700 writes/second at around 32 concurrent writers, flat to
128 - and latency does not blow up under 4x the load, it plateaus. Past
saturation the bounded 1024-deep channel applies backpressure through an
awaiting send rather than dropping work or growing without limit.
Single-writer latency is fsync-bound, not CPU-bound: one commit per
observation. Concurrency lets SQLite coalesce WAL commits, which is why
throughput rises 17x while per-write latency falls.
For a team this is not close to a constraint - several hundred
concurrently active agents at a pessimistic one tool call per agent per
second. No mitigation is warranted; the queue that would have been the
mitigation already exists.
The throughput test is #[ignore]d so a timing measurement never becomes a
flaky CI gate on a loaded runner. The backpressure test is not: it asserts
a property, that a 1500-write burst against a 1024-deep queue lands every
write.
docs/deploy.md carries the table, how to re-run it, and the two caveats -
the numbers are fsync-bound so slower or network storage will be lower,
and they measure the store rather than the HTTP front door.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(auth): add browser sessions and API credentials
* fix(auth): preserve 1.x compatibility paths
* fix(changelog): preserve released sections
* docs(auth): tie mirror triggers to 2.0 cutover
* docs(store): correct mirror migration version
* docs(changelog): link admin console PR
* fix(migrations): renumber human_auth/api_credentials to V54/V55
Main gained V52__purged_sessions_tombstones and V53__page_embed_failures
after this branch was opened, so the merge produced four migration files
across two version numbers. Git saw no conflict - the filenames differ -
and the collision only surfaces at runtime:
UNIQUE constraint failed: refinery_schema_history.version
on a fresh database, so the merged tree could not open a store at all.
Renumbered V52__human_auth -> V54 and V53__api_credentials -> V55, and
shifted the version numbers the tests pin: `run_to`/`open_to` targets, the
`schema_version` assertions, the rollback assertion, and the test names.
The pre-migration fixture point moves 51 -> 53 because main's V52/V53
create unrelated tables (purged_sessions, page_embed_failures) and touch
neither `users` nor any table these migrations rewrite.
Migration content is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
---------
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Compaction was only reachable by DELETING something - it existed solely
as `--compact` on the destructive commands. A store that had accumulated
free pages through a retention sweep, a forget-sweep, or months of
superseded page versions had no way to return the space except by purging
data worth keeping.
There was also no figure anywhere saying whether a VACUUM would reclaim
anything, which made scheduling one guesswork. `status` now reports
database size, reclaimable bytes and their share of the file, from three
pragma reads.
- `ai-memory compact --confirm` / `POST /admin/compact` reuse
`reclaim_freed_pages` from #540, so there is still one definition of
"reclaimed". They delete nothing.
- `--confirm` is required not because compaction is destructive but
because it holds an exclusive lock: every write blocks until it
finishes. That is an availability decision and should be deliberate.
- Compaction runs through the writer actor like every other exclusive
operation, so it cannot overlap a write.
- `status` advises compaction only above 20% of the file AND 64 MiB. The
raw numbers are always reported; the advice is not.
`database_bytes` is `page_count * page_size`, deliberately not the
filesystem size: a -wal sidecar holds committed pages not yet
checkpointed, so `metadata().len()` moves for reasons compaction has
nothing to do with.
docs/deploy.md documents why an unconditional nightly VACUUM is the wrong
default - exclusive lock, whole-file rewrite, and SQLite reuses free
pages so a store in steady use usually has nothing to reclaim - and gives
a conditional off-hours recipe gated on the reported figure instead.
Control verified: skipping the reclaim inside `compact` leaves 2.5 MB at
2.5 MB and fails the test.
Closes#549
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The spool health section shipped in #441 with no operator-facing
description. `docs/deploy.md` has the troubleshooting list an operator
actually reads when capture looks delayed, and it covered provider health,
embedding mismatches and restart loops but not the one signal that answers
"are my hook events reaching the server".
Documents the output shape, that a non-empty spool means delayed rather
than lost, that the numbers are client-side and so describe the machine
running the command rather than the server, and that the section is also
printed when the server is unreachable.
The sample output was taken from a real run against a seeded spool, not
written by hand.
Refs #428
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every chat provider hardcoded a 300s reqwest timeout. Slow hosted
gateways (observed with free aggregator tiers) stream long completions
past that ceiling, so every request fails with 'http: error sending
request' and LLM consolidation exhausts its retries.
The timeout now lives on each provider as a field applied per request
(RequestBuilder::timeout), defaults to DEFAULT_REQUEST_TIMEOUT_SECS
(300, unchanged), and is overridable end-to-end via the new
AI_MEMORY_LLM_TIMEOUT_SECS env var / llm_timeout_secs config key,
flowing through ProviderConfig::request_timeout_secs. The Copilot
token exchange and the shared OAuth refresh grant are bounded by the
same ceiling.
Google will discontinue gemini-2.5-flash on October 20, 2026.
Update the default model for the gemini provider to gemini-3.5-flash
and keep the thinking-budget workaround for both the legacy model and
the new default. Update docs and smoke-test defaults accordingly.
Tested against the live Gemini API in a Docker container.
Stale-docs audit across the operator-facing surface. None of the
underlying behaviour changed — these are doc fixes for behaviour that
already shipped over the last week of PRs (#60 move-project,
#65 base-path, #68 wikilinks, #70 openai-compat strict, plus the
audit-cleanup follow-ups in the 65682dc…0d32be1 range).
Operator docs
- `README.md`: status badge bumped from v0.2 to v0.8, project
derivation paragraph rewritten to describe the CLI's main-repo-root
resolution (worktrees share one project) vs the hook router's
basename default, write-page example switched to the H1-in-body
convention (passing `--title` still works but invites issue #67's
JSON-escape footgun).
- `docs/install.md`: bootstrap `--project` default is now described as
"derived from cwd" instead of the stale `"scratch"`. Added an
"Optional serve flags" table covering `--base-path`, `--web-slug`,
`--web-ui-dir`, `--cors-allow-origin`.
- `docs/deploy.md`: cross-link to the "Hosting under a subpath"
section in `https-via-proxy.md`.
Admin / lifecycle docs
- `docs/admission-webhooks.md`: payload sample includes the new
`partial_failure` field (purge-project only, skipped on the wire
when false). Spelled out that admission now fires BEFORE the SQL
destruction in both `/admin/purge-project` and the
`/admin/move-project` copy-purge path, so a `Reject` webhook on
`purge_project` leaves the source intact. Documented that
copy-purge fires two distinct webhook events from one request.
- `docs/lifecycle-ops.md`: failure-mode table covers
`WikiError::DestinationExists` (409) and the block-policy 409 with
`conflicts` body. `--on-conflict` documented alongside the JSON
`on_conflict` field for direct `/admin/move-project` callers.
Dedup naming cites the `DEDUP_FROM_TOKEN` const so the doc
round-trips with the source. Safety-matrix `move-project` row
notes the reject-policy escape hatch.
Frontend / web docs
- `docs/frontend-api.md`: removed two stale "Known gaps" claims —
Cache-Control + ETag and CORS both ship today; §5 documents the
cache headers and §9 is a new CORS section. Added `/api/v1/graph`
to the endpoint reference (§4.9). §6 covers the base-path
normaliser's safety rules (dot-segment rejection, fall-back-to-root
on unsafe chars, query-preserving trailing-slash redirect).
- `docs/usage.md`: wikilink description spells out the "literal in
code" guarantee (fence-glyph tracking, indented code, inline code)
and the external-scheme allowlist; cross-links to the base-path
flag docs.
- `docs/https-via-proxy.md`: subpath section now mentions the
dot-segment rejection and the silent root-fallback warning.
LLM / MCP docs
- `docs/llm-provider-comparison.md`: new section on
`AI_MEMORY_LLM_COMPAT_STRICT`, including the narrowed parse-shape
fallback (5xx / auth / transport errors propagate now), the
two-call cost of fallback, and a per-engine recommendation table.
"When to revisit" entry on strict-JSON-schema availability rewritten
to reflect that the feature exists today.
- `docs/mcp-install.md`: verify-it-works tool list now includes
`memory_read_page` and `memory_delete_page` (their omission would
have made users think their install was broken).
863 workspace tests still pass; fmt + clippy clean. No code touched.
Per Phase 2 decision to NOT terminate TLS in-binary: ai-memory
stays HTTP-on-loopback by default (zero change for existing users
on upgrade) and operators front it with a battle-tested reverse
proxy when they cross the deployment shapes that need TLS.
Three new files:
- docs/https-via-proxy.md — the deployment guide. Covers:
- When you DON'T need TLS (loopback + stdio cases honestly
don't, called out up front so we don't add ceremony where
it doesn't earn its keep).
- When you DO need TLS (multi-user, non-loopback bind, /web
from another machine, public exposure).
- Five deployment paths with copy-paste configs:
- Caddy + public domain + Let's Encrypt
- Caddy + internal CA (LAN-only)
- Cloudflare Tunnel (no open ports, TLS at the edge)
- External cert files (Caddy or nginx)
- Native (non-Docker) Caddy
- Each path documents what can go wrong + the trust-install
step for the internal-CA case (the load-bearing manual
step — skipping it produces security theatre).
- Final section "Don't paper over the security gap" calls out
three specific anti-patterns operators reach for when
things "don't work" (disable allowed-hosts, --insecure on
clients, --no-tls-verify on cloudflared).
- docker/compose.tls.caddy.yml — copy-paste compose template
with Caddy front. Three variants documented inline (LE,
internal CA, external cert files). Same service definition
across all three; only the Caddyfile differs.
- docker/compose.tls.cloudflared.yml — Cloudflare Tunnel sidecar
compose template. Walks through the one-time CF dashboard
setup, the .env.production shape (incl. CLOUDFLARE_TUNNEL_TOKEN),
and the client-config commands. No ports exposed; the tunnel is
outbound-only.
Code change (small):
- crates/ai-memory-cli/src/commands/serve.rs: extend the
existing non-loopback startup warning. Previously fired only
when bind wasn't loopback AND no auth token was configured;
now ALSO fires (with different message) when bind isn't
loopback AND auth IS configured — to remind the operator
that bearer tokens still travel cleartext on plain HTTP and
point them at docs/https-via-proxy.md. One-shot at startup,
not a refusal to serve (operators may already be behind a
proxy we can't reliably detect; refusing to bind would break
their flow). Single line of log output, links to the doc.
Cross-linking:
- README "Security" section gets a clear "Want HTTPS?"
paragraph pointing at docs/https-via-proxy.md + the two
compose templates. Spells out the loopback exception so the
single-user happy-path user doesn't think they're missing
something.
- README docs table updated: docs/deploy.md row reframed as
"pointers to the TLS guide"; new row added for
docs/https-via-proxy.md with a full description of what it
covers.
- docs/deploy.md's brief "encrypted transport" section
trimmed from inline two-paragraph table to pointers at the
new dedicated doc + the compose templates. Avoids drift
between two places documenting the same thing.
Why this shape (instead of in-binary TLS):
- Reverse proxies are the production-grade answer anyway.
- Caddy does Let's Encrypt better than we ever would; we don't
have to chase rcgen / instant-acme / rustls CVEs.
- Scope discipline — TLS termination is infrastructure, not
ai-memory's job. Bundling them would repeat cognee's
LiteLLM mistake at a smaller scale.
- The four "thinking you're secure when you're not" failure
modes all become the proxy's problem, not ours: silent ACME
expiry, trust-store dance, cert reload races, fail-open
fallback.
No behavior change for existing single-user installs.
720/720 tests pass.
Adds Google Gemini as a fourth LlmProvider alongside Anthropic, OpenAI,
and openai-compat. Uses native responseSchema for structured output,
inlining schemars $refs and stripping Draft-2020-12 keywords Gemini
rejects. Reads GEMINI_API_KEY (or GOOGLE_API_KEY) and selects via
AI_MEMORY_LLM_PROVIDER=gemini; defaults to gemini-2.5-flash when no
model is set. CLI llm-test gains a matching --provider gemini choice.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Triage from the install-docs audit. Four files, no code changes.
**README.md**
- Promote Security to a top-level H2 right after Quick start (was
H3 buried under Configuring the CLI). Split into "When you need
bearer auth" checklist + "Enabling bearer auth (turn-key recipe)"
+ "DNS-rebinding guard" + "Browser access to /web". Used to be
one flat block.
- DNS-rebinding-guard (AI_MEMORY_ALLOWED_HOSTS) gets its own
subsection in Security instead of being tucked under "Option 3
— Local Ollama" where it was unfindable. Ollama section now
just cross-references it.
- Quick start gets a "Server on a different machine?" callout
(scenario C/D) pointing to docs/install.md walkthrough.
- `ai-memory upgrade` recipe adds a callout for scenarios C/D
explaining the upgrade only touches local wrapper + image +
hook scripts — the remote server needs separate redeploy.
- "Browse the wiki in a browser" section: fix stale claim that
the user should "pass Authorization: Bearer from the browser"
(browsers can't do that for normal navigation). Replaced with
the real flow: native Basic-auth dialog, leave username blank,
paste token as password, 30-day cookie.
- Configuring the CLI: clarify flag-vs-env precedence ("explicit
flags override env vars when both are set").
**docs/install.md**
- New "Server on a different machine" section with the complete
server-side + client-side recipe for scenario C/D, including
AI_MEMORY_ALLOWED_HOSTS.
- "Operating without auth" section rewritten — was using the
fragile setup-agent bind-mount path the rest of the file warns
against. Now uses the plain install-mcp / install-hooks defaults
(loopback + no token).
- "Running ai-memory without docker" prefixed with a note that
it's the advanced path; most users should stay on the docker
wrapper.
- Top of "Configuring other agent CLIs" gets a paragraph
explaining why install-mcp --server-url INCLUDES /mcp but
install-hooks --server-url omits it.
- LLM-tier table updated: Haiku 4.5 as recommended default to
match README; Sonnet 4.5 demoted, reasoning-mode carve-out
added.
**docs/deploy.md**
- LLM provider table re-synced with README (Haiku 4.5 default,
reasoning-mode warning).
- One-line preamble noting the deploy examples assume the wrapper
is installed per README quick-start.
**docs/mcp-install.md**
- Top-of-file callout explaining URL substitution + Bearer header
shape for scenarios C/D. Per-client snippets unchanged (the
callout points the user at what to swap).
Code unchanged; clippy + fmt clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses the three CRITICAL findings from the security audit run
against the LAN-bound homelab deploy. All three are now in effect
on the deployed server (running healthy, auth=true).
## 1. Bearer-token auth middleware (audit critical #1)
axum from_fn_with_state layer applied to /mcp, /hook, and /handoff.
When `AI_MEMORY_AUTH_TOKEN` env var (or `[auth].bearer_token` in
config.toml) is set, every request must carry
`Authorization: Bearer <token>` matching the configured value;
constant-time comparison via `subtle::ConstantTimeEq` so an attacker
on the same LAN can't time-side-channel-recover the token. When
unset, the layer is a no-op — preserves the zero-config local-dev
experience and the existing unit/e2e tests.
401 responses include `WWW-Authenticate: Bearer realm="ai-memory",
error="invalid_token"` per the MCP authorization spec, so conformant
clients can distinguish missing-token from wrong-token cleanly.
Wire is a static shared secret, NOT full OAuth 2.1. The spec mandates
OAuth for HTTP-authenticated servers, but for a single-user homelab
that's massive overkill (clients would need authorization-server
discovery, PKCE, token refresh). Every major MCP client — Claude Code,
Codex, OpenCode, Cursor, Claude Desktop via mcp-remote, Gemini CLI,
OpenClaw — accepts a static `Authorization` header in its config, so
the wire shape is spec-compatible and the client UX stays one extra
line per config.
A loud startup warning fires when the server binds to a non-loopback
address WITHOUT auth — closes the silent-exposure failure mode the
audit flagged.
## 2. Body size limit (audit critical #2)
`DefaultBodyLimit::max(10 MB)` on the outer router. Before this,
axum would stream unbounded bodies into memory; an attacker could
POST a 1 GB envelope to /hook and OOM the writer-actor task that
gets spawned after the 202 ACK.
## 3. Symlink filter in walk_markdown (audit critical #3)
Skip symlinks early in the reconciliation walk. Without this, anyone
with write access to wiki/ could plant a symlink to /etc/hosts or
~/.ssh/id_ed25519 and have the watcher index the target's content.
The sanitiser would still scrub credentials, but we'd be reading
files we shouldn't. New test in watcher.rs:
`walk_markdown_skips_symlinks`.
## Supporting changes
- `ai-memory generate-auth-token [--bytes 32]` subcommand prints a
fresh hex token via OS RNG. Used by the deploy walkthrough to
onboard the operator cleanly.
- All 21 hook scripts regenerated with a `post_hook`/`get_handoff`
helper that forwards `Authorization: Bearer $AI_MEMORY_AUTH_TOKEN`
when set in the hook's environment. Empty/unset = open-mode
pass-through.
- `install-hooks --auth-token <token>` embeds the token in the
rendered hook config's env block so Claude Code passes it to each
hook script automatically. Both Anthropic Claude Code's hooks file
and the Codex/OpenCode equivalents get it.
- `install-mcp --auth-token <token>` embeds the bearer header in
every client's snippet — Claude Code's `--header`, Codex's
`[mcp_servers.<name>.headers]` TOML, OpenCode's `headers` field,
Cursor's `headers` field, Claude Desktop's mcp-remote `--header`
arg, Gemini CLI's `headers` field, OpenClaw's `headers` field.
- e2e test (tests/e2e/handoff_smoke.sh) starts the server with auth
enabled and asserts: (a) unauthenticated request → 401, (b) wrong
token → 401, (c) right token → 200/expected. Plus the existing
9 recall + sanitisation assertions. Still 9/9 PASS with auth on.
- Pre-existing logging.rs UX bug: `ai-memory generate-auth-token`
on a fresh machine failed because logging::init required the data
dir. Now falls back to a tempdir appender so pure-stdout
subcommands work without an `init` step.
- docs/deploy.md: new "Security — bearer-token auth + encrypted
transport" section walking through token generation, sync to
homelab, client update, plus a comparison of cloudflared vs Caddy
for TLS.
## Audit items deliberately NOT addressed in this pass
- Writer-actor timeouts (audit medium): no current exploit; SQLite
WAL doesn't deadlock under load. Defer until concrete need.
- Reset race condition (audit medium): CLI-only, requires local
shell access. Out of scope for the LAN-exposure threat model.
- Unicode normalisation on PagePath (audit medium): low impact,
paths are flat keys not used for filesystem traversal.
- Malformed YAML DoS (audit medium): only reachable if attacker
can already write to wiki/, which auth + symlink filter already
block.
Deployed:
ssh akitaonrails@192.168.0.90 docker ps → ai-memory running (healthy)
curl http://192.168.0.90:49374/handoff → 401 with WWW-Authenticate
curl ... -H "Authorization: Bearer <t>" → 200 / handoff body
startup log: auth=true, body_limit_mb=10
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the deploy plumbing modelled on ~/Projects/frank_mega's bin/deploy
pattern: a single shell script that builds the image, pushes to a
registry, ssh's the homelab, and `docker compose pull/down/up`s.
Lab-specific values (server, deploy dir, image tag, API keys) live in
.gitignored local files; the repo ships only `.example` templates.
Files added (templates, committed):
- bin/deploy — build/push/restart script; sources bin/deploy.env
- bin/deploy.env.example — SERVER / DEPLOY_DIR / IMAGE template
- docker/docker-compose.prod.yml.example — LAN-binding compose with
`security_opt: label:disable` (required on SELinux-enforcing hosts:
openSUSE MicroOS, Fedora, RHEL) and env_file for secrets
- docker/.env.production.example — API key + provider env template
- docs/deploy.md — first-time setup walkthrough + routine ops
Files added to .gitignore (live copies stay local):
- /bin/deploy.env
- /docker/docker-compose.prod.yml
- /docker/.env.production
Supporting changes surfaced while deploying:
- Default LLM model per provider now lands in factory.rs
(anthropic=claude-sonnet-4-6, openai=gpt-4o-mini; openai-compat
still requires explicit AI_MEMORY_LLM_MODEL). M11 referenced this
in its commit message but the actual code change never made it
into factory.rs; this fixes that gap.
- Stale `claude-sonnet-4-7` references swept to `claude-sonnet-4-6`
across cli.rs / payload.rs / README.md / ARCHITECTURE.md.
Anthropic's API returns 404 for `claude-sonnet-4-7`; the current
Sonnet generation is 4.6.
- Dockerfile was the lone file the earlier port-rename sed missed
(no extension matched my `--include`). EXPOSE + CMD now both use
49374; the container was binding to 7777 silently.
- max_tokens bumped 1500→4000 (single-page consolidator) and
2000→4000 (lint contradiction pass) to accommodate reasoning
models. Kimi 2.6 burns ~2k tokens on hidden reasoning before any
visible JSON, which truncated both paths at the old limits. Non-
reasoning models stop early and don't pay extra for the higher cap.
- Removed `#[serde(deny_unknown_fields)]` from Config. Figment's
Env::prefixed("AI_MEMORY_") provider pulls every AI_MEMORY_-
prefixed var into the struct deserializer, including
AI_MEMORY_LLM_*, AI_MEMORY_EMBEDDING_*, etc. — which are read
directly by factory.rs. Strict rejection crashed every deploy
that configured an LLM via env vars.
End state: container running on the homelab at
http://<homelab>:49374/mcp, healthy, configured for Kimi 2.6 (via
OpenRouter / openai-compat) + text-embedding-3-small hybrid retrieval.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>