`ai-memory run` auto-wires hooks whose command is the wrapper's native
client path. Keeping that client under ~/.cache meant a cache flush broke
every hook of Claude Code, Codex, Kimi Code, Command Code, Kiro CLI v3,
Grok and Antigravity CLI. Keep it in $XDG_DATA_HOME/ai-memory instead.
The wrapper also deleted the release's hooks/ bundle after extraction,
so auto-wire failed with "could not locate hooks directory" for
script-based harnesses on a host where install-hooks never ran. Keep
the bundle beside the client, where install-hooks looks for it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012PoLrSBy2WePNoHhD8BJHD
figment's Env provider parses each var's string with its loose-value
parser, whose bare/unquoted branch calls .trim() (figment 0.10.19's
src/value/parse.rs:78, verified against the vendored source) — so
AI_MEMORY_EMBEDDING_QUERY_PREFIX="query: " reached Config::embedding_query_prefix
as "query:", silently dropping the publisher-significant trailing space.
Config::load now overlays the raw (untrimmed) env value for the two prefix
keys via figment::providers::Serialized (which hands figment an
already-typed value, bypassing the string parser), merged after the
generic Env::prefixed pass so it still wins over a config.toml value.
Presence, not non-emptiness, is the signal: a variable set to "" is a
deliberate override clearing a config.toml-configured prefix, distinct
from the variable being absent. Both wrapper scripts (bin/ai-memory,
bin/ai-memory.ps1) now forward these two keys on presence for the same
reason; a non-empty override was already forwarded correctly, only the
empty-override case was silently dropped by the [ -n ] check every other
forwarded var correctly uses.
The overlay itself lives in overlay_embedding_prefixes(), a pure function
taking the env values as parameters rather than reading std::env::var
itself, so it stays directly unit-testable without mutating process
environment or the current directory — the same pattern
ai-memory-cli/src/commands/path_util.rs's agent_config_home and
ai-memory-hooks's drain_with_live_token already use. std::env::set_var is
unsafe under edition 2024 and forbidden workspace-wide, and
figment::Jail calls std::env::set_current_dir on the real process
internally, racing every other test in this crate's multi-threaded lib
test binary that relies on cwd. Four pure unit tests build a minimal
in-memory Figment and assert on the merged Config directly. One
additional test exercises the real Config::load end to end through a
TOML file and an explicit absolute path; since this process's own
environment is shared with every other test in the binary and could
already carry one of the two prefix vars from the test runner's shell,
that test re-execs this same test binary filtered to just itself as a
genuinely separate child process, with both vars removed via
Command::env_remove — real process isolation rather than an in-process
assumption about the ambient environment.
Plus: a wiremock transport test proving OpenAiEmbedder (not just
OpenAiCompatEmbedder) sends the configured prefix on the wire; a
regression test for the memory_query embed_query fix from the prior
commit, using a task-aware fixture embedder whose
embed/embed_document/embed_query methods return distinguishable vectors
so a regression back to calling the wrong one fails a direct equality
assertion rather than an inferred ranking change; fake-Docker argument
tests for the wrapper's presence-based forwarding across
unset/empty/whitespace-only/non-empty, with both env vars explicitly
removed from the child environment before each case so the "unset" case
cannot silently inherit an ambient export from the test runner.
bin/ai-memory.ps1 also notes, briefly, the PowerShell/.NET version an
operator needs for `$env:NAME = ''` to reach the wrapper as set-but-empty
rather than deleted — no pwsh runtime was available to exercise it live.
Also narrows the query/document-prefix docs (config.rs, docs/llm-providers.md,
CHANGELOG.md) with exact, separate templates: Nemotron-3-Embed
(nvidia/Nemotron-3-Embed-1B-BF16) and base E5 use "query: "/"passage: ";
e5-mistral-7b-instruct wants "Instruct: {task}\nQuery: " (trailing space);
Qwen3-Embedding wants "Instruct: {task}\nQuery:" (no trailing space) — both
leave documents plain.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds `embedding_query_prefix` / `embedding_document_prefix` config keys
(env: AI_MEMORY_EMBEDDING_QUERY_PREFIX / AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX)
applied by the openai and openai-compat embedders: an optional string
prepended to query/document text before the existing truncation, so
truncation still bounds the whole input. Asymmetric embedding models
(Nemotron-3-Embed, base E5, Qwen3-Embedding) need a "query: " /
"passage: " instruction their publisher specifies; the OpenAI-compatible
/v1/embeddings wire format has no field for it. Empty by default, and not
trimmed so a publisher's trailing space is preserved.
Also fixes memory_query's embed_query helper (ai-memory-mcp/src/server.rs),
found while auditing every embed/embed_query/embed_document call site: it
called the generic Embedder::embed instead of Embedder::embed_query, so an
asymmetric embedder (google's task-typed embeddings, or these new prefixes)
embedded the search query on the document side.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Windows Docker wrapper forwarded a shorter env allowlist than the
POSIX wrapper, so a host-exported Gemini, Copilot, or OpenCode key never
reached Config::load inside the helper container. Match the POSIX list
and forward OPENCODE_API_KEY from both wrappers.
#698: the Docker wrapper's -e forwarding allowlist carried every provider
credential except GEMINI_API_KEY / GOOGLE_API_KEY, so a gemini provider or
embedder selected inside the container never saw its key and failed with
"provider not configured". Add both to the allowlist; guard with a packaging
test naming every forwarded provider key.
#695: the OKF conformance scan that feeds the pre-migration backup gate
flagged auto-improve `_pending/` staging sidecars (no frontmatter, never
migrated) as nonconformant, so once the backup receipt's archive was deleted
every boot re-archived the whole data dir. Skip the `_pending/` subtree,
matching the watcher indexer and the #669 ledger skip.
Closes#698Closes#695
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
In bin/ai-memory, remote_digest() extracted the first sha256 found in manifest inspect output via grep | head -n 1. Because multi-arch image manifest lists list linux/arm64 first, the comparison against local_digest() always failed on x86_64 and Podman hosts, producing a perpetual false-positive 'a newer image is available on Docker Hub' warning.
Match against the host platform architecture (amd64 or arm64), check against all locally known repo digests, preserve volume mount modes (such as :Z) in emit_docker_run_script, and filter out transient container runtime environment variables.
Forward-merge of the nine PRs that landed on main (the 2.0.4 batch: #642 auth
stale-bearer, #646/#640 LoginLimiter, #644 Cursor attribution, #650/#647
reindex manifest, #638 MCP routing, #652 CI docs, #645 dev-loop/build) into the
2.1 feature train, so release/2.1 carries every fix before 2.1.0 is cut.
Conflicts resolved:
- crates/ai-memory-wiki/src/wiki.rs: 2.1's per-page write lock (page_locks,
#607) and main's manifested_scopes memo (#650) are independent additions to
the same struct/imports/constructor — kept both; imports merged to
{HashMap, HashSet}.
- CHANGELOG.md: [Unreleased] now carries 2.1's ### Added features above main's
### Changed + ### Fixed (the 2.0.4 fixes), Keep-a-Changelog order, single
[2.0.3] section preserved.
- crates/ai-memory-llm/tests/extra_headers_on_the_wire.rs (2.1's #606 test)
relocated into tests/suite/ and declared in mod.rs to satisfy #645's
one-test-binary-per-crate harness convention (caught by the repo_layout guard).
fmt, clippy -D warnings, llm harness, and the repo_layout guard all green.
The edit-to-result loop was ~380s for the workspace on macOS and needed an
environment variable on every command. This makes `cargo t` the whole story
on macOS, Linux, and Windows, with numbers measured along the way.
Build
- `[profile.dev]` keeps only line tables (full debuginfo put ~190 MB of
DWARF in each test binary and made the build linker-bound); dependencies
build at opt-level 1 with no debuginfo; proc macros and build scripts at
opt-level 3, since they are run once per dependent crate.
- Test binaries: 78 to 11 in the everyday loop (13 under `--workspace`).
Each one is a link and, on macOS (Gatekeeper) and Windows (Defender), a
first-run malware scan of the whole file, paid serially by nextest's list
phase before the first test starts. Integration tests now live in
`tests/suite/` and compile into their crate's own test harness (`mod.rs`,
included from `src/lib.rs` under `#[cfg(test)]`, with `extern crate self`
so they keep addressing the public API by crate name). Only the CLI keeps
a separate `suite` target, because its tests run the built executable.
The evals harness leaves `default-members`, so a bare `cargo t` skips its
two binaries while `--workspace` (CI, the hook, `cargo tf`) still builds
them. A repo-layout test fails on an undeclared suite file, a stray
top-level `tests/*.rs`, or a `mod.rs` that `lib.rs` never includes.
- `ai-memory-cli` gains a lib target; `main.rs` is a shim. 806 tests that
lived in the bin are reachable, and `--lib` runs skip the 127 MB binary.
- The web crate's vendored `static/tailwind.css` is the default on every
build, so nothing needs `TAILWIND_SKIP=1` any more: every release, Docker,
and CI path already used the vendored file, and the download branch only
ever ran for developers who forgot the flag (and then rewrote the source
tree as a side effect). `TAILWIND_BUILD=1 cargo build -p ai-memory-web`
regenerates it explicitly. CI runs that on Linux and fails if the
committed file is stale, a check that did not exist before; the committed
file reproduces byte for byte today.
- `tokenizers` aligned on one version instead of the 0.21 pin plus the 0.22
candle pulled in.
Test tiers
- `.config/nextest.toml`: the `default` profile skips any test whose module
path has a segment starting with `slow` or `stress` (`packaging::slow::*`
drives real wrapper scripts and fake container engines at 10-20s each;
`stress_*` modules hammer concurrency), reports every failure in one run,
and marks anything over 5s in its summary so a new slow test is visible
the day it lands. `full` runs everything. `ci` keeps its retries and
writes JUnit.
- `.cargo/config.toml` holds two aliases and nothing else: `cargo t` (default
members) and `cargo tf` (`--workspace -P full`). `cargo t -p <crate>`
builds just that crate. Neither passes `--all-targets`: there are no
examples or benches, and it only added harnesses for two `test = false`
targets.
- `scripts/install-git-hooks.sh` installs an opt-in pre-push hook that runs
the full tier, touching only its own marked block. Two independent things
run the skipped tier: that hook, and CI, which uses `cargo test` and never
reads the nextest config.
Slow tests fixed rather than tiered
- `project_observations` in the consolidator trimmed an over-budget
projection one observation at a time, re-rendering the whole text and
re-scoring every remaining candidate after each removal. Each score scans
the body, so 256 observations of 4k chars cost ~65k body scans per prompt:
14s in production consolidation, exactly as in the unit test. Scores and
per-block sizes are now computed once and the prune subtracts; output is
unchanged and pinned by the existing tests. 13.9s to 0.18s.
- Windows takes ~2s to refuse a loopback connect, so every hook test that
posted to a closed port paid 2s per request. `dead_http_endpoint()` in the
new `ai-memory-test-support` crate accepts and closes instead, with a
fallback to the closed port where binding is denied. devin hook tests:
4.2s to 0.15s each.
- The store unit fixture opened a file-backed SQLite with the default
rollback journal and synchronous=FULL, so ~120 parallel fixtures fsynced
every transaction. journal_mode=MEMORY + synchronous=OFF: 242s to 89s of
test time, p90 1.6s to 0.5s.
- Windows-only tests resolve `powershell.exe` or `pwsh.exe` once per process
and the auto-improve eval fixtures are `.ps1` scripts instead of cmd.exe
batch files; a post-bind settle sleep is gone; the two unpinned
multi-thread tokio tests pin `worker_threads = 4`. The four copies of the
PowerShell resolver and the mcp suite's duplicated `post`/`get` helpers are
now one each.
Not done, with the numbers in AGENTS.md: nextest vs in-process libtest is a
wash per crate and a rout for the workspace (20s vs 309s); the
`local-embeddings` default feature costs ~50s of cold build and ~27 MB per
binary but under a second per relink, so it stays a product default.
Measured: workspace loop ~380s to ~150s on macOS; on a 32-thread Windows box
the warm everyday run is 20s of test time across 2919 tests in 11 binaries,
and the rebuild after a core edit is 13s of cargo with lld plus the
first-run scans.
ai-memory.ps1 sets $ErrorActionPreference = 'Stop' script-wide. In
Windows PowerShell 5.1 a REDIRECTED native stderr (the repo-root
probe's `git rev-parse --show-toplevel 2>$null`) is converted to a
terminating error, so running `ai-memory status` — or any command —
outside a git repository aborted with
`git.exe : fatal: not a git repository … NativeCommandError`.
The probe now runs under a localized SilentlyContinue, gates on
$LASTEXITCODE, and restores the preference in a finally, falling back
to the working directory exactly as before. The wrapper's unredirected
`& $Docker` calls pass their stderr straight to the console and were
never affected, which is why only this probe tripped it.
Verified with the real PowerShell parser (syntax) and a behavioral
reproduction under `$PSNativeCommandUseErrorActionPreference = $true`
(the PS 7 switch that mimics 5.1's Stop-on-native-stderr): the old
pattern throws, the new pattern survives and falls back. Native runtime
proof on Windows PowerShell 5.1 itself is unavailable in this
environment.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
* feat(cli): add interactive workstream resume picker
* chore(changelog): restore the blank line before [1.36.0]
The branch dropped one blank line separating the end of [1.37.0] from the
[1.36.0] heading. Whitespace only, but it sits inside an already-released
section, so `check-changelog-frozen.sh` rejects it - correctly: the guard
cannot tell a stray edit from a rewritten entry, and should not try.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
---------
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Workstream names were fixed at `run --new` time and had no correction path,
so a typo outlived the work it labelled and the only escape was starting a
new workstream and abandoning the ledger attached to the old one.
The rename selects by current name or by the stable id `workstreams` prints,
resolving the same (workspace, project, repository, worktree) identity `run`
selects with. Both selectors repeat the checkout predicate: `workstreams.id`
is globally unique, so without it a caller holding an id from another
checkout could retitle a workstream their request never named. An id outside
the resolved scope reads as absent rather than renamable.
It is metadata only. `workstream_events`, `managed_runs`, and
`workstream_native_sessions` all key on `workstreams.id`, so the rename is a
single-row update with nothing to cascade, and `selected_at` and
`updated_at` are deliberately left untouched — relabelling is not activity.
Neither the discovery listing order nor the workstream a bare `ai-memory
run` resumes moves as a side effect of fixing a typo.
The destination is validated exactly like a `--new` name and refused with a
named conflict when another workstream in the same checkout already holds
it, turning the UNIQUE constraint into a typed error instead of a bare
SQLite failure. Renaming a workstream to the name it already has writes
nothing and is not an error.
The new `/workstream/rename` route requires NormalWrite rather than the
NormalRead its sibling discovery read uses, and resolves scope through
`lookup_existing_scope` so a rename never creates the workspace or project
it names. The Docker shell wrapper routes the command through its native
host client alongside `run`, `show`, `continue`, and `workstreams`, since
repository identity is a host resource the helper container cannot see.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Embeddings are already independently configurable — provider, model,
dimension and base URL each have their own setting — but there was no key
to go with them. `openai_embedding_api_key` returned `OPENAI_API_KEY`
before the base-URL check ran, so pointing `AI_MEMORY_EMBEDDING_BASE_URL`
at a second provider sent it the chat model's credential. The existing
`LLM_API_KEY` fallback only fires when `OPENAI_API_KEY` is absent, which
also takes the `openai` chat provider down, since it reads that same
variable.
`voyage` and `google`/`gemini` already name their own key and are
untouched. Scoped to the two embedders that borrowed another role's:
`openai` resolves EMBEDDING_API_KEY -> OPENAI_API_KEY -> LLM_API_KEY (the
last still only with a custom base URL), and `openai-compat` resolves
EMBEDDING_API_KEY -> LLM_API_KEY, staying keyless when neither is set.
With the variable absent, resolution is byte-identical to before.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
List the managed workstreams that `run --workstream` can select from the
current checkout, so branching between lines of work no longer requires
remembering names or reading the database by hand.
The list resolves the same (workspace, project, repository, worktree)
identity `run` selects with, puts the current selection first, then orders
by recent activity. Rows carry the linked harnesses and the stable
workstream id that `workstream-search --workstream-id` accepts.
Discovery is a read of workstream metadata only: checkout paths, repository
and worktree fingerprints, and native session ids never travel back to the
client. The new `/workstream/recent` route requires NormalRead, resolves
scope through `lookup_existing_scope` (no-create, fails closed), and bounds
the limit server-side.
The Docker shell wrapper routes the command through its native host client
alongside `run`, `show`, and `continue`, since repository identity is a host
resource the helper container cannot see.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
The PowerShell wrapper runs every CLI command inside a short-lived helper
container, so the CLI's default http://127.0.0.1:49374 resolved to that
helper rather than to the Windows host. A healthy loopback-published
server was therefore unreachable from `ai-memory status` with
"Connection refused (os error 111)", while curl from PowerShell worked.
Docker Desktop gives Linux containers no host networking on Windows, the
same constraint the POSIX wrapper's Darwin arm already handles for macOS
(issue #107). Port that semantics rather than rewriting the address
globally: thin-client commands get host.docker.internal injected, while
install-mcp/install-hooks/setup-agent keep rendering the loopback URL
into host-side agent config, where host.docker.internal does not resolve.
An explicit AI_MEMORY_SERVER_URL still wins, so remote/homelab servers
are unaffected.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
`bin/regen-hooks` claimed in its header to be "the single source of
truth" for `hooks/*/*.sh`. It has not been true for months, and running
it today is destructive.
The script templates a flat 3 agents x 7 events matrix. The bundles it
claims to own now number 11 agents / 75 POSIX scripts with per-agent
event vocabularies (`CLAUDE_CODE_EVENTS`, `KIMI_CODE_EVENTS`,
`KIRO_CLI_V*_EVENTS`, Devin's `post-compaction`), a `.ps1` peer each,
bodies delegating to `hooks/_lib.sh` / `hooks/lib/ai-memory-hook.ps1`,
and per-agent behaviour the template has no notion of.
Running it in a scratch worktree rewrites 21 files across claude-code,
codex and opencode, reverting them to a pre-`_lib.sh` shape: no marker
walk-up, no `cwd`/`workspace`/`project` query params, no `session_id`,
and Claude Code's `hookSpecificOutput.additionalContext` wrapper
replaced by bare text. It emits no `.ps1`, so the survivors also break
`bundled_posix_and_powershell_hooks_stay_in_parity`.
Nothing invokes it — not CI, not `bin/release`, not the packaging
scripts. `AGENTS.md` inventories `bin/` as "`ai-memory`, `deploy`,
`release`" and never mentions it; the CHANGELOG never announced it.
The real contract lives in `render_shared.rs`, whose doc comments
already tell contributors to hand-write both files and lean on the
parity test. So delete the footgun rather than rebuild `render_shared.rs`
in bash.
That leaves the parity test as the guard, and it had the same drift
bug in miniature: a hardcoded list of 9 agent dirs, silently skipping
`command-code` and `kiro-cli` since they shipped. Enumerate `hooks/*/`
instead so the next bundle is covered on arrival — the class of defect
closes rather than one instance of it.
Verified: full `-p ai-memory-cli` suite (805 passed), clippy clean,
`tests/hooks/test_lib.sh` 44 passed, and the enumerating assert
confirmed to fail on a bundle with a `.sh` and no `.ps1`.
`bin/deploy` builds with plain `docker build`, which produces an image for
the architecture of the machine running it. Pushing that to a tag whose
current manifest is a multi-arch list replaces the list outright, and every
host on the other architecture fails the next pull with `exec format error`.
This is not hypothetical — it happened to `:latest` on 2026-08-18 and was
reported within hours as #427. The release workflow builds amd64 and arm64
and joins them into a manifest; a homelab deploy from an x86 workstation
then silently flattened that to amd64-only. The registry has been repaired
by re-pointing `:latest` at the `:1.28.1` manifest list.
Two changes so it cannot recur:
- `bin/deploy` inspects the target tag first and exits 65 if it already
resolves to a manifest list, naming the private-tag fix and the
pull-only alternative for deploying a released image.
- `bin/deploy.env.example` shipped `IMAGE="akitaonrails/ai-memory:latest"`,
so the documented setup pointed the operator's build at the public
release tag by default. It now defaults to `:homelab`.
Refs #427
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ai-memory upgrade` could not recreate a container started with plain
`docker run`, so it told the operator to stop and remove it and rebuild
the command from memory — losing ports, mounts and environment in the
process. The old comment claimed the original args were unknowable; they
are not, `docker inspect` has them.
Reconstructs that container's own stop/remove/run — name, restart policy,
published ports, mounts, operator-set environment, and an overridden
command — and writes it for review before it is run.
Written to a 0600 file rather than echoed: a real install's environment
carries provider API keys and AI_MEMORY_AUTH_TOKEN, and printing those
into terminal scrollback or a piped install log is the kind of leak this
project exists to prevent. Environment already baked into the image is
omitted, so the new image's defaults are not frozen to the old ones.
Verified round-trip against live containers: the generated command
recreates an identical container, including an environment value
containing a space, a double quote and a `$`.
Refs #407
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The wrapper decided two container-security adjustments from a single probe:
DOCKER_SECURITY_OPTIONS=$("${DOCKER}" info --format '{{.SecurityOptions}}' 2>/dev/null || true)
`.SecurityOptions` is a Docker-only field. Podman — including through the
podman-docker `docker` shim — fails that template with "can't evaluate
field SecurityOptions in type system.infoReport" and exits 125, which the
`|| true` swallows into an empty string. Both gates then read as "not
rootless, no SELinux", so on rootless podman the helper runs as an
unmapped subordinate UID under a confined SELinux label and every host
file access fails:
$ ai-memory install-mcp --client claude-code --apply
Error: inspecting symlink /home/user/.claude.json
Caused by:
Permission denied (os error 13)
Both conditions are required. Verified by running the helper by hand on
Fedora (rootless podman, SELinux Enforcing):
uid 1000:1000, no label=disable -> Permission denied
uid 0:0, no label=disable -> Permission denied
uid 1000:1000, label=disable -> Permission denied
uid 0:0, label=disable -> updated ~/.claude.json
So the SELinux relaxation added in #218 never fired under podman either:
it hung off the same broken probe, which that PR hoisted into a variable
precisely so both gates could share it. Podman exposes both facts under
its own keys (`.Host.Security.Rootless`, `.Host.Security.SELinuxEnabled`),
which are now consulted when the Docker probe comes back empty. Real
Docker is untouched — its probe answers, so the fallback never runs.
`bootstrap` needed the same treatment. It only reads host files (the repo
bind-mounted at /work, `.ai-memory.toml` markers under $HOME), but an
unmapped UID blocks reads as hard as writes, and it was not in
WRITES_HOST_FILES. The symptom was misleading: it degraded silently to
"no .git found at /work; bootstrapping from README/docs/rules only" and
then died with "Permission denied (os error 13)". Hence READS_HOST_FILES
and TOUCHES_HOST_FILES = WRITES || READS, used by both gates.
The Darwin branch still keys off WRITES_HOST_FILES: Docker Desktop is not
affected by this failure mode and I have no macOS host to verify a change
there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both wrappers' env passthrough lists carried ANTHROPIC_API_KEY and
OPENAI_API_KEY but not ANTHROPIC_OAUTH_TOKEN or CLAUDE_CODE_OAUTH_TOKEN,
so a subscription-token setup (`claude setup-token` plus
AI_MEMORY_LLM_PROVIDER=anthropic-oauth) lost its credential at the
container boundary.
It surfaces as a contradiction: `status` reports the provider error and
recommends `llm-test --provider anthropic-oauth`, and that exact command
then fails with a missing token — because `llm-test` runs client-side, in
the helper container, not on the server that already holds the token in
its own environment. The CLI itself was never the problem: config.rs
reads both names, with CLAUDE_CODE_OAUTH_TOKEN as the documented fallback.
The regression test asserts both wrappers, since each keeps its own
hand-maintained list and both had the same hole. Those lists have drifted
further apart — the PowerShell one is also missing the whole Copilot
credential set, CLAUDE_CONFIG_DIR, and the workstream/hook variables — but
that is a separate defect and is left for its own change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The docker wrappers only passed `-it` when stdin *and* stdout were a
terminal, so every piped or redirected invocation reached the container
with a closed stdin. `ai-memory write-page --body -` therefore read an
empty string and persisted a page with frontmatter and no body — while
still printing the success line, so the data loss is silent. Any other
stdin reader (a hook fed by a pipe) has the same failure mode.
Pass `-i` whenever stdin is not a terminal, keeping `-t` for real
terminals only. `AI_MEMORY_NO_TTY=1` keeps its documented meaning: it
suppresses the TTY, not the stdin attachment.
Add a regression test that spawns the wrapper with stdin on a pipe
against the existing fake-docker harness and asserts the resulting flag
set; it fails without the wrapper change.
Audit follow-up on the PR branch: uninstall edits the same host
agent-config files the install-* commands write, and backup writes its
tarball to a host path — both go through the same $HOME/$PWD bind
mounts, so they hit the same subordinate-UID failure under rootless
Docker. Add both to the WRITES_HOST_FILES set and to the rootless
regression test's subcommand loop; extend the CHANGELOG entry.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Rootless Docker maps container UID 0 back to the real host UID via
rootlesskit's primary user-namespace mapping, but any non-zero UID
passed via -u (including the host UID the wrapper always passes) is
instead routed through a separate subordinate-UID range
(/etc/subuid, typically 100000+) that owns nothing on the host
filesystem. Every bin/ai-memory subcommand that writes to a host
bind-mounted path (~/.claude/settings.json, ~/.claude.json, the hook
staging dir, $PWD/CLAUDE.md, skill directories) failed under
rootless Docker with Permission denied or a misleading "does not
exist" error.
Detect rootless Docker via `docker info --format
'{{.SecurityOptions}}'` and run as -u 0:0 only for the subcommands
that write host-side files (install-mcp, install-hooks, setup-agent,
install-instructions, install-skills). Thin-client commands (status,
bootstrap, search, ...) are untouched since they only touch the
/data named volume, which isn't host-visible.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The README Docker quick-start produced a non-working native-agent setup on
macOS. Three independent breakages are fixed here:
1. The macOS wrapper baked the container-only host.docker.internal URL into
the *host* agent config (the MCP url + AI_MEMORY_HOOK_URL on every hook),
so MCP and capture failed silently. install-mcp, install-hooks and
setup-agent now render the host-reachable http://127.0.0.1:49374,
decoupled from the host.docker.internal URL the wrapper still uses for
its own in-container thin-client commands.
2. Those thin-client commands (status, ...) were rejected with
"403 forbidden host" because the server's loopback-only Host allowlist
excluded host.docker.internal. The Docker image now ships it in the
default AI_MEMORY_ALLOWED_HOSTS; native installs stay loopback-only and
exposed deployments still override the env entirely.
3. setup-agent/install-hooks could not locate the hooks bundle that ships
beside the binary in the release tarball (the probe derived a bogus
/private/hooks/...). The binary-sibling hooks/ dir is now on the search
path, so --source is no longer required.
Issue point 1 (:latest single-arch) is already fixed on main: the release
manifest job publishes :latest as a multi-arch list; it ships with the next
tagged release.
`for_bash_runner` now returns `PosixNative` on native macOS/Linux (mirroring
Windows), so Claude Code hooks invoke `ai-memory hook` — getting the spool +
OIDC fallback — instead of the `.sh` script that POSTs via curl. The Docker
wrapper (`bin/ai-memory`) forces `AI_MEMORY_HOOK_PLATFORM=posix`, so its
host-rendered config keeps the `.sh` scripts (the host has no local binary).
Validated in a Linux container (runtime-source image): install-hooks renders
the binary command by default and `.sh` when `posix` is forced; the rendered
hook spools to `<data_dir>/hook-spool` (0600), drains on session-end, and the
server writes the session page.
Follow-up to #51 fixing two defects the merged diff introduced:
1. Linux silently lost `-u`. The case had `Linux)`, `Darwin)`, `*)`
but only Darwin and the catch-all assigned USER_ARGS, so Linux
fell through with the pre-case `USER_ARGS=()` default and ran the
container as its internal uid 1000. Regressing for any Linux user
with a non-1000 host UID or a bind-mounted `AI_MEMORY_DATA_DIR`.
Fix: invert the default — USER_ARGS=(-u $(id -u):$(id -g)) at the
top, then only Darwin overrides to (). Eliminates the Linux/`*`
duplication too.
2. "${USER_ARGS[@]}" was not guarded for set -u. The PR carefully
applied ${arr[@]+"${arr[@]}"} to TTY_ARGS / NETWORK_ARGS / ENV_ARGS
for set-u compat on macOS's default /bin/bash 3.2 — but the new
USER_ARGS line used the plain quoted form, which re-hits the exact
same unbound-variable crash on Darwin (where USER_ARGS is empty by
design). Fix: use the same ${USER_ARGS[@]+"${USER_ARGS[@]}"}
pattern as the surrounding arrays.
Verified via simulated case-arm expansion: Linux → `-u 1026:1026`,
Darwin → empty (no `-u`), other (FreeBSD/etc.) → `-u 1026:1026`.
`bash -n` clean.
On macOS, Docker Desktop manages file-sharing permissions via its
gRPC/SSH layer. Passing -u <host-uid>:<host-gid> to one-shot wrapper
containers causes a UID mismatch: the data volume is owned by the
container's internal uid 1000 (the ai-memory user), but the host
UID on macOS is typically 501/502.
This causes one-shot commands (status, install-mcp, install-hooks,
etc.) to crash with:
thread 'main' panicked at .../rolling.rs:156:14:
initializing rolling file appender failed: InitError {
context: "failed to create log file",
source: Os { code: 13, kind: PermissionDenied, message: "Permission denied" }
}
Fix: only pass -u on non-macOS platforms. On macOS, omitting -u lets
the container run as its default uid 1000 which matches the volume
owner. Docker Desktop's file sharing handles the rest.
Also fixes unbound variable errors (set -u) for TTY_ARGS, NETWORK_ARGS,
and ENV_ARGS arrays that may be empty.