Add a FutureInfra example to the openai-compat section of docs/install.md, list it in docs/llm-providers.md, and add an [Unreleased] CHANGELOG entry. Docs only.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
marker-file.md gains a section on the `server` key: registering
profiles, the fail-closed rules, inheritance and the walk boundary with
the allowlist-mode caveat that comes with it, check-capture output,
which integrations route or drop, what `uninstall` removes, that two
profiles may share a URL with different tokens (one server, several
identities under #708), and two known limits (older binaries draining a
shared spool; MCP and `ai-memory run` still using the install server).
install.md, cookbook.md, security.md and the README point to it,
security-boundaries.md adds row 11d, and CHANGELOG records the feature
and the backfill fix; the Kimi `{}` fix reached release/2.5 through the
cherry-pick to main (#996), so its commit is no longer part of this branch.
The submitted text stated a specific '15–60% less' range the vendor site does
not substantiate; reword to a claim we can stand behind (below list price)
without quoting an unverified figure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The watcher's reconcile pass only ever handled create/modify events, so a
page whose file vanished from disk stayed indexed forever until an operator
ran `ai-memory delete-page` (#929). A prior attempt at a hard-delete
reconcile-diff (merged as #958's embed half only) was reviewed as unsafe:
three false-positive classes (indexed-but-unwalked reserved paths, a
walk/list race, and partial walks on NotFound) could each make a live page
look "deleted".
This adds the safer design proposed on the issue instead: a new opt-in,
default-off `[maintenance] reconcile_tombstones_deleted_pages` flag lets the
watcher soft-tombstone (`is_latest = 0` + `superseded_at`, mirroring decay
eviction) a page only after its file has been missing on two consecutive
reconcile passes, with the DB-side candidate snapshot taken before the walk
runs and re-verified live immediately before acting, reserved/pending/session
paths filtered out of candidacy entirely, partial walks skipped for that
pass, and a circuit breaker that refuses a whole scope when more than
max(3, 50%) of its candidates look missing at once (or when a complete walk
finds nothing at all).
The tombstone is picked up by the SAME aged-tombstone hard-delete sweep decay
eviction uses, not exempted from it: what actually protects a false positive
is that `upsert_page`'s insert path now re-links a fresh write at a
tombstoned path onto that chain via `supersedes` and clears the old row's
`superseded_at`, turning it into an ordinary protected supersession-chain
member instead of an orphan the sweep would otherwise destroy. It never runs
the blocking admission gate, but does fire-and-forget any non-blocking
observer/mirror webhook. `sessions/*.md` pages are excluded from candidacy
entirely, since a same-workspace `move-session` re-home can leave one with a
correct row and no file (a separate, pre-existing bug, not fixed here). With
the flag off (the default), behavior is unchanged.
(cherry picked from commit 321e76c439)
Hermes runs a configured hook command through shlex.split with the event JSON on stdin and no shell, so its integration is exec-form (like Zero and ZCode) rather than a .sh bundle: AgentChoice::Hermes (alias hermes-agent) makes install-hooks, setup-agent and finalize-session accept it, and no script subdir is staged.
build_hermes_hooks_yaml emits the ready-to-paste hooks: block for ~/.hermes/config.yaml with the two events ai-memory can act on - pre_tool_call -> pre-tool-use and post_tool_call -> post-tool-use. Their tool_name / tool_input payload is the envelope the router already maps for agent=hermes, which is what gives Hermes sessions tool observations, tool-family titles and capture exclusions.
install-hooks --agent hermes never writes the config file: it is YAML the operator also edits and Hermes gates user hooks behind its own acceptance prompt (hooks_auto_accept), so the block is printed for pasting - the same choice made for Pool. Session lifecycle is deliberately left to the memory-provider plugin, so no hook-driven session-end is installed and a session cannot be closed twice.
Verified against Hermes v0.21.4: the block parses through Hermes own hermes_yaml, _parse_hooks_block registers both events, split_command_line yields an argv that runs the native hook subcommand, and that command spools the event with rc=0. Docs updated in docs/install.md, docs/support-matrix.md and docs/mcp-install.md. Follow-up to #623.
(cherry picked from commit 18155089cd)
A named profile (`--profile`, `OMP_PROFILE`, or the legacy `PI_PROFILE`) owns
`~/.omp/profiles/<name>/agent` and ignores `PI_CODING_AGENT_DIR`, as OMP does;
`install-hooks` wrote the extension into `PI_CODING_AGENT_DIR` instead. Profile
names are normalized and refused like OMP does, `PI_CONFIG_DIR` renames the
`~/.omp` root, and on Linux and macOS sessions move to `$XDG_DATA_HOME/omp`
when that directory exists. `install-mcp`, auto-wire, session import,
`backfill`, `doctor` and the `uninstall` sweep follow the same agent dir.
Refs #820
Bootstrap estimates tokens as bytes / 4 and filled --max-input-tokens
and --chunk-input-tokens to the last estimated token. That estimate
undercounts non-English text and source code (about 40% on Portuguese
mixed with code, as #884 measured for consolidation), so a chunk sized
to fit a model's context window could overflow it on input alone.
Fill 80% of each budget by the estimate, the default consolidation
uses for the same undercount, in both the prune and the chunk packing.
No flag or config key is added; a run may plan more chunks than before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hbGY9oRywtM3Nz2NcXxUF
build_chunk_request picked max_tokens from the chunk count: 16K when a
run had several chunks, 64K when it had one. A repository small enough
for a single chunk under default chunking therefore got the 64K cap
meant for --chunk-input-tokens 0, and a 64K-context model rejects that
request before reading any input.
Key the cap on the chunking mode instead: every call under chunking is
a chunk that fits the chunk budget and gets 16K; only a disabled
chunk budget (one call with the whole pruned bundle) keeps 64K. The
--max-input-tokens help also stops claiming that its 150K default
leaves room for 64K of output in a 200K window.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hbGY9oRywtM3Nz2NcXxUF
`ai-memory uninstall` left the `ai-memory run` auto-wire sentinels behind, so
the next managed launch skipped wiring and captured nothing. Removing hooks or
MCP now deletes every sentinel and lists them in the dry-run plan.
`--only mcp|instructions|skills` no longer deletes the stored hook bearer the
still-installed hooks read.
Refs #820
With turn checkpoints (#865), several live OpenCode sessions in one directory
each publish a baton. Each checkpoint retired the other live sessions'
batons, and SessionStart handed a new session the baton of a session still
in use, with another conversation's context.
A checkpoint now spares other open sessions' batons. Startup delivery
(startup_handoff) skips an open source that captured anything in the last
ten minutes, and the claim re-checks it in its own transaction. The claim's
sweep retires older batons of quiet open sessions but spares a busy one; an
explicit accept keeps sparing every open session.
`ai-memory run --env/--env-file` values reached the spawned harness and
native-session resolution but not first-launch auto-wire, so hooks and MCP
landed in the default config home, and a sentinel keyed only on agent and
version skipped a second account on the same version. Auto-wire now resolves
every install target from the launch environment, the Codex MCP entry follows
`CODEX_HOME`, and the sentinel also keys on the resolved hook and MCP paths.
Fixes found on the same paths: OMP profile and PI_CONFIG_DIR resolution,
blank relocation values, uninstall leaving sentinels behind, the Crush
context packet (default context files and the global crushrc), the Crush data
directory lookup, session discovery with concurrent launches in one checkout,
and Kiro v3 resume wiring. Spawned-server test fixtures now survive losing a
free port to another socket.
Refs #820
Post-merge hardening note per audit: the .sha256 is fetched from the same
base as the archive, so it only proves a consistent pair; an https base is
needed for transit integrity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
OpenCode 2.0.10 replaced the lifecycle events the generated plugin listened
for (session.idle and friends) with session.execution.{succeeded,failed,
interrupted}, and it runs one long-lived service behind every CLI: closing a
terminal is not a session end, so sessions captured through the old binding
rarely produced a summary page or a baton. The plugin also re-ran two `git`
processes per event, lost startup context after the first model request,
dropped content-only tool results, and shared queue and spool state across
the per-location instances OpenCode 2 loads.
Plugin (install_hooks.rs, render_shared.rs, the node host fixture):
- binds session.execution.*, session.text.ended, session.moved,
session.deleted, session.compaction.started and the context/prompt/tool
hooks; every name was checked against the OpenCode 2.0.14 binary;
- claims startup context once per root session and re-injects it on every
model request; child sessions never claim it or publish a baton;
- keeps queue, spool and cleanup state per location instance; spool names
can no longer collide within a millisecond, and a torn-down host cancels
what is still in flight only after its final session-ends had the drain
budget (all generated TypeScript integrations share this runtime now);
- opt-in assistant capture (--capture-assistant) hands the last completed
text to the native hook, which sanitizes and caps it before the spool or
the wire, and falls back to the plain stop hook if that binary cannot run.
Server (router.rs, ops.rs, reader.rs, writer.rs):
- a completed root turn is a turn checkpoint: sessions/<id>.md and the
automatic baton are refreshed deterministically (no LLM) in the session
row's own scope, keeping one open baton per live session, refreshed in
place and audited; a checkpoint that lost the race with the session's end
touches nothing;
- an explicit, keyed session.moved rebinds the live session row to its new
directory on first delivery (compare-and-set on the cwd it left), so the
session's end and checkpoints follow it;
- a SessionEnd whose resolved scope drifted under the same cwd (a
.ai-memory.toml appeared mid-session) ends the session instead of
stranding it open; scope-drifted ordinary events are recorded as upstream
records any other drifted event;
- the latest captured assistant excerpt rides in the automatic baton; it is
not rendered into the git-tracked session page.
Docs: install.md, support-matrix.md, auto-scope.md, SECURITY.md and
DATA_HANDLING.md describe the checkpoint, routing and capture behavior.
Verified: cargo fmt --all -- --check; cargo clippy --workspace --all-targets
-- -D warnings; cargo test --workspace --all-targets (3569 passed, on release/2.5); the node
host fixture runs under cargo test (node >= 22.6) and fails if unload aborts
deliveries before its session-ends. Live on Windows 11 with OpenCode 2.0.14:
after a 64 -> 66 store migration, the running OpenCode service loaded the
regenerated plugin and one real turn through it (a Code Mode call to memory_status)
recorded its prompt, tool events and stop, logged "turn checkpoint written;
native session remains open", left the session open, wrote
sessions/<id>.md and exactly one open baton carrying the captured
assistant excerpt, which the session page does not contain.
figment's Env provider parses each var's string with its loose-value
parser, whose bare/unquoted branch calls .trim() (figment 0.10.19's
src/value/parse.rs:78, verified against the vendored source) — so
AI_MEMORY_EMBEDDING_QUERY_PREFIX="query: " reached Config::embedding_query_prefix
as "query:", silently dropping the publisher-significant trailing space.
Config::load now overlays the raw (untrimmed) env value for the two prefix
keys via figment::providers::Serialized (which hands figment an
already-typed value, bypassing the string parser), merged after the
generic Env::prefixed pass so it still wins over a config.toml value.
Presence, not non-emptiness, is the signal: a variable set to "" is a
deliberate override clearing a config.toml-configured prefix, distinct
from the variable being absent. Both wrapper scripts (bin/ai-memory,
bin/ai-memory.ps1) now forward these two keys on presence for the same
reason; a non-empty override was already forwarded correctly, only the
empty-override case was silently dropped by the [ -n ] check every other
forwarded var correctly uses.
The overlay itself lives in overlay_embedding_prefixes(), a pure function
taking the env values as parameters rather than reading std::env::var
itself, so it stays directly unit-testable without mutating process
environment or the current directory — the same pattern
ai-memory-cli/src/commands/path_util.rs's agent_config_home and
ai-memory-hooks's drain_with_live_token already use. std::env::set_var is
unsafe under edition 2024 and forbidden workspace-wide, and
figment::Jail calls std::env::set_current_dir on the real process
internally, racing every other test in this crate's multi-threaded lib
test binary that relies on cwd. Four pure unit tests build a minimal
in-memory Figment and assert on the merged Config directly. One
additional test exercises the real Config::load end to end through a
TOML file and an explicit absolute path; since this process's own
environment is shared with every other test in the binary and could
already carry one of the two prefix vars from the test runner's shell,
that test re-execs this same test binary filtered to just itself as a
genuinely separate child process, with both vars removed via
Command::env_remove — real process isolation rather than an in-process
assumption about the ambient environment.
Plus: a wiremock transport test proving OpenAiEmbedder (not just
OpenAiCompatEmbedder) sends the configured prefix on the wire; a
regression test for the memory_query embed_query fix from the prior
commit, using a task-aware fixture embedder whose
embed/embed_document/embed_query methods return distinguishable vectors
so a regression back to calling the wrong one fails a direct equality
assertion rather than an inferred ranking change; fake-Docker argument
tests for the wrapper's presence-based forwarding across
unset/empty/whitespace-only/non-empty, with both env vars explicitly
removed from the child environment before each case so the "unset" case
cannot silently inherit an ambient export from the test runner.
bin/ai-memory.ps1 also notes, briefly, the PowerShell/.NET version an
operator needs for `$env:NAME = ''` to reach the wrapper as set-but-empty
rather than deleted — no pwsh runtime was available to exercise it live.
Also narrows the query/document-prefix docs (config.rs, docs/llm-providers.md,
CHANGELOG.md) with exact, separate templates: Nemotron-3-Embed
(nvidia/Nemotron-3-Embed-1B-BF16) and base E5 use "query: "/"passage: ";
e5-mistral-7b-instruct wants "Instruct: {task}\nQuery: " (trailing space);
Qwen3-Embedding wants "Instruct: {task}\nQuery:" (no trailing space) — both
leave documents plain.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds `embedding_query_prefix` / `embedding_document_prefix` config keys
(env: AI_MEMORY_EMBEDDING_QUERY_PREFIX / AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX)
applied by the openai and openai-compat embedders: an optional string
prepended to query/document text before the existing truncation, so
truncation still bounds the whole input. Asymmetric embedding models
(Nemotron-3-Embed, base E5, Qwen3-Embedding) need a "query: " /
"passage: " instruction their publisher specifies; the OpenAI-compatible
/v1/embeddings wire format has no field for it. Empty by default, and not
trimmed so a publisher's trailing space is preserved.
Also fixes memory_query's embed_query helper (ai-memory-mcp/src/server.rs),
found while auditing every embed/embed_query/embed_document call site: it
called the generic Embedder::embed instead of Embedder::embed_query, so an
asymmetric embedder (google's task-typed embeddings, or these new prefixes)
embedded the search query on the document side.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Codex/ChatGPT backend behind `openai-oauth` restricts the selectable
model set to a small server-defined list and rejects other ids
(including `gpt-5-mini`) with a deterministic 400. The TIP blocks in
`docs/llm-providers.md` and `docs/install.md` recommended setting
`AI_MEMORY_LLM_MODEL=gpt-5-mini`, which fails on that backend.
Advise leaving the provider default (`gpt-5.5`) for `openai-oauth`/`codex`,
keep `claude-haiku-4-5` for `anthropic-oauth`, and qualify `gpt-5-mini`
for `copilot` as unverified rather than asserting it works.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Agents without a true session-end event (Antigravity CLI, Kiro, ZCode,
Pool) keep capturing under the same session id after a manual finalize,
but the finalize-session discovery step only matched open sessions, so a
second finalize silently found nothing.
Add an opt-in --reopen flag (requires --session-id) that threads
include_ended=true through GET /admin/open-sessions for the exact-id
lookup only; the bulk listing still excludes ended sessions and
include_ended without session_id is a 400. The server-side session-end
path already distinguishes ReEnd from AlreadyEnded, so a re-run with
nothing new stays a harmless no-op while new observations get a fresh
summary page (via supersession), handoff, and opt-in consolidation.
The accessory app ships the ai-memory binary and hooks tree, starts the
existing LaunchAgent, and opens /web, status, config, and logs. Durable
data stays in Application Support so replacing the .app is an update.
Enhance CHANGELOG and installation documentation to include the new download size limit of 128 MiB for upgrade requests and clarify the trust boundary for the `release_base_url` configuration. This ensures users are informed about the implications of using custom release bases.
Introduce `release_base_url` in the configuration to allow users to specify a custom base URL for GitHub Releases during upgrades. This feature supports hermetic tests and mirrors, enhancing flexibility in deployment scenarios. Update documentation and tests to reflect this change. (#801)
Post-audit of release/2.3 (v2.2.2..HEAD) before the 2.3.0 release.
Fix (S1, the one code finding): install-hooks re-applied with no
--capture-assistant flag silently stripped an existing assistant-capture
opt-in — most visibly through the new `ai-memory run` auto-wire, which always
re-applies without the flag. install-hooks now preserves an already-baked
--capture-assistant on a bare apply (Claude Code + Codex, via the same
existing-config read that prompt-capture already uses); an explicit flag still
forces it on. Regression tests: bare re-apply preserves the opt-in; a fresh
install without the flag stays off. (The underlying non-preservation is also
latent on main and can be backported to a 2.2.x patch if desired.)
Docs (release-readiness):
- README: add the agent-messaging.md row to the Docs table (new feature + 4 MCP
tools shipped without a README entry); note run-as-preferred/auto-wire.
- ARCHITECTURE config reference: document capture_assistant, backfill_on_start,
run_autowire + their AI_MEMORY_* env overrides.
- install.md: add an "Upgrading to 2.3.0" note (the two default-on behaviors +
opt-outs) and a forward-only/backup-before-rollback note.
- design-boot-backfill.md: reconcile the now-answered "Open questions" with the
as-built implementation (path B, on-by-default, capped).
- design-rules-promotion.md: drop the stale "(2.1)" target label (untargeted,
still unimplemented).
- CHANGELOG: add the install-hooks preservation Fixed entry.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Extend the assistant-capture closed table with (Codex, Stop) → last_assistant_message
(verified on codex-cli 0.154.0), behind the existing double opt-in (client
--capture-assistant + server capture_assistant=true). The install gate
capture_assistant_allowed now admits Codex on a native hook platform (it
previously refused it), and the Codex render path threads --capture-assistant
onto the Stop command. The runtime read + sanitize/bound/backstop/store pipeline
is already agent-agnostic and keys off the closed table, so no new capture code.
Only (Codex, Stop) is added — Codex has no SubagentStop.
Tests: closed-table set updated to admit Codex Stop; a positive
client-transform capture test for Codex; the install-flag gate test admits
Codex; a render test asserting --capture-assistant bakes onto the Codex Stop
command (and only Stop) when opted in. Docs: support-matrix + install.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm