217 Commits
Author SHA1 Message Date
AkitaOnRails e2d7d17d90 fix(run): allow same-owner lease recovery
Refs #795
2026-10-01 16:18:44 -03:00
AkitaOnRails 62196091fa Merge PR #996: per-repository server profiles (fixes #992) 2026-09-30 15:55:27 -03:00
AkitaOnRailsandClaude Opus 4.8 0a7554c86f docs: fill 2.5.0 config reference gaps + fix stale Hermes/tool-count
Pre-release doc-staleness sweep for 2.5.0:
- ARCHITECTURE.md config reference: add the [maintenance] section
  (enabled/forget_sweep/lint/embedding_backfill intervals +
  reconcile_tombstones_deleted_pages, #929/#964) and release_base_url (#801).
- README support matrix: Hermes Agent Community -> Supported, aligning the
  compact table with the authoritative docs/support-matrix.md (#933 hooks +
  #942 skills are first-party).
- research-codebase-memory-mcp.md: correct the stale '19 tools' -> 23.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-30 15:44:55 -03:00
Talysson de Oliveira Cassiano a9e166f357 feat(cli): add server profile registry and ai-memory server (#992)
Named server profiles live in <data_dir>/servers.toml, with each token in
its own owner-only file under <data_dir>/auth-tokens/. A strict parser
refuses the whole registry on any unknown key, profile names are
validated so they can never become a path (or a Windows device), and
resolve() fails closed: unknown profile, missing token, outside roots,
or unrooted once several profiles exist.

`server add` keeps registered roots when --root is omitted and discards
a token whose profile moved to another URL, so rotating a token cannot
lift the roots restriction or pair one server's token with another.
`server list` never prints a token.

`uninstall` removes the stored profile tokens together with the hook
token, since the hooks it removes are their only readers, and keeps the
registry of URLs, so a later reinstall lists each profile as tokenless
rather than forgetting it.
2026-09-30 18:18:24 +00:00
AkitaOnRails 7e17939881 Merge PR #990: status discriminant on memory_handoff_accept (fixes #988) 2026-09-30 00:04:10 -03:00
Vinícius Andrade ea9c5e6750 feat(mcp): add status discriminant to memory_handoff_accept (fixes #988)
- Adds HandoffAcceptStatus enum ('claimed' | 'consumed_by_hook' | 'none_pending')
  returned alongside 'handoff' in memory_handoff_accept (#988, split from #920).
- Adds ReaderPool::handoff_claimed_by_live_session to verify if the caller's
  live session already received the handoff via SessionStart hook.
- Reports 'consumed_by_hook' only when the caller forwards its session id
  (e.g., Claude Code with --session-aware or OpenCode 2) and matches the
  scope, live session, and owner filters.
- Reports 'none_pending' when no handoff was pending or for static clients
  where the hook claim cannot be confirmed for that exact session.
- Updates documentation, routing skill, ARCHITECTURE.md, security boundaries (row 4g),
  and CHANGELOG.md.
2026-09-29 17:17:53 -03:00
AkitaOnRailsandClaude Opus 4.8 80888360e2 feat(run): warn before --yolo, offer ai-jail, add Claude true-yolo (#983)
ai-memory run --yolo now, on an interactive TTY only (never in hook/CI/
detached paths):

- warns that --yolo runs every tool call unconfirmed, [Y/n] default-yes;
- offers to re-run inside ai-jail when it is installed, re-execing the
  original argv under `ai-jail --network --agent-state --env <NAME>…`
  (--network shares the host net namespace so the loopback ai-memory server
  stays reachable; only already-set credential/config env is forwarded);
- skips both prompts when already inside ai-jail (Linux hostname ai-sandbox
  / macOS PS1 (jail) ; fails open to showing the warning; Windows never).

Opt-in Claude "true yolo" (--true-yolo / [config] claude_true_yolo,
AI_MEMORY_CLAUDE_TRUE_YOLO) silences the pauses --dangerously-skip-permissions
leaves: sets the CLAUDE_CODE_DISABLE_*_RM_* env vars and injects
--settings bypassPermissions. Claude-only, off by default, sandbox-first.

Detection and argv assembly are pure/OS-explicit (ai-memory-workstream::jail)
with unit tests for every branch; a CLI integration suite asserts the argv
against the real ai-jail via --dry-run (skips cleanly when ai-jail/bwrap are
absent). Design: docs/design-yolo-safety-ai-jail.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-29 15:24:14 -03:00
AkitaOnRails 8aeb422a9c Merge PR #921: repair-backfill-timestamps restores original session start/end
# Conflicts:
#	CHANGELOG.md
#	crates/ai-memory-store/src/lib.rs
#	docs/ARCHITECTURE.md
#	docs/lifecycle-ops.md
#	docs/security-boundaries.md
2026-09-29 12:37:40 -03:00
AkitaOnRails b3d8e99746 Reapply "feat(consolidate): add input_token_safety_margin to bound approximate token budget (#884)"
This reverts commit 0fc60ffc24.
2026-09-28 21:19:04 -03:00
AkitaOnRails 4829c14d43 Merge main into release/2.5
Forward-merge to restore main ⊆ release/2.5: brings the 2.4.x fixes onto the
2.5 feature line. Retains the 2.5 features main reverted for the 2.4.1 patch
(#884 input_token_safety_margin, #873 jev docs); #904 stays reverted as on main.

# Conflicts:
#	.github/workflows/ci.yml
#	CHANGELOG.md
#	crates/ai-memory-cli/src/commands/backfill.rs
#	crates/ai-memory-cli/src/commands/doctor.rs
#	crates/ai-memory-cli/src/commands/render_shared.rs
#	crates/ai-memory-cli/tests/suite/removal.rs
#	crates/ai-memory-hooks/src/router.rs
#	crates/ai-memory-mcp/src/server.rs
#	crates/ai-memory-store/src/lib.rs
#	crates/ai-memory-wiki/src/wiki.rs
#	docs/jev-reranker-adapter.md
#	docs/security-boundaries.md
2026-09-28 21:18:08 -03:00
AkitaOnRails 9d5c41e8ce Merge PR #925 into release/2.5
Route captures by repository identity, not folder name (#708)
2026-09-28 20:37:21 -03:00
AkitaOnRailsandClaude Opus 4.8 954548ffe1 fix(mcp): union _global for single-project queries, not just unscoped ones
memory_query gated the reserved _global preferences union on 'no named
workspace/project'. The routing doctrine tells static MCP clients to pass
workspace+project on every call, which set that gate false, so static clients
never received global_scope_hits despite the documented contract (#930). The
union now keys on single-project resolution (scopes empty); an explicit
multi-scopes set is the only opt-out, and global=true/as_of are unaffected. The
double-search guard now resolves the queried project (named or active) rather
than the active-project default. The union remains keyed strictly to the
reserved _global scope and never leaks another project's pages (adversarial
test + security-boundaries row 1b).

Closes #930

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-28 15:02:38 -03:00
luisfnicolauandClaude Opus 5.5 d9f8e2f5ec docs: changelog, marker file and boundary inventory for repository identity (#708)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aqFKAuGVkuBoewmpA3cx9
2026-09-28 09:43:31 -04:00
Maxsuel Einstein 7170c46762 feat(search): configurable FTS stopword list for non-English installs
Option A from #953 (proposed by @rntjr): [search.fts] stopwords config
key. Absent keeps the built-in English list byte-identical; [] disables
the filter; a non-empty list replaces the default outright. Case folds
with full Unicode case-folding (matching the unicode61 remove_diacritics
2 content tokenizer) but never strips diacritics, so an accented capital
still matches a configured lowercase entry while "e"/"é" stay distinct.

Threaded from Config::load through a new ReaderPool::set_fts_stopwords /
FtsStopwords, mirroring how [retrieval] already reaches ReaderPool.
Env override AI_MEMORY_SEARCH_FTS_STOPWORDS (CSV, matching
allowed_hosts/cors_allow_origins/trusted_proxy_cidrs) is applied as data
in Config::load rather than through the usual __-split figment Env
layer, since that layer merges raw values before any deserializer could
tell a present-but-blank var (must mean "unset") apart from a real
empty-list override.

Closes #953.
2026-09-28 09:31:00 -03:00
AkitaOnRails 6add9b727b Merge PR #927 into release/2.5
feat(cli): reclaim-ledger-versions drops the pre-#660 ledger residue (#914)

# Conflicts:
#	docs/ARCHITECTURE.md
2026-09-27 12:52:54 -03:00
Felipe Neves RicardoandClaude Opus 5.5 7fd00eb766 fix(consolidate): skip multi-page updates to pinned pages (#934)
A multi-page batch writes each update at a path the model chose. When
that path named an existing pinned page, apply_batch replaced its body
and wrote the new version with pinned = 0, because neither the batch
loop nor the wiki write path looks at the page being replaced. Pinned
pages are documented as immutable to automation.

Skip such updates with a warning, next to the existing invariant-slot
skip. _slots/ pages are pinned automatically and keep their own
state/invariant regime, so they are not skipped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012hbGY9oRywtM3Nz2NcXxUF
2026-09-26 22:09:27 -03:00
Yi-111-a d2b9896876 feat(cli): reclaim-ledger-versions drops the pre-#660 ledger residue
#660 stopped the indexer from superseding an OKF-conformed ledger on every
hook append, but the rows it already wrote stay and nothing removed them:
`compact` deletes nothing, `forget-sweep` only hard-deletes decay
tombstones (only `decay` writes `superseded_at`), and `reindex` loses the
DB-only state. One reported store holds 6,539 versions of 101 live pages,
95% of its rows.

Adds `ai-memory reclaim-ledger-versions` (`POST
/admin/reclaim-ledger-versions`), an online command that deletes only the
superseded versions of paths whose *content* opens with a hook log entry.
Dry run unless `--confirm`; `--drop-latest` also removes each ledger's live
row; `--compact` rebuilds the FTS index and VACUUMs to return the bytes.

The content gate is the load-bearing part and is now one definition:
`ai_memory_core::log_ledger` owns the predicates, `ai_memory_wiki::ledger`
keeps the file-reading half and delegates, and the store decides from a
`pages.body` with no filesystem access. A real page a human named
`log-2026-09.md` keeps its whole version chain. Only `is_latest=0 AND
superseded_at IS NULL` rows are eligible, so decay-owned rows stay with the
sweep that reasons about them.

Derived FTS/entity/vector/link rows go with the page through the existing
ON DELETE CASCADE. The FTS delete trigger is stood down for the bulk delete —
its DDL read back from sqlite_master and re-executed, so it cannot drift
from the schema — and pages_fts is rebuilt wholesale, which is what keeps
this from re-tokenizing tens of gigabytes of ledger body row by row.

Reading the gate needs a bounded prefix, not the whole body: GATE_PREFIX_BYTES
lives with the predicate, shared by the file reader and the store.

Refs #914
2026-09-26 14:16:04 +08:00
Maxsuel Einstein 25fabe4a13 feat(cli): repair-backfill-timestamps restores original start/end of backfilled sessions
Sessions imported by `backfill` before it carried per-event timestamps all
have started_at/ended_at on the import day, which flattens every
time-ordered view of that history. This adds an admin repair for sessions
already imported; it is independent of the companion fix to `backfill`
itself (#919).

The repair runs online, through the server's single writer, because it does
not need direct disk access (invariant 9). POST /admin/repair-session-times
validates a caller-supplied batch of (session_id, started_at, ended_at)
candidates in one transaction, scoped to (workspace, project) like
purge-session. It only rewrites a row whose current started_at postdates
the candidate's own transcript end, the bug's signature, so a correctly
hook-captured session or a re-run is a no-op (`not_flattened` /
`unchanged`). Without confirm it is a real dry run: the write runs and
rolls back. Out-of-scope sessions are reported not-found and left
untouched; negative, inverted or future times are refused; an open session
never gets an end time; and a confirmed batch writes one audit_log row with
every repaired session's before/after times.

`ai-memory repair-backfill-timestamps` is the thin client. It reuses
backfill's session discovery and the workstream transcript reader, takes
each session's earliest and latest event time, resolves native ids with the
new `SessionId::from_native` (now shared with the hook router instead of a
private copy), chunks the batch at the server's cap, and prints each
repaired session's old and new times.
2026-09-25 20:51:40 -03:00
luisfnicolauandClaude Opus 5.5 9da6f52e37 docs: changelog, users guide and design status for #708 slice 3
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aqFKAuGVkuBoewmpA3cx9
2026-09-25 18:51:09 -04:00
AkitaOnRails 0fc60ffc24 Revert "feat(consolidate): add input_token_safety_margin to bound approximate token budget (#884)"
This reverts commit 7ff429935f.
2026-09-25 17:56:18 -03:00
AkitaOnRailsandClaude Opus 4.8 dde5806b23 Merge PR #912 into release/2.5
feat(web): add a root-only /web/pending triage page (#855)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-25 17:29:50 -03:00
Maxsuel Einstein 191a8eea11 fix(backfill): store each event's original time instead of the import time
Every session/observation `ai-memory backfill` imported was stamped with
started_at/ended_at/created_at at import time, discarding each transcript
event's own timestamp even though the workstream adapters already captured
it (`NewWorkstreamEvent::occurred_at`) — a session imported today from a
transcript recorded weeks ago showed up as having happened today.

`NewSession` and `NewObservation` gain an optional `occurred_at`
(microseconds); the store falls back to `now()` when it is `None`, so live
hook capture is unaffected. This includes `admit_hook_session_event`'s own
session-row INSERT (the real path a live `/hook` session-start event takes,
separate from `begin_session_row`) — missing that one meant a backfilled
session's `ended_at` (original, past) could sit before its `started_at`
(import time).

The `/hook` body accepts an RFC 3339 `occurred_at`, read from the top level
of the body only (not the nested `payload`/`event`/`properties`/`info`/`path`
search other hook fields use, so a harness payload that happens to carry an
`occurred_at` key elsewhere in its own structure is never mistaken for this
field). It is client-controlled input arriving over the hook endpoint, so
`HookEnvelope::occurred_at_micros` bounds it (must be > 0 and no more than
five minutes ahead of server time) before it is trusted; anything else —
missing, unparsable, or out of bounds — resolves to `None` rather than
erroring, keeping hooks fire-and-forget. It is numeric metadata, not text, so
it never goes through the sanitizer.

Backfill validates each transcript event's own timestamp (an unparsable one
is treated as missing) and threads the resolved time through `map_event` and
into the session-start/session-end items: a missing timestamp inherits the
nearest preceding valid one, an event before the first valid timestamp
inherits that first one, and the session's start/end times are the earliest/
latest valid event time in the transcript rather than assuming it is already
time-sorted.

Because a backfilled session's `ended_at` can land well in the past, it can
sit below the auto-improve review watermark and the cross-session
experience-pass anchor (both keyed on `ended_at`), so a freshly imported
session may not get an automatic review pass until a newer session moves
those forward; an opt-in retention window measured from an observation's own
time can also make an old backfilled observation immediately prunable rather
than only after it ages in place; and the "most recently active project"
restart fallback, which looks at how recent observations are, may not pick a
project that was just backfilled. These are documented consequences, not
regressions introduced here — they follow directly from timestamps now being
honest.
2026-09-25 15:27:03 -03:00
Jayson ReisandClaude Opus 5.5 a84d66be60 feat(web): add a root-only /web/pending triage page (#855)
Add a server-rendered page under /web that lists the pending
auto-improvement proposals of all projects. The page has a project
filter and a sort. It shows the rationale, the proposed body, and a
warning for proposals that write under _rules/ or replace a page.

The approve and reject buttons post from the browser to the existing
/admin/pending-writes/{id}/approve|reject handlers with the session
cookie and the CSRF header. Admission, audit, attribution, and the
single writer stay the same. ai-memory-web adds no write route.

The page uses the same Capability::Admin decision as /admin, and
serve passes the trusted-proxy setting to it. A non-root session gets
a 403 page. The HTML auth redirect does not send that session to the
change-password form.

Two store reads feed the page. One counts pending proposals per
project, so the total and the project filter include every project.
The other lists proposals, optionally for one project, with a cap of
500 rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SwU4vj4tZLYVLR7vEKeh2W
2026-09-25 18:11:02 +02:00
Éverton Toffanetto 407d55d0f8 Merge remote-tracking branch 'upstream/release/2.5' into fix/autowire-run-env 2026-09-25 00:44:04 -03:00
AkitaOnRails 8f55486e5f Merge branch 'main' into release/2.5
Forward-merge the 2.4.x fix batch (#877,#891,#893,#888,#820,#885,#886,#884,#890,#894,#895) so main stays a subset of release/2.5.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

# Conflicts:
#	CHANGELOG.md
#	crates/ai-memory-cli/src/commands/run.rs
#	crates/ai-memory-store/tests/suite/multi_session.rs
2026-09-25 00:22:09 -03:00
AkitaOnRails fe5b6ca789 Merge fix/885-886-884-consolidation into main
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

# Conflicts:
#	CHANGELOG.md
2026-09-25 00:07:00 -03:00
AkitaOnRails ff3b515947 Merge fix/895-894-synth-log into main
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

# Conflicts:
#	CHANGELOG.md
2026-09-25 00:05:53 -03:00
AkitaOnRailsandClaude Opus 4.8 7ff429935f feat(consolidate): add input_token_safety_margin to bound approximate token budget (#884)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-25 00:03:30 -03:00
AkitaOnRailsandClaude Opus 4.8 1aad7a5d4b fix(logging): quiet the 30s reconcile summary and the rmcp SDK in the default log (#894)
Two routine INFO lines made up most of the default server log:

- the wiki watcher's "reconciliation pass complete" summary fired every
  RECONCILE_INTERVAL (30s) regardless of activity; drop it to debug. Its
  failure signals stay loud (the per-page warn!, the watcher_degraded
  error!, and the info! recovery transition).
- the external rmcp MCP SDK logs per-request lifecycle at info, which the
  default filter did not cap. Prepend `rmcp=warn` to the default filter so
  an operator can restore it via log_level (e.g. "info,rmcp=info") or
  RUST_LOG, while `tracing_appender=warn` stays appended and thus
  non-overridable — the invariant #15 feedback-loop guard.

Extract the filter string into a pure `default_filter` helper and unit-test
the ordering: the default suppresses both targets, log_level can restore
rmcp, and log_level cannot lower tracing_appender below warn.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-24 23:22:57 -03:00
AkitaOnRailsandClaude Opus 4.8 18e7936ca2 fix(consolidate): reconcile the session job row after a manual memory_consolidate (#890)
A manual memory_consolidate MCP call wrote the page directly through the
consolidator and never touched session_consolidation_jobs, so a session
whose automatic SessionEnd job had reached the terminal `failed` state
(which claim_next never re-picks) kept showing `failed` even though the
operator had just consolidated it — a two-sources-of-truth inconsistency.

Add a writer method reconcile_session_consolidation_completed(session_id)
that flips a `failed`/`pending`/`superseded` row for the session to
`completed`. It deliberately never touches a `running` row: that is an
active worker lease, and stomping it would violate invariants #2/#16, so
the automatic path's claim-guarded complete/fail stays the only lease
transition. The memory_consolidate handler calls it best-effort after a
successful, non-dry consolidate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-24 23:11:27 -03:00
Éverton Toffanetto 7b20d1d734 Merge remote-tracking branch 'upstream/release/2.5' into fix/autowire-run-env
# Conflicts:
#	CHANGELOG.md
2026-09-24 22:40:14 -03:00
Éverton Toffanetto 6295dd814b fix(run): wire hooks and MCP where the launch environment points
`ai-memory run --env/--env-file` values reached the spawned harness and
native-session resolution but not first-launch auto-wire, so hooks and MCP
landed in the default config home, and a sentinel keyed only on agent and
version skipped a second account on the same version. Auto-wire now resolves
every install target from the launch environment, the Codex MCP entry follows
`CODEX_HOME`, and the sentinel also keys on the resolved hook and MCP paths.

Fixes found on the same paths: OMP profile and PI_CONFIG_DIR resolution,
blank relocation values, uninstall leaving sentinels behind, the Crush
context packet (default context files and the global crushrc), the Crush data
directory lookup, session discovery with concurrent launches in one checkout,
and Kiro v3 resume wiring. Spawned-server test fixtures now survive losing a
free port to another socket.

Refs #820
2026-09-24 19:03:13 -03:00
Test 25b3481bbf Merge upstream/release/2.5 into issue/801-native-cli-upgrade. 2026-09-23 21:32:07 +03:00
AkitaOnRailsandClaude Opus 4.8 115895cbe1 feat(lint): make the A5 contradiction-similarity band configurable
The zero-LLM contradiction detector in memory_lint used a hardcoded
absolute cosine band (0.4-0.75). On a single-language or single-domain
store the background similarity floor is elevated, so the fixed band
admits many same-domain-but-unrelated pairs as noise, up to the 25-finding
cap.

Add contradiction_band_min / contradiction_band_max config keys (env:
AI_MEMORY_CONTRADICTION_BAND_MIN / _MAX), defaulting to the historical
0.4 / 0.75 so existing installs see no behavior change. Validated at
config load (finite, 0.0 <= min < max <= 1.0). Threaded through
LintOptions and into the three call sites that build it (admin HTTP
lint, the memory_lint MCP tool, and the scheduled lint tick).

Closes #853.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-23 14:58:04 -03:00
Ruben Licio Reis cf90793569 refactor(upgrade): improve code formatting and enhance readability
- Reformatted code in `upgrade.rs` for better readability, including consistent indentation and line breaks.
- Simplified the `refuse_oversized_content_length` function signature for clarity.
- Updated test cases for improved formatting and consistency in assertions.
- Adjusted the `ARCHITECTURE.md` documentation to reflect the new order of commands.
2026-09-23 20:52:47 +03:00
AkitaOnRails 7552d4fd39 Merge branch 'pr-859' into release/2.5
# Conflicts:
#	CHANGELOG.md
2026-09-23 14:17:53 -03:00
seathatflowsinourveinsandClaude Sonnet 5 0fff203e8c feat(embedding): add query/document prefix support for openai and openai-compat embedders
Adds `embedding_query_prefix` / `embedding_document_prefix` config keys
(env: AI_MEMORY_EMBEDDING_QUERY_PREFIX / AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX)
applied by the openai and openai-compat embedders: an optional string
prepended to query/document text before the existing truncation, so
truncation still bounds the whole input. Asymmetric embedding models
(Nemotron-3-Embed, base E5, Qwen3-Embedding) need a "query: " /
"passage: " instruction their publisher specifies; the OpenAI-compatible
/v1/embeddings wire format has no field for it. Empty by default, and not
trimmed so a publisher's trailing space is preserved.

Also fixes memory_query's embed_query helper (ai-memory-mcp/src/server.rs),
found while auditing every embed/embed_query/embed_document call site: it
called the generic Embedder::embed instead of Embedder::embed_query, so an
asymmetric embedder (google's task-typed embeddings, or these new prefixes)
embedded the search query on the document side.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 12:00:54 -04:00
AkitaOnRails e3b2de1bfb Merge branch 'pr-822' into release/2.5
# Conflicts:
#	CHANGELOG.md
2026-09-22 15:02:30 -03:00
AkitaOnRails 35086354d9 Merge branch 'pr-837' into release/2.5
# Conflicts:
#	CHANGELOG.md
2026-09-22 15:01:33 -03:00
AkitaOnRailsandClaude Opus 4.8 00aa6ee863 fix(serve): enable TCP keepalive on accepted connections (#792)
ai-memory serve leaked one file descriptor per dead hook/MCP peer.
A client whose connection dies without sending FIN (laptop sleep, a
VPN/Tailscale flap, an abrupt kill) leaves the accepted socket
ESTABLISHED forever, since the OS keepalive default is off. Over ~2-3
days of normal churn that exhausts the 1024-fd default and breaks the
healthcheck -- an unauthenticated availability/DoS. The rmcp
session-table half of this leak was already fixed in 2.4.0 by the
rmcp 2.x bump; this closes the remaining half-open-socket half.

Accepted sockets now get TCP keepalive via socket2, wired through
axum::serve::ListenerExt::tap_io (a hand-rolled axum::serve::Listener
newtype was tried first, but into_make_service_with_connect_info's
Connected<IncomingStream<'_, L>> bound is only implemented by axum for
its own TcpListener and, generically, for TapIo<L, F> -- never for an
arbitrary third-party L, and the orphan rule blocks implementing it
ourselves since neither Connected, SocketAddr, nor IncomingStream is
local to this crate). tap_io keeps the real peer SocketAddr flowing to
ConnectInfo while still touching every accepted stream.

New tcp_keepalive_secs config field (default 60s; AI_MEMORY_TCP_KEEPALIVE_SECS
env override; 0 disables keepalive entirely), read once through the
existing Config::load path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-22 14:48:48 -03:00
Abhishek Sharma bdff55c568 feat(auto-improve): make the reviewer's readable folders configurable
`load_patchable_pages` hardcoded `_rules/` and `procedures/` as the only folders
whose page bodies reach the reviewer. Every other page arrives through
`render_recent_pages` as one line — path, title, kind, updated_at — with no
content, so the model cannot tell that what it is about to propose is already
written down.

For a project whose durable knowledge lives in `decisions/` or `gotchas/`, with
no `_rules/` pages at all, the filter returns an empty vector and the reviewer
sees no page body whatsoever. It then re-proposes invariants that already exist,
at high confidence, because the claim really is well evidenced by the session —
confidence tracks the evidence, and could never track "the wiki already says
this".

The workaround is to move normative content into `_rules/`. That is good hygiene
independently, but it makes folder choice silently decide model visibility, and
the wiki's own conventions encourage `decisions/` and `gotchas/`.

`patchable_page_prefixes` defaults to the historical pair, so every existing
config behaves exactly as before; a project can add the folders where its
knowledge actually lives instead of restructuring its wiki.

Two boundaries are deliberate and tested. An empty list sends no page bodies —
"no folders are patchable" must not mean "every folder is". An empty string in
the list is ignored, because `"".starts_with` matches every path and would turn
the filter off silently. Prefixes carry their trailing slash, so `decisions/`
does not also match `decisions-archive/`.

This is the narrowest of the three options in #834. It does not change the
retrieval strategy (embedding-nearest dedup context) or the recency ordering
that lets `sessions/` churn crowd the recent-page list; both are noted in the
docs as still open.

3382 workspace tests pass, clippy -D warnings and fmt clean. Mutation-tested:
adding `decisions/` to the default fails two tests, and letting an empty prefix
match fails another.

Closes #834
2026-09-21 15:06:18 -07:00
Éverton Toffanetto 32b320825e feat(hooks): support external lifecycle capture ownership 2026-09-20 23:27:45 -03:00
AkitaOnRailsandClaude Opus 4.8 53eee9b43d feat(consolidate): B2/B3/B4 opt-in LLM "dream" pass (off by default)
Add the LLM "dream" pass: cross-session rewrite/merge of cold clusters,
scheduled on idle with cancel-on-activity, surprisal-first ordering
(docs/design-memory-aging.md B2/B3/B4).

B2 — cross-session LLM merge. `dream::run_dream_pass` clusters the SAME
bounded cold set the forget sweep materialises (reuses A3's adaptive_eps /
dbscan; new pub(crate) `sweep::materialize_cold_set` shares the scoring,
invariant #2) and hands each multi-member cluster to the provider to rewrite
into ONE coherent page via JSON-schema structured output (invariant #7,
`DreamMergedPage`). Routes through the gated apply path
(`preflight_admission(Consolidate)` before the LLM, then `Wiki::apply_batch`)
with dry_run first. Never deletes a source (invariant #16): the survivor is
rewritten, every merged-away member is superseded with a merge-note stub
(reachable via git + supersession chain), and `page_evidence`
(reconsolidation + `b2_dream:<id>`) records the sources. No provider / no
embedder / flag off ⇒ a clean no-op; the zero-LLM A3 path is untouched
(invariant #13). No migration.

B3 — event + idle scheduling with cancel-on-activity. New `[dream]` config
and a scheduled job in serve.rs that runs the pass only after a configurable
idle window (read from the MCP tool router's shared `ActivityClock`, bumped on
every tool call) and cancels it the moment activity resumes (a cheap
`DreamCancel` polled between clusters; a watcher flips it on renewed activity).
Bounded to max_clusters_per_run per run (invariant #5), all writes via the
single-writer actor (invariant #2). Every run yields an observable
`DreamReport` — no silent window.

B4 — surprisal-first ordering. The work queue is ordered by each cluster's
embedding distance to the nearest existing (non-cold, latest) page —
most-novel first. Pure ordering, unit-tested.

Opt-in LLM, OFF by default, R2-gated before default-on. Tests use a fake
LlmProvider (zero live calls). CHANGELOG + ARCHITECTURE updated;
comparison.md / competitive-parity.md left for the closing docs pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 14:24:05 -03:00
AkitaOnRailsandClaude Opus 4.8 e9b9d9cfe2 feat(store): B1 belief-strength confidence into ranking authority (default-OFF)
Turn the dormant page_evidence substrate (V63) into a read-time, zero-LLM
belief-strength `confidence` that can shade retrieval ranking as ONE bounded
factor — exposed inertly by default, foldable into authority only behind a
default-OFF, R2-gated flag (design-memory-aging.md B1 /
design-hindsight-borrowings.md §3).

- New pure `ai_memory_store::belief` module (mirrors decay.rs): confidence is a
  bounded [0.0, 0.95] function of distinct supporting sessions (breadth, not
  raw count), recency of the newest sighting, and live `contradicts` count.
  Anti-entrenchment baked in: distinct-session weighting, a residual cap so
  non-session volume can't dominate, recency shading to a floor, and a hard
  confidence cap so evidence alone never pins authority at the ceiling.
- New batched reader `page_belief_inputs` (one evidence-aggregate query + one
  live-contradiction query, never per-hit; invariant #2). Fetched only when an
  explained query needs it or the weight is on, so the hot default path is
  untouched. Replaces the now-subsumed `page_evidence_counts`.
- SearchExplain gains `confidence` and `belief_factor` beside `evidence_count`;
  all populated on the explained path (inert). memory_status gains
  `evidence_rows`.
- Folded into PageAuthority inside the existing [0.55, 1.50] clamp — one more
  bounded factor, no new multiplier tower — behind
  `[retrieval] belief_authority_weight` (default 0.0 = OFF, byte-identical
  ranking). A supersession always wins regardless of evidence: a superseded
  version's stale evidence is never boosted, and confidence never gates whether
  a write/correction takes (invariant #16).

R2: the authority factor ships OFF; enabling it is gated on a positive R2 delta
(retrieval triple / QA), not yet performed.

Tests: pure monotonicity/bound/cap/recency/contradiction unit tests; default-OFF
byte-identity proof; ranking-on boost; inert explain/status exposure; live
contradiction weakening; supersession-always-wins. Full gate green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 13:49:23 -03:00
AkitaOnRailsandClaude Opus 4.8 4bf20d28cc feat(lint): A5 zero-LLM contradiction detection via cosine-similarity band
Flag likely-conflicting pages through the existing memory_lint. Cold
semantic/procedural pages whose already-stored embeddings sit in the
0.4-0.75 cosine-similarity band ("same topic, not a near-duplicate" -- the
shape of a likely contradiction; >=0.75 is A3 dedup territory, <0.4
unrelated) get an advisory `contradiction` finding naming both pages with
newer-wins timestamp advice.

Zero generative LLM (invariant #13): reads only existing embeddings + cosine.
No embedder configured, or no embeddings for the (provider, model, dim)
triple, is a clean no-op -- never an error, never a provider call. Advisory
only (invariant #16): emits a finding, never deletes/edits/supersedes a page.

No persistence, no migration: `contradicts` edges live in the body-derived
`links` table (replace_links_in_tx rewrites it from the page body on every
write), so a programmatic edge would be silently wiped -- hence a lint-time
detector, not a stored edge. Bounded (invariant #2): one embeddings load
over the capped cold set, capped findings, deterministic ordering. No new
MCP tool (still 23).

Runs on the user-invoked memory_lint (MCP + admin) off the configured
embedder's triple; the automatic scheduled lint stays rule-based.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 13:06:22 -03:00
AkitaOnRailsandClaude Opus 4.8 3c227ade1f feat(consolidate): A3 cold-cluster dedup + A4 entropy pre-filter (opt-in)
Two independent, opt-in, non-destructive "consolidation hygiene" features
(docs/design-memory-aging.md buckets A3 + A4). Both off by default, zero
generative LLM, and reuse existing tables (no new migration).

A3 — cold-cluster dedup. The forget-sweep can cluster near-duplicate cold
episodic pages by embedding (cosine DBSCAN, adaptive k-distance eps clamped
to a conservative ceiling, minPts=2) over the bounded cold set it already
materialises, and collapse each cluster to one survivor: the highest-retention
member, its body the extractive union of the cluster's keep-tokens (every
member's durable facts survive), with the other members superseded by a merge
note pointing at the survivor. No source is hard-deleted (invariant #16 — merged
members stay reachable via supersession + git, recoverable with restore-page);
merge provenance is recorded in page_evidence (reconsolidation kind, a3_merge:
source-id prefix — the closed vocabulary has no A3 kind and adding one would need
a CHECK-constraint migration). Reads only already-stored embeddings, so it is a
clean no-op with no embedder or no matching (provider,model,dim) triple
(invariant #13). New pure module cold_cluster.rs (DBSCAN + adaptive eps). Gated
by [decay] dedup_cold_clusters (default false), wired into the scheduled sweep;
results surface in SweepReport.merged.

A4 — entropy / boilerplate pre-filter. A pure, zero-LLM Shannon-entropy +
boilerplate gate (entropy_filter.rs) that skips low-information session pages
(near-empty, whitespace, single-character, highly-repetitive) from the
cross-session experience consolidation pass before they reach the prompt, eval
gate, or apply_batch. Advisory (skip, never delete; invariant #16), off by
default via [auto_improve.scheduler.experience_entropy_filter] with conservative
validated thresholds tuned so a terse-but-informative note is KEPT.

R2: both ship opt-in/off, so R2 is not a merge blocker; the recall
no-regression proof is the gate before any future default-on. No R2 number was
run for this change. No new MCP tool (still 23).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 11:35:26 -03:00
AkitaOnRailsandClaude Opus 4.8 ec18487a5f feat(decay): extractive tier-down of cold episodic pages (A2)
Add opt-in extractive tier-down (docs/design-memory-aging.md §A2): when
`[decay] compact_cold_episodic` is on, the forget-sweep COMPACTS a cold
episodic page instead of evicting it — keeping the L0 frontmatter
`abstract:`, an L1 first-paragraph summary, and an L2 regex-mined
keep-token set (paths, URLs, code spans, error codes, UPPER_SNAKE
constants, long identifiers), dropping the prose body. Tier-down beats
eviction because the durable facts survive.

Off by default, so an upgrade changes nothing. Fully zero-LLM (regex
only, #13). Reversible and non-destructive: the rewrite goes through the
wiki layer (write_page/apply_batch), so the full pre-compaction body
stays in git history and the supersession chain and is recoverable with
restore-page (#16). The compaction pass runs BEFORE the decay-eviction
pass and batches its writes (#2).

New V65 migration adds a nullable `pages.compacted_at` marker: additive
`ADD COLUMN` only, no backfill — populated lazily by the sweep, derived
at the single upsert choke point from a `compacted: true` frontmatter
mirror so the marker and body commit in one transaction (#3, #10). The
sweep and curator skip a marked page, so it is never re-compacted,
re-evicted, or re-reported as cold. Downgrade guard preserved
(set_abort_missing(true)).

Ships opt-in/off; the R2 recall no-regression proof is the gate before
any future default-on. No new MCP tool (still 23).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 10:55:42 -03:00
AkitaOnRailsandClaude Opus 4.8 6f664e86ae feat(decay): per-tier retention half-life curves (A1)
Replace the single global decay rate λ with an optional per-Tier
half-life lookup, config-driven, so an operator can keep episodic
history longer and working-tier scratch shorter (the
mcp-memory-service 365/180/90/30 shape) instead of one λ for every
tier.

The default is byte-identical to today: `DecayParams::tier_lambda`
is an empty `TierLambdas` (a Copy struct of four `Option<f64>`, not
a HashMap, so `DecayParams` stays `Copy`). `lambda_for(tier)` returns
the scalar `lambda` unchanged for any unset tier — no days↔λ
round-trip that could perturb an unconfigured store's scores — so an
upgrade changes no score and mass-evicts nothing on the first
post-upgrade forget-sweep. The load-bearing property test asserts
byte-for-byte (to_bits) equality with the historical scalar-λ formula
across a grid of inputs for every tier.

`retention_score`/`retention_score_with_breadth` now take the page's
tier; the two retention-scoring callers (sweep, curator) pass
`c.tier`. Config exposes an opt-in `[decay.half_life_days]` table in
DAYS (converted to λ = ln(2)/days), wired through `DecaySettings`,
`decay_params()`, and load-time validation (each override finite and
> 0). Pure math + config: no new column, no migration, no new MCP
tool (still 23).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 02:31:17 -03:00
AkitaOnRailsandClaude Opus 4.8 5bac224a58 feat(mcp): reinforce retention on read_page, related-walk, and explore (C1)
Close the access-reinforcement gap so memory surfaced through more MCP read
paths resists decay. Previously only memory_query and memory_recent bumped
pages.access_count + last_accessed_at (the M8 decay reinforcement term);
memory_read_page, its include_related graph walk, and memory_explore
reinforced nothing.

Wire the existing sanctioned spawn_access_bump into all three:
- memory_read_page bumps the resolved page (via latest_page_id_by_ids), and,
  when include_related is set, the walked neighbours too — fired only on the
  success paths so an errored read reinforces nothing.
- the related walk now carries each node's PageId (serde-skipped, so the wire
  payload is byte-identical) for the walk-bump.
- memory_explore resolves the surfaced snapshot pages (rules, slots, recent,
  pinned, settled) to ids in one batched read (latest_page_ids_by_paths, no
  per-page N+1) and bumps them.

Strictly additive: reuses the throttled (<=1 per (page,operator) per 60s),
FTS-exempt, fire-and-forget single-writer path, so invariants #2/#3/#13/#16
hold. Response payloads unchanged; no new MCP tool (still 23).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-20 01:58:30 -03:00
AkitaOnRailsandClaude Opus 4.8 78eda74826 Merge remote-tracking branch 'origin/release/2.4' into fix/757-scope-observability
Rebase-via-merge of PR #774 (memory_status scope reporting) onto the
current release/2.4. Integrated the traced scope-resolution path
(resolve_read_args_traced / ScopeSource / ReadPointer) with the
error-classification and reasoning/pinned/answer features already on
release/2.4. Resolved conflicts in CHANGELOG.md (kept release/2.4's
Added/Changed lists and appended #774's memory_status entry, ref #757,
#774) and docs/ARCHITECTURE.md (kept #774's memory_status row plus
release/2.4's memory_briefing/memory_explore rows).

docs(changelog): reference #774 for memory_status scope reporting

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-19 16:36:18 -03:00