memory_query gated the reserved _global preferences union on 'no named
workspace/project'. The routing doctrine tells static MCP clients to pass
workspace+project on every call, which set that gate false, so static clients
never received global_scope_hits despite the documented contract (#930). The
union now keys on single-project resolution (scopes empty); an explicit
multi-scopes set is the only opt-out, and global=true/as_of are unaffected. The
double-search guard now resolves the queried project (named or active) rather
than the active-project default. The union remains keyed strictly to the
reserved _global scope and never leaks another project's pages (adversarial
test + security-boundaries row 1b).
Closes#930
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Grok Build CLI posts Claude Code's snake_case tool fields (tool_name /
tool_input / tool_use_id), but AgentKind::Grok was absent from both
closed_tool_agent (payload.rs) and the tool_observation_metadata agent
match (capture_policy.rs), so tool body extraction returned None and
every Grok PostToolUse observation was stored with an empty body while
still being captured. Grok now shares the Claude Code tool mapping.
Closes#931
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Fixes-only patch on the 2.4.x line; additive items (input_token_safety_margin,
HTTP-exposure status field, status --workspace/--project scoping, jev adapter
docs) are reverted here and live on the release/2.5 feature line to keep the
2.5.0 version reserved for it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Appending the extension to an empty final component fabricated `.md` from a
missing path, which slipped past auto-improve's empty-path rejection and
surfaced as `unsupported_path_prefix` instead of `invalid_path`. Skip the
append when the final component is empty so a missing/invalid path stays empty
and is rejected upstream as before.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Two routine INFO lines made up most of the default server log:
- the wiki watcher's "reconciliation pass complete" summary fired every
RECONCILE_INTERVAL (30s) regardless of activity; drop it to debug. Its
failure signals stay loud (the per-page warn!, the watcher_degraded
error!, and the info! recovery transition).
- the external rmcp MCP SDK logs per-request lifecycle at info, which the
default filter did not cap. Prepend `rmcp=warn` to the default filter so
an operator can restore it via log_level (e.g. "info,rmcp=info") or
RUST_LOG, while `tracing_appender=warn` stays appended and thus
non-overridable — the invariant #15 feedback-loop guard.
Extract the filter string into a pure `default_filter` helper and unit-test
the ordering: the default suppresses both targets, log_level can restore
rmcp, and log_level cannot lower tracing_appender below warn.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Every closed-tool agent stores a call's title as `tool file` / `tool
non-file` / … (a partition of the calls, not a tool name), while the
reserved-protocol path uses the bare `file` / … spelling. Three readers
handled only one spelling or none:
- automatic handoffs listed the labels verbatim as `Tools used: tool
file, tool non-file`;
- the "session ended without a normal stop while working with files"
warning matched only the bare `file` spelling, so it never fired for a
real (closed-tool) session — the heuristic was effectively dead;
- a session whose only non-prompt observation was such a call took the
label as its page title.
Add one shared recognizer, `tool_family_from_title`, that maps both
spellings to a `ToolFamily`, and use it to drop the labels from the
handoff tool list and to drive the file-activity warning from either
spelling; reject `is_safe_tool_title` labels in the title fallback.
Synthesis-only; no page-read filtering (invariant #16 untouched).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
A manual memory_consolidate MCP call wrote the page directly through the
consolidator and never touched session_consolidation_jobs, so a session
whose automatic SessionEnd job had reached the terminal `failed` state
(which claim_next never re-picks) kept showing `failed` even though the
operator had just consolidated it — a two-sources-of-truth inconsistency.
Add a writer method reconcile_session_consolidation_completed(session_id)
that flips a `failed`/`pending`/`superseded` row for the session to
`completed`. It deliberately never touches a `running` row: that is an
active worker lease, and stomping it would violate invariants #2/#16, so
the automatic path's claim-guarded complete/fail stays the only lease
transition. The memory_consolidate handler calls it best-effort after a
successful, non-dry consolidate.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The multi-worktree/multi-target-dir release flow leaves stale 100-180 MB test
binaries that have exhausted disk; record cargo clean after every release as a
mandatory maintenance rule alongside the Homebrew tap bump.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Adds a delimited "Support Tier: Experimental → Supported Exit Criteria"
section enumerating what native Windows support requires and its current
status from repository evidence: CI trigger coverage (done), the hook-bundle
Windows job (done), the #[cfg(windows)] regression coverage (in-progress),
.exe code-signing (deferred-pending-policy — cert + CI-secret decision, App
Control/Smart App Control implications), and a native upgrade path (#801/#802).
It is the checklist a maintainer uses to decide promotion; it does not claim one.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The Windows tier promise no longer depends on a human remembering the
`windows` label. A leading `changes` job (dorny/paths-filter, SHA-pinned)
detects pull requests touching path handling, file locking, git plumbing,
and the hook bundle, and each Windows job runs when that path gate OR the
`windows` label fires. The `pull_request` trigger keeps no `paths:` filter
so all PRs still enter the workflow and the `labeled` override keeps working
for PRs outside those paths; schedule and workflow_dispatch are unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
fix(hooks): derive project name from raw cwd so Windows casing does not split projects (#871)
fix(windows): document owner-account service guidance and surface wiki git owner failures (#872)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
# Conflicts:
# CHANGELOG.md
A LocalSystem Windows service over a user-owned data dir trips libgit2's
dubious-ownership guard (CVE-2022-24765): every wiki commit fails with
code=Owner, but the failure was WARN-only, so capture kept working while
the wiki git history silently stopped advancing.
- Map an Owner-code git2 error to a distinct WikiError::GitOwner (the git
CLI fallback is skipped: it enforces the same guard and would refuse
identically). The owner check itself stays enabled (disabling it is
unsafe and reopens the CVE).
- Surface the startup baseline-checkpoint owner failure at ERROR with the
remedy (run the service as the owning user), classified through a small
testable seam.
- docs/windows.md Scenario E: recommend a <serviceaccount>, correct the
"only the data directory is account-sensitive" claim, and note the
WinSW error-1069 / stale-password gotcha for Microsoft-account / PIN /
Hello users.
A doctor wiki-git probe was considered but not added: doctor is
HTTP-only and never opens the store directly, so a probe would need a new
server endpoint — out of scope for a targeted fix.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The hook router derived a project's name from the cwd after
normalize_project_path_key had ASCII-lowercased the entire Windows
drive-letter/UNC path (basename included), so a session in
`D:\...\Default Project` was captured under `default project` while the
CLI kept `Default Project`. get_or_create_project matches names
case-sensitively, so one folder minted two projects.
Take the name from the raw cwd; the cache key and cwd-prefix match keep
the case-folded path so #806 handoff stickiness is unaffected.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Per the AGENTS.md mandatory-boundary-test rule, #866 touched purge/ownership
selection (a per-project/session isolation guard), so record the guard, its
enforcing code, and the adversarial cross-scope tests that would fail if it
regressed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Post-merge polish per audit: the example is stdlib-only and not meant to be
executed in place (100644), and the golden-set numbers are the contributor's
own measurement over a private wiki, so attribute rather than state as a
project benchmark.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Each test actively attempts a boundary violation and asserts refusal, with
a legitimate control case, so a future regression that removes the guard
fails CI:
- handoff_admission: a non-admin proxied caller's `any_owner: true` on
memory_handoff_accept is refused by require_admin_capability and does not
claim the baton; root's any_owner accept is the control.
- multi_session: a page with a non-null author_id is still readable by a
different operator via both search_pages_for_project and
page_body_by_ids, pinning that author_id is attribution and never a read
filter (invariant #16).
- agent_messages: the recipient cannot cancel the sender's message by exact
id (the specific-id sibling of the existing whole-outbox test), and two
concurrent pop_message calls on one message deliver it exactly once,
pinning the state='pending' CAS in pop_message_in_transaction.
- removal: reset/restore/reindex/uninstall --purge-data each refuse and
leave the data dir untouched when a sibling ai-memory process is
detected. Adds a minimal, test-only injection seam to
process_guard::sibling_processes() (AI_MEMORY_TEST_FORCE_SIBLING_PIDS)
so the refusal path is deterministically testable without spawning a
real sibling process; production behavior is unchanged since the
variable is never set outside a test harness.
Each test was verified to fail when its guard is removed and pass when
restored.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Add docs/security-boundaries.md: the source-of-truth map of every isolation/
security guard (per-project + workspace isolation, auth ladder, handoff single-
claim + any_owner gate, invariant-#16 pages-shared/author_id-never-a-filter +
supersession, active-project pointer, sanitizer boundary, messaging scope,
scope-resolution fail-closed, destructive-op guards, hook backpressure, network
posture, + #708 as FUTURE) → the adversarial test that fails if the guard is
removed, with a keep-current protocol. Add the AGENTS.md standing rule: touching
boundary code (or adding a raw-id/unscoped/cross-scope entry point) requires an
adversarial boundary-breaking test proven to bite, and a matching inventory
update, in the same change; subsumes the older scope/permissions test bullets.
Also record that CI now runs doctests (--all-targets excludes them).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The `//!` doc comment in commands/setup_agent.rs had an indented `docker run`
block, which rustdoc treats as a Rust code block and tried to compile —
"unknown start of token: \" — a broken doctest on both main and release/2.5.
`cargo test --workspace --all-targets` (what the gates and CI ran) excludes
doctests, so it slipped through. Fence the block as ```text, and add a
`cargo test --workspace --doc` step to CI so a broken doctest can't ship again.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The single authorize_project choke point at scope resolution is necessary but
not sufficient: unscoped/cross-project reads (memory_query global, global
recent, web search) and raw-id entry points (session/run/page id,
page_evidence_counts) never resolve a single scope and would bypass it. Document
that unscoped reads must filter by readable project in SQL before LIMIT (post-
filter leaks hit existence/counts) and that every bare-id path must resolve
id->project then authorize; recommend making unguarded lookups crate-private so
new call sites are compiler-forced through the gate. Add ship-inert / never-
fail-closed safeguards (empty grants + access_mode 'open' = no-op; restricted
with zero grants still admits root+creator; unreadable grants table degrades to
open with a warning) and extend the verification plan with unscoped-read-leak,
raw-id-authz, and ship-inert tests. Fix the stale release/2.3 -> release/2.5
migration target. Still design-only; no code.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
A no-scope memory_message_pop/list resolves the shared active-project slot
(whichever project published last), so two same-operator agents with no
session id can have a pop land on a different inbox than the on-start notice
counted — "you have mail" then a silent {"message": null}. Add
ScopeSource::is_inferred() (broader than is_fallback: includes SharedSlot),
and when a message read is empty AND the scope was inferred, return the
resolved_scope (workspace+project), scope_source, and a hint to re-run with
explicit scope. Explicit/session-scoped empty reads are unchanged; the
mis-scoped pop consumes nothing. Regression tests reproduce the divergence
(mail in B, no-scope pop resolves A -> diagnosed null + hint; explicit pop of
B delivers) for both pop and list.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Resolve issue #810 and supersede the stub PR #828 with fair, first-party-sourced
assessments of five agent-memory projects, kept in sync across all three
comparison docs per the AGENTS.md rule.
Each project was researched from its own repo/README (and, where the name could
mislead, its LICENSE and the GitHub metadata endpoint) on 2026-09-22, not from a
name search:
- Engram (MIT, Go, SQLite+FTS5, MCP, single binary) — closest new sibling;
agent-driven MCP capture and SQLite-as-truth vs our hook capture + git wiki.
- EverOS (Apache-2.0, Python, Markdown truth + SQLite/LanceDB) — file-first
sibling, but LLM-required for the core flow.
- memU (Apache-2.0, Python, Markdown skill distillation, no MCP) — file-first
LLM-leaning; hosted option.
- Caura (Apache-2.0 + managed, Postgres+pgvector) — governed multi-agent fleet
memory; different buyer; its trust-tier governance is an honest gap for us.
- TencentDB Agent Memory (MIT, Node, SQLite default, proxy fan-in) — needs no
Tencent DB despite the name.
comparison.md: 4 sourced camp-table rows (replacing the vague PR #828 rows),
maturity rows, and migration notes. research-2026-landscape.md: §3 camp-table
rows + inline entries per the no-standalone-doc convention + §6 sources +
popularity rows. competitive-parity.md: migration verdicts + a fleet-governance
honest-gap note. All benchmark numbers attributed to their named benchmark and
flagged self-reported.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
consolidate_session_multi's build_update built each WritePageRequest.path
straight from LLM output with no portability sanitize. PagePath::new is
deliberately tolerant, so a path with a Windows-illegal character (e.g. a
`:` copied from a conventional-commit subject) passed construction, entered
the atomic apply_batch call, and only failed ensure_portable at write time
- aborting the whole per-session consolidation batch and losing every other
page. This is the same class of bug already fixed for bootstrap in #847.
Factor slugify_page_path (+ PATH_ILLEGAL_CHARS) out of bootstrap.rs into a
new shared crate::path_sanitize module, and call it from build_update
before PagePath::new so rule-routing, per-user slot placement, and the
req.path == anchor comparison in consolidate_session_multi all see the same
sanitized path. Add an ensure_portable() final guard where the batch is
assembled that skips (warns on) a page slugify couldn't fix, instead of
aborting the batch, mirroring bootstrap's #847 fix.
Fixes#848.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Document that ai-memory run resolves the native session dir and installs
hooks from its own environment, so a custom CLAUDE_CONFIG_DIR set only
inside a harness wrapper causes a split-brain and "native transcript
import failed" (#820). Export per-account config dirs before invoking
ai-memory run. The launch-configuration ergonomics half of #820 remains a
separate design-first feature.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
ai-memory serve leaked one file descriptor per dead hook/MCP peer.
A client whose connection dies without sending FIN (laptop sleep, a
VPN/Tailscale flap, an abrupt kill) leaves the accepted socket
ESTABLISHED forever, since the OS keepalive default is off. Over ~2-3
days of normal churn that exhausts the 1024-fd default and breaks the
healthcheck -- an unauthenticated availability/DoS. The rmcp
session-table half of this leak was already fixed in 2.4.0 by the
rmcp 2.x bump; this closes the remaining half-open-socket half.
Accepted sockets now get TCP keepalive via socket2, wired through
axum::serve::ListenerExt::tap_io (a hand-rolled axum::serve::Listener
newtype was tried first, but into_make_service_with_connect_info's
Connected<IncomingStream<'_, L>> bound is only implemented by axum for
its own TcpListener and, generically, for TapIo<L, F> -- never for an
arbitrary third-party L, and the orphan rule blocks implementing it
ourselves since neither Connected, SocketAddr, nor IncomingStream is
local to this crate). tap_io keeps the real peer SocketAddr flowing to
ConnectInfo while still touching every accepted stream.
New tcp_keepalive_secs config field (default 60s; AI_MEMORY_TCP_KEEPALIVE_SECS
env override; 0 disables keepalive entirely), read once through the
existing Config::load path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Clarify that AI_MEMORY_LLM_BASE_URL redirects only the LLM endpoint;
AI_MEMORY_EMBEDDING_BASE_URL must be set separately or embedding traffic
keeps going to the embedding provider default. Both env vars already exist
and are consumed (config.rs figment mapping); this addresses the confusion
in #843 (the vars themselves were already present and documented).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
`ai-memory bootstrap` returned a 500 when the LLM emitted a page path
containing a character Windows refuses (e.g. `:` copied verbatim from a
conventional-commit subject like `build(sandbox): orchestrate`). The
write loop validated model-produced paths with the deliberately
tolerant `PagePath::new` only; the bad path entered the batch and only
failed later at `ensure_portable` inside `Wiki::apply_batch`, which is
atomic — one bad page aborted every page in the run.
Add `slugify_page_path`, which replaces Windows-illegal filename
characters and ASCII control bytes with `-` in each `/`-separated path
component, preserving the directory shape. Run it before
`PagePath::new`, then call `ensure_portable` as a final guard; a path
that still fails after sanitizing is skipped with the existing `warn!`
rather than aborting the batch.
Checked the other LLM -> PagePath write funnels: `memory_write_page`
(server.rs) validates a single explicit write and returns a clear
error to the caller, not a batch, so it isn't the same failure mode.
`consolidator.rs::consolidate_session_multi` builds `PagePath`s from
LLM output via `build_update` and also calls `Wiki::apply_batch`, so it
has the same atomic-abort exposure — flagged for a follow-up ticket
since fixing it safely means threading sanitization through
`build_update`'s slot-placement/rule-routing logic and several test
call sites, which is a bigger, separate change.
Regression test `bad_windows_path_is_sanitized_not_aborted` fails on
the prior write loop and passes with the fix; unit test
`slugify_page_path_replaces_illegal_chars_and_keeps_slashes` covers the
helper directly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Disabling btrfs CoW on the VM storage dir cut the I/O-bound warm suite
~25% (1600s->1193s), confirming fsync-on-CoW was the bottleneck, but the
full suite still doesn't beat windows-latest (~1000s). Verdict unchanged:
the dockur VM is a focused-iteration / full-local-coverage tool, not a
faster full gate; no Phase 2 full-suite automation on this hardware.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2.4.0 shipped from release/2.4; bring main up to it so post-release fixes
(2.4.1) cut from main and features branch off a 2.4.0 base, per AGENTS.md
trunk-based releases. main keeps the #844 Windows cross-build gate on top;
CHANGELOG [Unreleased] is empty (its pre-release entries are in [2.4.0]);
workspace version is 2.4.0. release/2.4 now goes dormant.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Land the design study that ci.yml's windows-cross job references, and
append the measured results: Phase 0 (cross-build gate) shipped and green;
Phase 1 (dockur Windows Server 2025 VM) runs the full suite green but the
warm full run (1600s) is slower than windows-latest (~1000s) because the
I/O-heavy tests hit btrfs CoW — so the VM's value is on-demand focused
iteration, not a faster full gate. Records the host-specific dockur
accommodations (SELinux :Z + label:disable, Defender-off, QEMU-monitor
provisioning) and revises the recommendation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
workflows_keep_fixed_rust_jobs_on_the_fixed_toolchain counts ci.yml jobs
pinned to Rust 1.95. windows-cross is now the second (after source-install),
so bump the expected count from 1 to 2. The guard's structural checks
(# 1.95 -> with: -> toolchain: "1.95") already pass for the new job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
rust-toolchain.toml pins 1.95 and overrides the active toolchain inside
the repo, so `cargo xwin build` runs under 1.95. dtolnay/rust-toolchain
had added the msvc target to `stable` instead, so the build failed with
"can't find crate for `core`" (the target's std was on the wrong
toolchain). Pin the job to 1.95 and add the target there, mirroring the
source-install job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
A #[cfg(windows)] compile or link break currently surfaces only in
windows.yml (nightly / label-gated, ~1000s) or at release time. Add a
Linux job that cross-builds the whole workspace --all-targets for
x86_64-pc-windows-msvc via cargo-xwin (clang-cl + lld-link against an
auto-downloaded MSVC CRT/SDK), so bundled SQLite and vendored libgit2
build for the target and the #[cfg(windows)] test binaries compile — at
Linux CI speed, on every PR.
Build-only by design: it proves the code compiles and links for Windows,
not that it behaves correctly there. Wine is deliberately not used to
fake a runtime (native PowerShell, NTFS case-folding, Win32 file-locking,
and verbatim \\?\ path handling are not faithfully reproduced); runtime
correctness stays with windows.yml and the local dockur VM loop
(docs/design-windows-ci.md).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
powershell_utf8::powershell_hook_posts_json_as_utf8_bytes flaked on
windows-latest: a cold PowerShell start (spawn + JIT + dot-sourcing the hook
lib) exceeded the receiver's 10s accept deadline while output.status.success()
still passed — i.e. the hook worked, the mock server just gave up too early.
Raised the deadline to 60s. It only bounds failure detection (a working hook
connects in seconds; a broken one is caught by the status assertion), so this
removes the false negative without masking a real regression. Test-only.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Keeps it consistent with the project's fast-CI-per-merge rule (Windows off
per-PR feedback by default; opt in with the `windows` label, same as `test`).
Still runs nightly and on manual dispatch. The PowerShell hook-branch coverage
this job adds is real; it just follows the same opt-in gate as the rest of the
Windows workflow rather than running a windows-latest runner on every PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The auto-improve reviewer only loads `_rules/`/`procedures/` page bodies,
and its recent-page context is a recency-ordered, one-line-per-page list
drawn from the shared project briefing. That list is dominated by
`sessions/` pages, so durable pages in `decisions/`/`gotchas/` outside
`_rules/` are effectively invisible to the reviewer, which causes
duplicate proposals.
Filter `sessions/` pages out of the recent-page list the reviewer builds
its prompt, patchable set, and existing-page index from. This is scoped
to the auto-improve reviewer only: it operates on the already-fetched
`briefing.recent_pages` and does not touch the shared
`briefing_for_project` reader, so the SessionStart briefing and
`memory_briefing` — where session pages legitimately belong — are
unchanged.
Also document the reviewer's context limits in
docs/auto-improvement-loop.md, with configurable patchable prefixes and
embedding-nearest dedup called out as deferred future work.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The proposal-staging loop returned `StoreError::InvalidState` for a
`Create` whose target already exists and for an `Update`/patch whose
target is missing. Because those returns fire inside the staging
transaction, one misclassified proposal discarded the entire run — every
sibling proposal and the run row. Create/update misclassification is an
ordinary probabilistic LLM error, not corrupt state.
Convert both arms to the existing skip machinery (push a `SkippedProposal`
with an accurate reason and `continue`), matching how a pending-target
UNIQUE collision is already handled and surfaced. The self-contradicting
case (two proposals in one run targeting the same path) stays a hard
error, and a create-on-existing is never coerced to an update — the page
could be pinned, so auto-applying would violate auto-improvement Safety
Invariant #10.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
release/2.4 already ships V65__page_compacted_marker (A2). Keeping this fix at
V65 would collide when main forward-merges into release/2.4 (two different V65
migrations -> refinery version conflict, a missing column on 2.4 stores, and a
broken 2.3->2.4 upgrade path). Renumbered to V66 so the forward-merge yields a
contiguous V65 (compacted) + V66 (claim attempts); the schema-version pin is
bumped to 66. main carries a harmless V64->V66 gap (it never ships a 2.3.x with
this migration; it reaches users via release/2.4).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
The Codex/ChatGPT backend behind `openai-oauth` restricts the selectable
model set to a small server-defined list and rejects other ids
(including `gpt-5-mini`) with a deterministic 400. The TIP blocks in
`docs/llm-providers.md` and `docs/install.md` recommended setting
`AI_MEMORY_LLM_MODEL=gpt-5-mini`, which fails on that backend.
Advise leaving the provider default (`gpt-5.5`) for `openai-oauth`/`codex`,
keep `claude-haiku-4-5` for `anthropic-oauth`, and qualify `gpt-5-mini`
for `copilot` as unverified rather than asserting it works.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
A confirmation local re-run on the same 2.4 commit bounced the run-1 per-slice
dips back (knowledge-update 0.875->0.903, ss-user 0.734->0.750, overall
0.815->0.821), confirming the dips were cross-run noise, not a regression.
Documents the variance band so a single small-n slice from one run is not
over-read.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Full-dataset run on release/2.4 (commit 79fea400, 470 scored): FTS baseline vs
local-embeddings candidate. Populates the previously-empty R2 A/B baselines
(local +0.149 hit@5 / +0.254 recall@10 over zero-LLM FTS, ~90ms p50 cost),
records the baseline-vs-baseline determinism check (all deltas 0.000), and
confirms no default-ranking regression vs the 2026-09-01 snapshot.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm