Merge remote-tracking branch 'upstream/main' into feature/admission-webhooks

# Conflicts:
#	crates/ai-memory-mcp/src/server.rs
This commit is contained in:
Djalma Júnior
2026-05-30 10:02:34 -03:00
20 changed files with 1759 additions and 113 deletions
+114 -1
View File
@@ -7,6 +7,118 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **`memory_query { global: true }` — cross-project global search** that
reaches every project in every workspace in one call, with each hit
annotated by its workspace + project so the agent can tell where it
came from. Use when the agent doesn't know which project holds a
cross-cutting note (shared infra/ops, a sibling app). Mutually
exclusive with `scopes`/`project`/`workspace`. Routing snippet +
`MEMORY_INSTRUCTIONS` now teach both broadening modes (`scopes` for
named siblings, `global=true` for unknown locations) and explicitly
warn that `memory_query` returns snippets — use `memory_read_page`
for full bodies. The prompt-surface contradiction the original PR
shipped ("there is no global 'search everything' mode" right after
the bullet advertising `global=true`) was caught in the post-merge
audit and rewritten; the prompt regression test now refuses any
variant of that legacy phrasing
([#56], thanks @djalmajr).
- **Cross-project wiki links + dependency graph.** Wikilinks gain an
explicit scope qualifier: `[[project:path.md]]` for a sibling project
in the same workspace, `[[workspace/project:path.md]]` for another
workspace. Bare links are unchanged (resolve within the source's own
project). `links.to_workspace` / `links.to_project` join the primary
key so the same `to_path` can land in two different projects without
colliding. `memory_lint` now reports dangling cross-project refs
(typo'd project vs missing/renamed target page), `memory_briefing`
exposes `cross_project_dependents` / `cross_project_dependencies`
per project, and `GET /api/v1/graph` returns the resolved cross-
project edges for a graph view. Migration V13 rebuilds the `links`
table preserving existing rows as `(to_workspace=NULL,
to_project=NULL)` — same "local" semantics as before
([#57], thanks @djalmajr).
### Changed
- **FTS5 queries OR-join bare multi-word inputs** instead of the
pre-existing AND default. A natural-language query like
`"have we discussed cross project search strategy"` previously
required every word to co-occur in one page — near-zero recall for
multi-word queries, which the caller silently mistook for "never
recorded". OR + BM25 ranking (callers already `ORDER BY rank`) keeps
the best-matching pages at the top of the list, so the user-visible
top-N is still AND-ish; OR just adds a relevant tail instead of
returning nothing. Explicit FTS5 syntax (`OR`/`AND`/`NOT`/`NEAR`,
quoted phrases, parens) is detected and preserved verbatim so the
exact-match escape hatch stays available. 5 new unit tests guard the
preservation contract (post-merge audit). Migration V12 rebuilds the
FTS tables with `unicode61 remove_diacritics 2` so accent-free
Portuguese queries (`"descricao da sessao"`) match accented stored
text (`"descrição da sessão"`); contentless FTS — source rows
untouched ([#58], thanks @djalmajr).
- **MCP write tools now honour the session's project (and create
named projects on demand).** Three correctness fixes on
`memory_write_page` / `memory_lint` / `memory_forget_sweep`:
- A `memory_write_page { project: "X" }` for a project name that
doesn't exist used to silently fall through to the session's
active project (find-only resolution); writes meant for a fresh
project polluted the current one. A new `write_target_ids`
helper uses **get-or-create** for an explicit project name, so
a named write always lands where the agent asked.
- `memory_lint` + `memory_forget_sweep` previously always targeted
the server's baked `--project` regardless of the session, so a
cross-project lint or retention sweep could never reach the
project the user was actually working in. Both now resolve
through the same find-only `effective_ids_for_read_args` path
the read tools use, with the hook-published active project as
the fallback.
- Both `lint` / `sweep` and the new `write_page` add explicit
`workspace` + `project` args (defaulted to current session,
documented with the v0.5.2 "**Omit unless the user explicitly
names a *different* project.**" tail). 2 regression tests cover
"Bug B" (explicit-project write must create + land) and
"Bug C" (sweep must evaluate the named project, not the baked
default) ([#59], thanks @djalmajr).
## [0.7.1] - 2026-05-29
### Fixed
- **`install-hooks --agent codex` no longer panics with `index not found`**
when `~/.codex/config.toml` carries an `[mcp_servers]` table that has other
MCP servers (context7, node_repl, …) but no `ai-memory` entry — a
perfectly valid setup since ai-memory can integrate via hooks alone.
`infer_codex_mcp_config` used `toml_edit`'s panicking `Index` impl with
bare `[]` chains; it now walks the table via `.get()` and returns `None`
on any missing key. Mirrors the safe pattern the JSON variant has used
all along. Adds 4 regression tests covering missing-entry,
missing-table, empty-doc, and bare-entry inputs
([#53], thanks @Otavio-Machado-Santos).
- **`install-hooks --agent claude-code` no longer silently stages 0 scripts
and points `settings.json` at an empty directory.** On macOS — and any
install where the binary lives outside the repo and the system package
paths (`/usr/local/share`, `/usr/share`) are absent — `resolve_hooks_dir`
fell through to the data-local candidate, which was *also* the staging
destination. The wipe-then-copy flow inside `stage_hook_scripts_in` then
deleted the very scripts it was about to read, leaving 0 copied; the
caller proceeded to rewrite `settings.json` anyway, disabling capture
with no error. The function now (a) canonicalizes source and destination
paths, skips the wipe + copy when they match and verifies in-place,
preserving any scripts a prior `setup-agent` run extracted there, and
(b) bails with an actionable error pointing at `--hooks-dir` or
`ai-memory setup-agent` whenever zero scripts are present in either
branch. Adds 3 regression tests
([#52], thanks @Otavio-Machado-Santos).
- **macOS thin-client wrapper no longer crashes with "Permission denied" in
the log file appender.** The `bin/ai-memory` wrapper passed
`-u $(id -u):$(id -g)` to the one-shot helper container, which on macOS
collides with the data volume owner (uid 1000 inside the container vs
uid 501/502 on the host). The wrapper now skips `-u` on Darwin so the
container runs as its default uid 1000 — Docker Desktop's file-sharing
layer handles host ownership transparently — while Linux and other
Unix systems continue to receive `-u`. Same change also hardens the
`${TTY_ARGS[@]}` / `${NETWORK_ARGS[@]}` / `${ENV_ARGS[@]}` /
`${USER_ARGS[@]}` expansions for `set -u` compatibility on macOS's
default bash 3.2 ([#51], thanks @abnersajr; supersedes [#50]).
## [0.7.0] - 2026-05-29
### Added
@@ -453,7 +565,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- Consolidator used server startup default project instead of the
session's actual project.
[Unreleased]: https://github.com/akitaonrails/ai-memory/compare/v0.7.0...HEAD
[Unreleased]: https://github.com/akitaonrails/ai-memory/compare/v0.7.1...HEAD
[0.7.1]: https://github.com/akitaonrails/ai-memory/releases/tag/v0.7.1
[0.7.0]: https://github.com/akitaonrails/ai-memory/releases/tag/v0.7.0
[0.6.1]: https://github.com/akitaonrails/ai-memory/releases/tag/v0.6.1
[0.6.0]: https://github.com/akitaonrails/ai-memory/releases/tag/v0.6.0
Generated
+10 -10
View File
@@ -31,7 +31,7 @@ dependencies = [
[[package]]
name = "ai-memory-cli"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -71,7 +71,7 @@ dependencies = [
[[package]]
name = "ai-memory-consolidate"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-core",
"ai-memory-llm",
@@ -92,7 +92,7 @@ dependencies = [
[[package]]
name = "ai-memory-core"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"jiff",
"regex",
@@ -106,7 +106,7 @@ dependencies = [
[[package]]
name = "ai-memory-eval"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -125,7 +125,7 @@ dependencies = [
[[package]]
name = "ai-memory-hooks"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -147,7 +147,7 @@ dependencies = [
[[package]]
name = "ai-memory-llm"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-core",
"anyhow",
@@ -169,7 +169,7 @@ dependencies = [
[[package]]
name = "ai-memory-mcp"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -195,7 +195,7 @@ dependencies = [
[[package]]
name = "ai-memory-store"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-core",
"anyhow",
@@ -215,7 +215,7 @@ dependencies = [
[[package]]
name = "ai-memory-web"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-core",
"ai-memory-store",
@@ -238,7 +238,7 @@ dependencies = [
[[package]]
name = "ai-memory-wiki"
version = "0.7.0"
version = "0.7.1"
dependencies = [
"ai-memory-core",
"ai-memory-llm",
+9 -9
View File
@@ -16,7 +16,7 @@ members = [
]
[workspace.package]
version = "0.7.0"
version = "0.7.1"
edition = "2024"
rust-version = "1.95"
license = "MIT"
@@ -25,14 +25,14 @@ authors = ["Fabio Akita <boss@akitaonrails.com>"]
[workspace.dependencies]
# Inter-crate dependencies.
ai-memory-core = { path = "crates/ai-memory-core", version = "0.7.0" }
ai-memory-store = { path = "crates/ai-memory-store", version = "0.7.0" }
ai-memory-wiki = { path = "crates/ai-memory-wiki", version = "0.7.0" }
ai-memory-mcp = { path = "crates/ai-memory-mcp", version = "0.7.0" }
ai-memory-hooks = { path = "crates/ai-memory-hooks", version = "0.7.0" }
ai-memory-llm = { path = "crates/ai-memory-llm", version = "0.7.0" }
ai-memory-consolidate = { path = "crates/ai-memory-consolidate", version = "0.7.0" }
ai-memory-web = { path = "crates/ai-memory-web", version = "0.7.0" }
ai-memory-core = { path = "crates/ai-memory-core", version = "0.7.1" }
ai-memory-store = { path = "crates/ai-memory-store", version = "0.7.1" }
ai-memory-wiki = { path = "crates/ai-memory-wiki", version = "0.7.1" }
ai-memory-mcp = { path = "crates/ai-memory-mcp", version = "0.7.1" }
ai-memory-hooks = { path = "crates/ai-memory-hooks", version = "0.7.1" }
ai-memory-llm = { path = "crates/ai-memory-llm", version = "0.7.1" }
ai-memory-consolidate = { path = "crates/ai-memory-consolidate", version = "0.7.1" }
ai-memory-web = { path = "crates/ai-memory-web", version = "0.7.1" }
# (Workspace shared deps follow below)
+18 -4
View File
@@ -248,6 +248,10 @@ done
# 127.0.0.1 would mean "this helper container", so thin-client commands like
# `status` and `bootstrap` could not reach the default server.
NETWORK_ARGS=()
# Default: map the container process to the host user so files written
# through bind mounts (~/.claude/settings.json, $PWD/, …) stay editable
# by the invoking user. macOS is the one exception (see Darwin arm).
USER_ARGS=(-u "$(id -u):$(id -g)")
case "$(uname -s 2>/dev/null || true)" in
Linux)
if [ -z "${AI_MEMORY_SERVER_URL:-}" ]; then
@@ -258,6 +262,16 @@ case "$(uname -s 2>/dev/null || true)" in
if [ -z "${AI_MEMORY_SERVER_URL:-}" ]; then
ENV_ARGS+=(-e "AI_MEMORY_SERVER_URL=http://host.docker.internal:49374")
fi
# On macOS, Docker Desktop handles file-sharing permissions via its
# gRPC/SSH layer. Passing -u <host-uid>:<host-gid> causes a UID
# mismatch: the data volume is typically owned by the container's
# internal uid 1000 (the ai-memory user), but the host UID on macOS
# is usually 501/502. The one-shot wrapper container then cannot
# create log files or write to the data dir, crashing with
# "Permission denied" in the rolling file appender.
# Omitting -u lets the container run as its default (uid 1000)
# which matches the volume owner.
USER_ARGS=()
;;
esac
@@ -292,14 +306,14 @@ fi
# binary can derive the *real* host-side project name from it. The
# `commands::resolve_project_name` helper checks this env var first
# and falls back to the cwd basename only if it's unset.
exec "${DOCKER}" run --rm "${TTY_ARGS[@]}" \
"${NETWORK_ARGS[@]}" \
exec "${DOCKER}" run --rm ${TTY_ARGS[@]+"${TTY_ARGS[@]}"} \
${NETWORK_ARGS[@]+"${NETWORK_ARGS[@]}"} \
-v "${HOME}:${HOME}" \
-v "${PWD}:/work" \
-w /work \
-e HOME="${HOME}" \
-e AI_MEMORY_HOST_CWD="${PWD}" \
-u "$(id -u):$(id -g)" \
${USER_ARGS[@]+"${USER_ARGS[@]}"} \
"${DATA_ARGS[@]}" \
"${ENV_ARGS[@]}" \
${ENV_ARGS[@]+"${ENV_ARGS[@]}"} \
"${IMAGE}" "$@"
@@ -257,13 +257,22 @@ fn infer_json_mcp_config(
fn infer_codex_mcp_config(content: &str) -> Option<InferredMcpConfig> {
let doc: toml_edit::DocumentMut = content.parse().ok()?;
let server = &doc["mcp_servers"]["ai-memory"];
let hook_server_url = server["url"]
.as_str()
// `toml_edit::Item`'s `Index` impl panics on missing keys, so this
// walks the table chain with `.get()` instead. A user with
// `[mcp_servers.context7]` but no `[mcp_servers.ai-memory]` is a
// perfectly valid hooks-only Codex setup (issue #53) — return None
// rather than abort the whole install with a stack trace.
let server = doc.get("mcp_servers")?.get("ai-memory")?;
let hook_server_url = server
.get("url")
.and_then(|v| v.as_str())
.and_then(hook_server_url_from_mcp_url);
let auth_token = server["http_headers"]["Authorization"]
.as_str()
.or_else(|| server["headers"]["Authorization"].as_str())
let auth_token = server
.get("http_headers")
.and_then(|h| h.get("Authorization"))
.or_else(|| server.get("headers").and_then(|h| h.get("Authorization")))
.and_then(|v| v.as_str())
.and_then(bearer_token_from_header);
if hook_server_url.is_none() && auth_token.is_none() {
return None;
@@ -1297,19 +1306,32 @@ fn stage_hook_scripts_in(
fs::create_dir_all(&dest_root)
.with_context(|| format!("creating staging dir {}", dest_root.display()))?;
// Wipe any previously-staged scripts that the current bundle
// no longer ships. Idempotent re-runs against an old install
// shouldn't leave stale entries pointed at by nothing.
if let Ok(entries) = fs::read_dir(&dest_root) {
for entry in entries.flatten() {
let p = entry.path();
if p.is_file() && is_hook_script_file(&p) {
fs::remove_file(&p).ok();
// When `resolve_hooks_dir` falls through to the data-local
// candidate (e.g. docker `setup-agent` already extracted the
// bundle into ~/.local/share/ai-memory/hooks/<agent>/, or a prior
// install left scripts in place), the source dir IS the
// destination dir. The wipe-then-copy flow below would delete the
// very scripts we mean to install before reading them, leaving 0
// copied and a settings.json pointing at an empty directory
// (issue #52). Detect that case via canonical paths and verify
// the existing layout in place instead of touching it.
let same_path = same_canonical_dir(source_dir, &dest_root);
if !same_path {
// Wipe any previously-staged scripts that the current bundle
// no longer ships. Idempotent re-runs against an old install
// shouldn't leave stale entries pointed at by nothing.
if let Ok(entries) = fs::read_dir(&dest_root) {
for entry in entries.flatten() {
let p = entry.path();
if p.is_file() && is_hook_script_file(&p) {
fs::remove_file(&p).ok();
}
}
}
}
let mut copied = 0_usize;
let mut count = 0_usize;
for entry in fs::read_dir(source_dir)
.with_context(|| format!("reading source bundle {}", source_dir.display()))?
{
@@ -1318,28 +1340,57 @@ fn stage_hook_scripts_in(
if !from.is_file() || !is_hook_script_file(&from) {
continue;
}
copy_hook_file(&from, &dest_root)?;
copied += 1;
if !same_path {
copy_hook_file(&from, &dest_root)?;
}
count += 1;
}
copy_support_hook_scripts(source_dir, &dest_root)?;
if !same_path {
copy_support_hook_scripts(source_dir, &dest_root)?;
// Stage the shared `_lib.sh` helper alongside the event scripts so
// they can `. "$(dirname "$0")/_lib.sh"` without depending on the
// user's PATH or repo layout. The helper lives ONCE in
// `hooks/_lib.sh` (one parent up from the agent-specific dir) —
// staging it here is what keeps every agent's runtime view
// consistent with the source of truth.
if let Some(shared) = source_dir.parent().map(|p| p.join("_lib.sh"))
&& shared.is_file()
{
copy_hook_file(&shared, &dest_root)?;
// Stage the shared `_lib.sh` helper alongside the event scripts so
// they can `. "$(dirname "$0")/_lib.sh"` without depending on the
// user's PATH or repo layout. The helper lives ONCE in
// `hooks/_lib.sh` (one parent up from the agent-specific dir) —
// staging it here is what keeps every agent's runtime view
// consistent with the source of truth.
if let Some(shared) = source_dir.parent().map(|p| p.join("_lib.sh"))
&& shared.is_file()
{
copy_hook_file(&shared, &dest_root)?;
}
}
eprintln!("✓ staged {copied} hook script(s) → {}", dest_root.display());
if count == 0 {
anyhow::bail!(
"no hook scripts found at {}.\n\
Refusing to install — pointing the agent's settings at an empty \
directory would silently disable all capture. Either pass \
`--hooks-dir <path>` to point at a populated source tree, or run \
`ai-memory setup-agent --agent <name>` first to extract the \
bundled scripts.",
source_dir.display()
);
}
let verb = if same_path { "verified" } else { "staged" };
eprintln!("✓ {verb} {count} hook script(s) → {}", dest_root.display());
Ok(dest_root)
}
/// `true` when `a` and `b` resolve to the same directory after symlink
/// canonicalization. Falls back to literal `==` if either canonicalize
/// call fails (e.g. dest hasn't been created yet on Windows, network
/// FS quirks). The caller has already `create_dir_all`'d both ends
/// in the staging flow, so the fast path almost always wins.
fn same_canonical_dir(a: &Path, b: &Path) -> bool {
match (a.canonicalize(), b.canonicalize()) {
(Ok(ca), Ok(cb)) => ca == cb,
_ => a == b,
}
}
/// Copy a single hook file (event script or shared `_lib.sh`) into the
/// staging dir, preserving the executable bit on Unix. Centralised so
/// the script bulk-copy and the `_lib.sh` companion follow the same
@@ -1629,6 +1680,62 @@ Authorization = "Bearer secret-token"
assert_eq!(inferred.auth_token.as_deref(), Some("secret-token"));
}
/// Regression for issue #53 — `install-hooks --agent codex` used to
/// panic with "index not found" when `~/.codex/config.toml` had an
/// `[mcp_servers]` table populated with *other* servers (context7,
/// node_repl, …) but no ai-memory entry. A perfectly valid setup —
/// ai-memory can live in Codex via hooks only without being an MCP
/// server — must return None, not abort the whole install.
#[test]
fn codex_mcp_inference_returns_none_when_ai_memory_entry_missing() {
let inferred = infer_codex_mcp_config(
r#"[mcp_servers.context7]
url = "http://localhost:9000/mcp"
[mcp_servers.node_repl]
command = "npx"
args = ["node-repl"]
"#,
);
assert!(
inferred.is_none(),
"missing [mcp_servers.ai-memory] must yield None, got {inferred:?}"
);
}
/// Same regression class — no `[mcp_servers]` table at all means
/// the user is on a hooks-only / fresh config; we should return
/// None rather than panic on the first index.
#[test]
fn codex_mcp_inference_returns_none_when_no_mcp_servers_table() {
let inferred = infer_codex_mcp_config(
r#"# fresh codex config
model = "gpt-5"
"#,
);
assert!(inferred.is_none());
}
/// And the empty-file edge case the parser still accepts.
#[test]
fn codex_mcp_inference_returns_none_for_empty_doc() {
assert!(infer_codex_mcp_config("").is_none());
}
/// An ai-memory entry that exists but ships neither a `url` nor an
/// `Authorization` header still falls back to None (caller infers
/// from defaults). Distinguishes "config absent" from "config
/// present but unhelpful" — both yield None, neither panics.
#[test]
fn codex_mcp_inference_returns_none_for_bare_ai_memory_entry() {
let inferred = infer_codex_mcp_config(
r#"[mcp_servers.ai-memory]
# intentionally empty — no url, no headers.
"#,
);
assert!(inferred.is_none());
}
#[test]
fn bundled_posix_and_powershell_hooks_stay_in_parity() {
let hooks_root = Path::new(env!("CARGO_MANIFEST_DIR"))
@@ -1794,6 +1901,86 @@ Authorization = "Bearer secret-token"
assert!(!staged.join("_lib.sh").exists());
}
/// Regression for issue #52 — when `resolve_hooks_dir` picks the
/// data-local dir as the source bundle (the docker `setup-agent`
/// flow extracts scripts there) AND the staging destination is
/// the *same* dir, the pre-fix wipe-then-copy loop would delete
/// every populated script and report `staged 0`. The same-path
/// branch must verify in place without wiping, so existing scripts
/// survive a re-run.
#[test]
fn stage_hook_scripts_preserves_in_place_scripts_when_source_equals_dest() {
let tmp = TempDir::new().unwrap();
let data_dir = tmp.path().join("data");
let agent_label = "stage-in-place";
// Simulate "scripts already extracted into the data-local
// hooks dir by a prior `setup-agent` run".
let in_place = data_dir.join("ai-memory/hooks").join(agent_label);
fs::create_dir_all(&in_place).unwrap();
stub_scripts(&in_place, &["session-start.sh", "post-tool-use.sh"]);
// Source == destination (this is what resolve_hooks_dir hands
// us when no other candidate exists).
let staged = stage_hook_scripts_in(&in_place, agent_label, &data_dir).unwrap();
assert_eq!(staged, in_place, "destination must canonicalize to source");
assert!(
staged.join("session-start.sh").is_file(),
"in-place script must survive the same-path branch (not be wiped)"
);
assert!(
staged.join("post-tool-use.sh").is_file(),
"in-place script must survive the same-path branch (not be wiped)"
);
}
/// Regression for issue #52 — the failure that the reporter actually
/// hit: `resolve_hooks_dir` resolved to a pre-existing but empty
/// data-local dir, so source == dest and there's nothing to verify.
/// The pre-fix code silently returned Ok with `copied = 0` and the
/// caller went on to rewrite `settings.json` against an empty dir,
/// disabling capture without any error. We must bail with an
/// actionable message instead.
#[test]
fn stage_hook_scripts_bails_when_source_equals_empty_dest() {
let tmp = TempDir::new().unwrap();
let data_dir = tmp.path().join("data");
let agent_label = "stage-empty-in-place";
let in_place = data_dir.join("ai-memory/hooks").join(agent_label);
fs::create_dir_all(&in_place).unwrap();
// Intentionally no scripts in `in_place`.
let err = stage_hook_scripts_in(&in_place, agent_label, &data_dir)
.expect_err("an empty source dir must produce a hard error, not Ok(0)");
let msg = format!("{err:#}");
assert!(
msg.contains("no hook scripts"),
"error should call out the empty source: {msg}"
);
assert!(
msg.contains("--hooks-dir") || msg.contains("setup-agent"),
"error should point at the workaround (--hooks-dir or setup-agent): {msg}"
);
}
/// Regression for issue #52 — same fail-on-zero guard applies even
/// when source and dest are different paths (e.g. user pointed
/// `--hooks-dir` at the wrong dir). Previously this also silently
/// returned Ok with `copied = 0`.
#[test]
fn stage_hook_scripts_bails_when_source_dir_is_empty() {
let tmp = TempDir::new().unwrap();
let bundle = tmp.path().join("hooks");
let agent_src = bundle.join("stage-empty-src");
fs::create_dir_all(&agent_src).unwrap();
// Source dir exists but has no scripts.
let data_dir = tmp.path().join("data");
let err = stage_hook_scripts_in(&agent_src, "stage-empty-src", &data_dir)
.expect_err("zero scripts should be an error, not a silent success");
assert!(format!("{err:#}").contains("no hook scripts"));
}
#[test]
fn hook_source_candidates_include_native_package_dir() {
let candidates = hook_source_candidates(
+32
View File
@@ -107,6 +107,38 @@ pub async fn run_lint(
let candidates = reader.decay_candidates(workspace_id, project_id).await?;
let mut findings = rule_based_findings(&candidates);
// Dangling cross-project links: a `[[project:path]]` dependency that does
// not resolve. A broken inter-project edge is high-signal — surface it
// even on the zero-LLM path.
for dangling in reader
.dangling_cross_project_links(workspace_id, project_id)
.await?
{
let target = match &dangling.workspace {
Some(ws) => format!("{ws}/{}:{}", dangling.project, dangling.path),
None => format!("{}:{}", dangling.project, dangling.path),
};
let message = if dangling.project_exists {
format!(
"Page {} links to {} but that page does not exist in project `{}` \
(missing, renamed, or deleted) — a broken cross-project dependency",
dangling.from_path, target, dangling.project,
)
} else {
format!(
"Page {} links to {} but project `{}` does not exist (typo or wrong name)",
dangling.from_path, target, dangling.project,
)
};
findings.push(LintFinding {
kind: "broken_link".into(),
severity: "warning".into(),
message,
pages: vec![dangling.from_path],
detail: None,
});
}
if use_llm && let Some(provider) = llm {
match contradiction_pass(
provider.clone(),
+1 -1
View File
@@ -89,7 +89,7 @@ id_newtype!(pub HandoffId, "Identifier for a cross-agent handoff record.");
/// Always uses `/` as the separator (POSIX-style), normalised on construction.
/// Never starts with a slash; never contains `..` or `.` components. This
/// invariant lets the store treat paths as flat keys without re-validating.
#[derive(Clone, PartialEq, Eq, Hash, Serialize, Deserialize)]
#[derive(Clone, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)]
#[serde(transparent)]
pub struct PagePath(String);
+1 -1
View File
@@ -27,6 +27,6 @@ pub use ids::{
AgentKind, HandoffId, ObservationId, PageId, PagePath, ProjectId, SessionId, WorkspaceId,
};
pub use observation::{NewObservation, NewSession, Observation, ObservationKind};
pub use page::{NewPage, Page, Tier};
pub use page::{LinkTarget, NewPage, Page, Tier};
pub use routing_snippet::{MARKER_END, MARKER_START, SNIPPET_BODY, full_block};
pub use sanitize::{SanitizeConfig, Sanitized, Sanitizer};
+47 -3
View File
@@ -52,9 +52,53 @@ pub struct NewPage {
pub pinned: bool,
/// Outgoing links discovered in the page body.
///
/// The store resolves these against the latest page rows in the same
/// project and keeps unresolved forward links as `to_page_id = NULL`.
pub links: Vec<PagePath>,
/// The store resolves these against the latest page rows in the target
/// project (the source's own project for bare links, or the named
/// project for a cross-project `[[project:path]]` link) and keeps
/// unresolved forward links as `to_page_id = NULL`.
pub links: Vec<LinkTarget>,
}
/// A link target discovered in a page body.
///
/// A bare `[[path]]` / `[label](path)` resolves within the source page's
/// own project (`workspace`/`project` both `None`). A `[[project:path]]`
/// link names a sibling project in the same workspace; a
/// `[[workspace/project:path]]` link crosses workspaces. Making
/// cross-project dependencies explicit edges is what turns the per-project
/// wikis into one graph.
#[derive(Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
pub struct LinkTarget {
/// Cross-workspace qualifier. `None` = same workspace as the source.
pub workspace: Option<String>,
/// Cross-project qualifier. `None` = same project as the source.
pub project: Option<String>,
/// Wiki path within the target project (root-relative).
pub path: PagePath,
}
impl LinkTarget {
/// A link resolving within the source page's own project.
#[must_use]
pub fn local(path: PagePath) -> Self {
Self {
workspace: None,
project: None,
path,
}
}
/// Whether this link names a different project than its source.
#[must_use]
pub fn is_cross_project(&self) -> bool {
self.project.is_some()
}
}
impl From<PagePath> for LinkTarget {
fn from(path: PagePath) -> Self {
Self::local(path)
}
}
/// Materialised view of a page row.
+21 -1
View File
@@ -60,7 +60,7 @@ match the intent to the tool. They do not need to name the tool.
| User says / situation | Tool |
|---|---|
| "have we discussed X?" / "search memory for Y" / before proposing architecture | `memory_query` |
| "have we discussed X?" / "search memory for Y" / before proposing architecture | `memory_query` (current project; `scopes` for named siblings; `global=true` to search every project) |
| "what's been going on" / "show recent activity" (light) | `memory_recent` |
| "is ai-memory healthy?" / "how big is the wiki?" | `memory_status` |
| "give me the stats" / structured snapshot for the agent to consume | `memory_briefing` |
@@ -80,6 +80,26 @@ going on" use case — it returns a prose digest whose verbosity
scales automatically to how long it's been since the last activity
(< 1 h → one line; > 30 days → full catchup).
### When the current project comes up empty — broaden the search
`memory_query` searches only the **current** project by default. If a
search comes back empty or thin, the knowledge may live in a **sibling
project** — shared `infra`, `ops`, or a related app. Don't conclude
"we never recorded it" after a single project misses; broaden instead:
- **Know which projects to check?** Re-run with explicit `scopes`, e.g.
`scopes: [{ "workspace": "default", "project": "infra" }]`.
- **Don't know where it lives?** Pass `global=true` to search every
project in every workspace at once. Each hit is annotated with its
workspace + project so you can tell where it came from. `global=true`
cannot be combined with `scopes`/`project`/`workspace`.
`memory_query` returns **snippets, not full page bodies** — an empty or
short snippet does **not** mean the page is empty (a large page can
match outside the snippet window). To read the whole page, use
`memory_read_page` (by `path`, or pass a `query` to fetch the top hit's
full body).
### When you write a project rule, write it here
If you're about to write a durable project rule ("always X", "never
+452 -10
View File
@@ -47,7 +47,10 @@ the conversation calls for them:\n\
\n\
- `memory_query` — when the user references prior work you don't \
recognise, or asks 'have we done / discussed X', or you're about \
to propose architecture (always check first).\n\
to propose architecture (always check first). Defaults to the \
current project; pass `scopes` to search named sibling projects, \
or `global=true` to search EVERY project at once when you don't \
know where the knowledge lives.\n\
- `memory_recent` — at session start, or when the user asks 'what's \
been going on lately'. Returns the N most-recent pages.\n\
- `memory_status` — when the user asks 'is ai-memory healthy' or \
@@ -95,6 +98,22 @@ the conversation calls for them:\n\
in the right rules file (Claude Code → CLAUDE.md, Codex / \
OpenCode / Cursor / Gemini → AGENTS.md).\n\
\n\
**When the current project comes up empty, broaden — don't stop.** \
`memory_query` searches only ONE project (the current one) by default. \
If a query returns nothing useful, the knowledge may live in a SIBLING \
project — shared `infra`, `ops`, or a related app. Two ways to \
broaden: (a) re-run with explicit `scopes: [{workspace, project}]` \
when you know which projects to check; (b) pass `global=true` to \
search EVERY project in EVERY workspace at once when you don't know \
where the knowledge lives — each hit then carries its workspace + \
project name. `global=true` cannot be combined with \
`scopes`/`project`/`workspace`. Don't conclude 'we never recorded \
it' after one project misses. Note also that `memory_query` returns \
SNIPPETS, not full page bodies — an empty or short snippet does NOT \
mean the page is empty (a large page can match outside the snippet \
window); to read the whole page use `memory_read_page` (by `path`, \
or a `query` for the top hit's body).\n\
\n\
The routing snippet this very text comes from can also be installed \
into the project's CLAUDE.md / AGENTS.md so the guidance survives \
across sessions. From the agent: ask 'install ai-memory routing'. \
@@ -174,6 +193,13 @@ struct QueryArgs {
/// knowledge. Cannot be combined with `workspace`/`project`.
#[serde(default)]
scopes: Vec<MemoryScopeArg>,
/// Search EVERY project in every workspace in one call (cross-project
/// global search). Use when you don't know which project holds the
/// knowledge — e.g. shared infra/ops notes. When true, omit
/// `project`/`workspace`/`scopes`; each hit is annotated with its
/// workspace + project so you can tell where it came from.
#[serde(default)]
global: Option<bool>,
}
#[derive(Debug, Serialize, Deserialize, schemars::JsonSchema)]
@@ -213,6 +239,10 @@ struct MemoryQueryResponse {
hits: Vec<ai_memory_store::PageHit>,
#[serde(skip_serializing_if = "Vec::is_empty")]
raw_hits: Vec<ai_memory_store::ObservationHit>,
/// Populated only by a `global=true` query: cross-project hits, each
/// carrying its workspace + project name.
#[serde(skip_serializing_if = "Vec::is_empty")]
global_hits: Vec<ai_memory_store::PageHitWithMeta>,
}
#[derive(Debug, Serialize)]
@@ -225,6 +255,14 @@ struct SweepArgs {
/// If true, preview only. Default false.
#[serde(default)]
dry_run: Option<bool>,
/// Project to sweep. Omit to target the project you're currently working
/// in (resolved from recent hook activity). **Omit unless the user
/// explicitly names a *different* project.**
#[serde(default)]
project: Option<String>,
/// Workspace the project lives in. Omit for the current workspace.
#[serde(default)]
workspace: Option<String>,
}
#[derive(Debug, Serialize, Deserialize, schemars::JsonSchema)]
@@ -237,6 +275,14 @@ struct LintArgs {
/// fast rule-based checks. Default false.
#[serde(default)]
no_llm: Option<bool>,
/// Project to audit. Omit to target the project you're currently working
/// in (resolved from recent hook activity). **Omit unless the user
/// explicitly names a *different* project.**
#[serde(default)]
project: Option<String>,
/// Workspace the project lives in. Omit for the current workspace.
#[serde(default)]
workspace: Option<String>,
}
#[derive(Debug, Serialize, Deserialize, schemars::JsonSchema)]
@@ -369,9 +415,16 @@ struct WritePageArgs {
#[serde(default)]
pinned: bool,
/// Project to write into. Omit to target the project you're currently
/// working in (resolved from recent hook activity). **Omit unless the user explicitly names a *different* project.**
/// working in (resolved from recent hook activity). When set to a name
/// that doesn't exist yet, the project is **created** — so writes always
/// land where you asked, never silently in the current project. **Omit
/// unless the user explicitly names a *different* project.**
#[serde(default)]
project: Option<String>,
/// Workspace to write into. Only honoured together with an explicit
/// `project`; created if it doesn't exist. Omit for the current workspace.
#[serde(default)]
workspace: Option<String>,
}
#[tool_router]
@@ -491,6 +544,40 @@ impl AiMemoryServer {
Ok((workspace_id, project_id))
}
/// Resolve the target for a WRITE, **creating** the workspace/project when
/// an explicit name doesn't exist yet. Distinct from [`Self::effective_ids`]
/// (find-only, for reads): a write to a named project must land there, not
/// silently fall back to the current project. With no explicit `project`,
/// the active-project-wins behaviour is preserved (issue #2).
async fn write_target_ids(
&self,
explicit_workspace: Option<&str>,
explicit_project: Option<&str>,
) -> Result<(WorkspaceId, ProjectId), McpError> {
let Some(project) = trimmed_opt(explicit_project) else {
// No explicit project → current project (hook-published active, or
// the baked default). Explicit workspace alone has nothing to scope.
return Ok(self
.active_project
.get()
.unwrap_or((self.workspace_id, self.project_id)));
};
let workspace_id = match trimmed_opt(explicit_workspace) {
Some(name) => self
.writer
.get_or_create_workspace(name.to_string())
.await
.map_err(|e| McpError::internal_error(e.to_string(), None))?,
None => self.workspace_id,
};
let project_id = self
.writer
.get_or_create_project(workspace_id, project.to_string(), None)
.await
.map_err(|e| McpError::internal_error(e.to_string(), None))?;
Ok((workspace_id, project_id))
}
async fn resolve_query_scopes(
&self,
scopes: &[MemoryScopeArg],
@@ -626,12 +713,42 @@ impl AiMemoryServer {
Returns up to `limit` pages with HTML-marked snippets and a rank \
score (lower rank = better match). Only latest page versions. \
If compiled wiki search misses, `raw_hits` contains bounded raw \
observation fallback matches.")]
observation fallback matches. Set `global=true` to search EVERY \
project at once (cross-project) when you don't know which project \
holds the knowledge — each hit then carries its workspace + \
project name.")]
async fn memory_query(
&self,
Parameters(args): Parameters<QueryArgs>,
) -> Result<CallToolResult, McpError> {
let limit = args.limit.unwrap_or(self.default_limit).clamp(1, 100);
if args.global.unwrap_or(false) {
if !args.scopes.is_empty()
|| args
.workspace
.as_deref()
.is_some_and(|s| !s.trim().is_empty())
|| args
.project
.as_deref()
.is_some_and(|s| !s.trim().is_empty())
{
return Err(McpError::internal_error(
"global cannot be combined with workspace/project/scopes",
None,
));
}
let global_hits = self
.reader
.search_pages_with_meta(args.query.clone(), limit)
.await
.map_err(|e| McpError::internal_error(e.to_string(), None))?;
return ok_json(&MemoryQueryResponse {
hits: Vec::new(),
raw_hits: Vec::new(),
global_hits,
});
}
if !args.scopes.is_empty()
&& (args
.workspace
@@ -699,7 +816,11 @@ impl AiMemoryServer {
} else {
Vec::new()
};
let response = MemoryQueryResponse { hits, raw_hits };
let response = MemoryQueryResponse {
hits,
raw_hits,
global_hits: Vec::new(),
};
ok_json(&response)
}
@@ -739,11 +860,14 @@ impl AiMemoryServer {
&self,
Parameters(args): Parameters<SweepArgs>,
) -> Result<CallToolResult, McpError> {
let (ws, proj) = self
.effective_ids_for_read_args(args.workspace.as_deref(), args.project.as_deref())
.await?;
let report = run_sweep(
&self.reader,
&self.writer,
self.workspace_id,
self.project_id,
ws,
proj,
&self.decay_params,
args.dry_run.unwrap_or(false),
)
@@ -767,12 +891,15 @@ impl AiMemoryServer {
None,
));
};
let (ws, proj) = self
.effective_ids_for_read_args(args.workspace.as_deref(), args.project.as_deref())
.await?;
let report = run_lint(
&self.reader,
wiki,
self.llm.as_ref(),
self.workspace_id,
self.project_id,
ws,
proj,
args.dry_run.unwrap_or(false),
!args.no_llm.unwrap_or(false),
)
@@ -843,7 +970,9 @@ impl AiMemoryServer {
.map_err(|_| McpError::internal_error(format!("unknown tier '{tier_name}'"), None))?;
let path = PagePath::new(args.path.clone())
.map_err(|e| McpError::internal_error(format!("invalid path: {e}"), None))?;
let (ws, proj) = self.effective_ids(args.project.as_deref()).await;
let (ws, proj) = self
.write_target_ids(args.workspace.as_deref(), args.project.as_deref())
.await?;
let mut fm = serde_json::Map::new();
if let Some(title) = &args.title {
@@ -1490,6 +1619,50 @@ mod tests {
}
}
#[test]
fn prompts_teach_cross_project_search_strategy() {
// Regression: a single-project miss must not read as "never recorded".
// Both surfaces must point the agent at `scopes` **and** at
// `global=true` (the two broadening modes), warn that query returns
// snippets (not full page bodies), and NOT contain the contradictory
// legacy "no global mode" phrasing that briefly shipped in #56.
// (Learned the hard way when cluster-access info lived in a sibling
// `infra` project.)
for prompt in [MEMORY_INSTRUCTIONS, ai_memory_core::SNIPPET_BODY] {
assert!(
prompt.contains("scopes"),
"prompt must teach broadening via `scopes`"
);
assert!(
prompt.contains("global=true") || prompt.contains("global = true"),
"prompt must also teach broadening via `global=true`"
);
assert!(
prompt.contains("sibling") || prompt.contains("SIBLING"),
"prompt must mention knowledge can live in a sibling project"
);
assert!(
prompt.contains("snippet") || prompt.contains("SNIPPET"),
"prompt must warn that query returns snippets, not full bodies"
);
// Guard against the contradiction: standalone prose must not say
// a global mode doesn't exist when the bullet/table-row above it
// advertises `global=true`.
let no_global_phrases = [
"no global \"search everything\" mode",
"NO global 'search everything' mode",
"no global 'search everything' mode",
"NO global \"search everything\" mode",
];
for phrase in no_global_phrases {
assert!(
!prompt.contains(phrase),
"prompt must not contain the contradictory phrase {phrase:?}"
);
}
}
}
#[test]
fn prompts_route_permanent_annotations_to_write_page_not_handoff() {
for prompt in [MEMORY_INSTRUCTIONS, ai_memory_core::SNIPPET_BODY] {
@@ -1558,6 +1731,7 @@ mod tests {
project: None,
scopes: Vec::new(),
workspace: None,
global: None,
}))
.await
.unwrap();
@@ -1606,6 +1780,7 @@ mod tests {
project: None,
scopes: Vec::new(),
workspace: None,
global: None,
}))
.await
.unwrap();
@@ -1660,6 +1835,7 @@ mod tests {
project: Some("unit-testing".into()),
scopes: Vec::new(),
workspace: Some("practice".into()),
global: None,
}))
.await
.unwrap();
@@ -1757,6 +1933,7 @@ mod tests {
},
],
workspace: None,
global: None,
}))
.await
.unwrap();
@@ -1774,6 +1951,93 @@ mod tests {
assert!(!text.contains("hidden.md"), "unexpected hidden hit: {text}");
}
#[tokio::test]
async fn memory_query_global_searches_all_projects() {
let (_tmp, store, server, ws, _pj) = setup_server().await;
let other = store
.writer
.get_or_create_project(ws, "infra", None)
.await
.unwrap();
let other_ws = store.writer.get_or_create_workspace("ops").await.unwrap();
let third = store
.writer
.get_or_create_project(other_ws, "runbooks", None)
.await
.unwrap();
for (w, p, path, body) in [
(ws, other, "cluster.md", "global_token lives in infra"),
(
other_ws,
third,
"deploy.md",
"global_token lives in ops runbooks",
),
] {
store
.writer
.upsert_page(NewPage {
workspace_id: w,
project_id: p,
path: PagePath::new(path).unwrap(),
title: path.into(),
body: body.into(),
tier: Tier::Semantic,
frontmatter_json: serde_json::json!({}),
pinned: false,
links: Vec::new(),
})
.await
.unwrap();
}
let result = server
.memory_query(Parameters(QueryArgs {
query: "global_token".into(),
limit: Some(10),
project: None,
scopes: Vec::new(),
workspace: None,
global: Some(true),
}))
.await
.unwrap();
let text = result
.content
.first()
.and_then(|c| c.as_text())
.map(|t| t.text.clone())
.unwrap();
// Both projects (across two workspaces) surface in one global call,
// each annotated with its project name.
assert!(text.contains("cluster.md"), "expected infra hit: {text}");
assert!(text.contains("deploy.md"), "expected ops hit: {text}");
assert!(
text.contains("infra"),
"hit must carry project name: {text}"
);
assert!(text.contains("global_hits"), "global hits field: {text}");
}
#[tokio::test]
async fn memory_query_global_rejects_explicit_scope() {
let (_tmp, _store, server, _ws, _pj) = setup_server().await;
let err = server
.memory_query(Parameters(QueryArgs {
query: "x".into(),
limit: Some(5),
project: Some("product".into()),
scopes: Vec::new(),
workspace: None,
global: Some(true),
}))
.await;
assert!(
err.is_err(),
"global must not combine with project/workspace/scopes"
);
}
#[tokio::test]
async fn memory_status_returns_counts() {
let (_tmp, _store, server, _ws, _pj) = setup_server().await;
@@ -1956,6 +2220,7 @@ mod tests {
tags: vec!["finance".into()],
pinned: true,
project: None,
workspace: None,
}),
rmcp::handler::server::tool::Extension(parts),
)
@@ -2006,7 +2271,6 @@ mod tests {
let wiki = Wiki::new(tmp.path(), store.writer.clone()).unwrap();
let server = AiMemoryServer::new(store.reader.clone(), store.writer.clone(), ws, proj)
.with_wiki(wiki);
let parts = || {
axum::http::Request::builder()
.uri("/mcp")
@@ -2027,6 +2291,7 @@ mod tests {
tags: vec![],
pinned: false,
project: None,
workspace: None,
}),
rmcp::handler::server::tool::Extension(parts()),
)
@@ -2077,6 +2342,92 @@ mod tests {
);
}
#[tokio::test]
async fn memory_write_page_creates_explicit_project() {
// Bug B regression: an explicit `project` that doesn't exist yet must
// be created and written to — NOT silently land in the current project.
let tmp = TempDir::new().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default")
.await
.unwrap();
let baked = store
.writer
.get_or_create_project(ws, "scratch", None)
.await
.unwrap();
let wiki = Wiki::new(tmp.path(), store.writer.clone()).unwrap();
let server = AiMemoryServer::new(store.reader.clone(), store.writer.clone(), ws, baked)
.with_wiki(wiki);
let parts = || {
axum::http::Request::builder()
.uri("/mcp")
.method("POST")
.body(())
.unwrap()
.into_parts()
.0
};
server
.memory_write_page(
Parameters(WritePageArgs {
path: "notes/elsewhere.md".into(),
body: "lands in `other`, not `scratch`".into(),
title: None,
tier: Some("semantic".into()),
tags: vec![],
pinned: false,
project: Some("other".into()),
workspace: None,
}),
rmcp::handler::server::tool::Extension(parts()),
)
.await
.unwrap();
// Visible in `other` (created), absent from the baked `scratch`.
let in_other = server
.memory_recent(Parameters(RecentArgs {
limit: Some(5),
project: Some("other".into()),
workspace: None,
}))
.await
.unwrap();
let other_text = in_other
.content
.first()
.and_then(|c| c.as_text())
.map(|t| t.text.clone())
.unwrap();
assert!(
other_text.contains("notes/elsewhere.md"),
"explicit project must be created + written; got {other_text}"
);
let in_scratch = server
.memory_recent(Parameters(RecentArgs {
limit: Some(5),
project: None,
workspace: None,
}))
.await
.unwrap();
let scratch_text = in_scratch
.content
.first()
.and_then(|c| c.as_text())
.map(|t| t.text.clone())
.unwrap();
assert!(
!scratch_text.contains("notes/elsewhere.md"),
"write must not leak into the current project; got {scratch_text}"
);
}
/// `memory_handoff_begin` must resolve the same project as
/// `memory_briefing` when hooks publish `ActiveProject` (issue #2).
#[tokio::test]
@@ -2213,6 +2564,8 @@ mod tests {
.memory_lint(Parameters(LintArgs {
dry_run: Some(true),
no_llm: None,
project: None,
workspace: None,
}))
.await
.expect_err("must reject when wiki is not attached");
@@ -2226,6 +2579,93 @@ mod tests {
);
}
#[tokio::test]
async fn memory_forget_sweep_targets_the_explicit_project() {
// Bug C regression: sweep must evaluate the project named in args (or
// the session's active project), NOT the baked default. An episodic
// page in `audited` is a sweep candidate only when the sweep points
// there — never when it runs against the baked `scratch`.
let tmp = TempDir::new().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default")
.await
.unwrap();
let baked = store
.writer
.get_or_create_project(ws, "scratch", None)
.await
.unwrap();
let wiki = Wiki::new(tmp.path(), store.writer.clone()).unwrap();
let server = AiMemoryServer::new(store.reader.clone(), store.writer.clone(), ws, baked)
.with_wiki(wiki);
let parts = || {
axum::http::Request::builder()
.uri("/mcp")
.method("POST")
.body(())
.unwrap()
.into_parts()
.0
};
server
.memory_write_page(
Parameters(WritePageArgs {
path: "log/ep.md".into(),
body: "episodic note".into(),
title: None,
tier: Some("episodic".into()),
tags: vec![],
pinned: false,
project: Some("audited".into()),
workspace: None,
}),
rmcp::handler::server::tool::Extension(parts()),
)
.await
.unwrap();
let sweep_count = |args: SweepArgs| {
let server = &server;
async move {
let out = server.memory_forget_sweep(Parameters(args)).await.unwrap();
let text = out
.content
.first()
.and_then(|c| c.as_text())
.map(|t| t.text.clone())
.unwrap();
serde_json::from_str::<serde_json::Value>(&text).unwrap()["candidates_evaluated"]
.as_u64()
.unwrap()
}
};
let audited = sweep_count(SweepArgs {
dry_run: Some(true),
project: Some("audited".into()),
workspace: None,
})
.await;
assert!(
audited >= 1,
"sweep of the named project must evaluate its episodic page, got {audited}"
);
let baked = sweep_count(SweepArgs {
dry_run: Some(true),
project: None,
workspace: None,
})
.await;
assert_eq!(
baked, 0,
"sweep of the baked project must not see another project's page, got {baked}"
);
}
/// `memory_handoff_accept` with no pending handoff returns a
/// happy-path `{"handoff": null}` payload (NOT an error). This
/// is the documented contract — the agent can call accept on
@@ -2268,6 +2708,7 @@ mod tests {
project: None,
scopes: Vec::new(),
workspace: None,
global: None,
}))
.await
.expect("oversized limit should be clamped, not refused");
@@ -2295,6 +2736,7 @@ mod tests {
project: None,
scopes: Vec::new(),
workspace: None,
global: None,
}))
.await;
// Either a tidy 0-hit Ok (FTS5 is occasionally lenient) or
@@ -0,0 +1,70 @@
-- Accent-insensitive full-text search (Portuguese-friendly).
--
-- The FTS5 tables were created with `unicode61 tokenchars '/_-'`, which keeps
-- diacritics: a query for "descricao" never matches stored "descrição". The
-- tokenizer is fixed at CREATE time, so the only way to change it is to drop
-- and recreate the FTS table (+ its sync triggers) and rebuild from the
-- content table. `remove_diacritics 2` folds diacritics across the full
-- Unicode range, so "descricao"/"descrição", "sessao"/"sessão", etc. unify.
--
-- Both FTS tables are contentless (`content='pages'` / `content='observations'`),
-- so the source rows are untouched; only the derived index is rebuilt.
-- ── pages_fts ────────────────────────────────────────────────────────────
DROP TRIGGER pages_fts_ai;
DROP TRIGGER pages_fts_ad;
DROP TRIGGER pages_fts_au;
DROP TABLE pages_fts;
CREATE VIRTUAL TABLE pages_fts USING fts5(
title, body,
content='pages',
content_rowid='rowid',
tokenize="unicode61 remove_diacritics 2 tokenchars '/_-'"
);
INSERT INTO pages_fts(pages_fts) VALUES('rebuild');
CREATE TRIGGER pages_fts_ai AFTER INSERT ON pages BEGIN
INSERT INTO pages_fts(rowid, title, body)
VALUES (new.rowid, new.title, new.body);
END;
CREATE TRIGGER pages_fts_ad AFTER DELETE ON pages BEGIN
INSERT INTO pages_fts(pages_fts, rowid, title, body)
VALUES ('delete', old.rowid, old.title, old.body);
END;
-- Matches the V08 narrowing: only re-index on title/body updates.
CREATE TRIGGER pages_fts_au AFTER UPDATE OF title, body ON pages BEGIN
INSERT INTO pages_fts(pages_fts, rowid, title, body)
VALUES ('delete', old.rowid, old.title, old.body);
INSERT INTO pages_fts(rowid, title, body)
VALUES (new.rowid, new.title, new.body);
END;
-- ── observations_fts ──────────────────────────────────────────────────────
DROP TRIGGER observations_fts_ai;
DROP TRIGGER observations_fts_ad;
DROP TRIGGER observations_fts_au;
DROP TABLE observations_fts;
CREATE VIRTUAL TABLE observations_fts USING fts5(
title, body,
content='observations',
content_rowid='rowid',
tokenize="unicode61 remove_diacritics 2 tokenchars '/_-'"
);
INSERT INTO observations_fts(observations_fts) VALUES('rebuild');
CREATE TRIGGER observations_fts_ai AFTER INSERT ON observations BEGIN
INSERT INTO observations_fts(rowid, title, body)
VALUES (new.rowid, new.title, new.body);
END;
CREATE TRIGGER observations_fts_ad AFTER DELETE ON observations BEGIN
INSERT INTO observations_fts(observations_fts, rowid, title, body)
VALUES ('delete', old.rowid, old.title, old.body);
END;
CREATE TRIGGER observations_fts_au AFTER UPDATE ON observations BEGIN
INSERT INTO observations_fts(observations_fts, rowid, title, body)
VALUES ('delete', old.rowid, old.title, old.body);
INSERT INTO observations_fts(rowid, title, body)
VALUES (new.rowid, new.title, new.body);
END;
@@ -0,0 +1,35 @@
-- Cross-project links.
--
-- A `[[project:path]]` / `[[workspace/project:path]]` wikilink carries an
-- explicit scope; `to_workspace` / `to_project` are NULL for the common case
-- (a link within the source page's own project). The scope joins the PRIMARY
-- KEY so one page can link to the same `to_path` in two different projects
-- without colliding. `to_page_id` stays a global PageId, so a resolved
-- cross-project link already surfaces as a backlink on its target with no
-- query change — the per-project wikis become one graph.
--
-- SQLite cannot alter a PRIMARY KEY in place, so rebuild the table. No other
-- table references `links`, so the drop/rename is safe with FKs enabled.
CREATE TABLE links_new (
from_page_id BLOB NOT NULL REFERENCES pages(id) ON DELETE CASCADE,
to_page_id BLOB REFERENCES pages(id) ON DELETE SET NULL,
to_workspace TEXT, -- NULL = source page's workspace
to_project TEXT, -- NULL = source page's project
to_path TEXT NOT NULL, -- root-relative path in the target project
link_type TEXT NOT NULL DEFAULT 'references',
PRIMARY KEY (from_page_id, to_workspace, to_project, to_path, link_type)
);
INSERT INTO links_new (from_page_id, to_page_id, to_workspace, to_project, to_path, link_type)
SELECT from_page_id, to_page_id, NULL, NULL, to_path, link_type FROM links;
DROP TABLE links;
ALTER TABLE links_new RENAME TO links;
-- Recreate the indexes the dropped table carried (V01/V07/V08) ...
CREATE INDEX idx_links_to ON links(to_page_id) WHERE to_page_id IS NOT NULL;
CREATE INDEX idx_links_unresolved_path ON links(to_path) WHERE to_page_id IS NULL;
CREATE INDEX idx_links_to_path ON links(to_path);
-- ... plus one for scoped (cross-project) resolution + refresh.
CREATE INDEX idx_links_scope ON links(to_project, to_path);
+94 -2
View File
@@ -10,13 +10,31 @@
///
/// Returns an empty string when `raw` is empty/whitespace-only; callers
/// should skip the SQL query in that case.
///
/// Bare multi-word queries are joined with **`OR`**, not the FTS5 default
/// (`AND`). A natural-language query like "cross project search strategy"
/// otherwise requires every word to co-occur in one page — near-zero recall
/// for anything but single keywords. With `OR` + bm25 ranking (callers
/// `ORDER BY rank`), the best-matching pages still surface first. When the
/// caller supplies explicit FTS5 syntax (`OR` / `AND` / `NOT` / `NEAR` /
/// quoted phrases / parens) we preserve it verbatim instead.
#[must_use]
pub fn prepare_fts5_query(raw: &str) -> String {
let explicit_syntax = raw.contains('"')
|| raw.contains('(')
|| raw.contains(')')
|| raw
.split_whitespace()
.any(|t| matches!(t, "OR" | "AND" | "NOT" | "NEAR"));
let tokens: Vec<String> = raw
.split_whitespace()
.flat_map(prepare_fts5_token)
.collect();
tokens.join(" ")
if tokens.is_empty() {
return String::new();
}
let separator = if explicit_syntax { " " } else { " OR " };
tokens.join(separator)
}
fn prepare_fts5_token(token: &str) -> Vec<String> {
@@ -58,8 +76,33 @@ mod tests {
#[test]
fn colon_is_not_column_syntax() {
// Bare multi-word → OR-joined (no explicit operator present).
let q = prepare_fts5_query("pick: handoff ai-memory");
assert_eq!(q, "\"pick\" handoff \"ai-memory\"");
assert_eq!(q, "\"pick\" OR handoff OR \"ai-memory\"");
}
#[test]
fn bare_multi_word_is_or_joined() {
// The recall fix: every word no longer has to co-occur.
assert_eq!(
prepare_fts5_query("cross project search strategy"),
"cross OR project OR search OR strategy"
);
}
#[test]
fn portuguese_accented_terms_or_join_and_keep_accents() {
// PT natural-language query: tokens preserved (accents intact),
// joined with OR so a page matching any term is found.
assert_eq!(
prepare_fts5_query("descrição testes commits"),
"descrição OR testes OR commits"
);
}
#[test]
fn single_word_has_no_or() {
assert_eq!(prepare_fts5_query("handoff"), "handoff");
}
#[test]
@@ -78,6 +121,55 @@ mod tests {
assert_eq!(prepare_fts5_query("quick OR slow"), "quick OR slow");
}
/// AND is the FTS5 default but operators can be explicit — when the
/// caller writes one, the OR-join must NOT mangle it into
/// `foo OR AND OR bar`. Same for NOT and NEAR. (The escape hatch from
/// the broad-recall default is what makes the OR-join safe to land.)
#[test]
fn explicit_and_operator_is_preserved() {
assert_eq!(prepare_fts5_query("foo AND bar"), "foo AND bar");
}
#[test]
fn explicit_not_operator_is_preserved() {
assert_eq!(prepare_fts5_query("foo NOT bar"), "foo NOT bar");
}
#[test]
fn explicit_near_operator_is_preserved() {
assert_eq!(prepare_fts5_query("foo NEAR bar"), "foo NEAR bar");
}
/// A query containing a quoted phrase is treated as explicit FTS5
/// syntax — `"exact phrase" baz` must not become
/// `"exact" OR "phrase" OR baz` (which destroys the phrase semantics).
/// The exact assertion is "space-joined, not OR-joined"; what the
/// individual tokens look like after `prepare_fts5_token` is a
/// separate concern (and unchanged from pre-#58 behaviour).
#[test]
fn quoted_phrase_query_is_not_or_joined() {
let q = prepare_fts5_query("\"exact phrase\" baz");
assert!(
!q.contains(" OR "),
"explicit quoted-phrase query must not get OR-joined; got {q}"
);
}
/// Same escape-hatch logic for parenthesised sub-expressions —
/// `(foo OR bar) AND baz` must survive unmangled.
#[test]
fn parenthesised_query_is_not_or_joined() {
let q = prepare_fts5_query("(foo OR bar) AND baz");
assert!(
!q.contains("OR (foo"),
"parens detection must skip OR-join entirely; got {q}"
);
assert!(
q.contains("AND"),
"explicit AND inside parens query must survive; got {q}"
);
}
#[test]
fn known_columns_are_preserved() {
assert_eq!(prepare_fts5_query("title:handoff"), "title:handoff");
+110 -3
View File
@@ -93,8 +93,8 @@ impl Store {
mod tests {
use super::*;
use ai_memory_core::{
AgentKind, NewObservation, NewPage, NewSession, ObservationId, ObservationKind, PagePath,
ProjectId, SessionId, Tier, WorkspaceId,
AgentKind, LinkTarget, NewObservation, NewPage, NewSession, ObservationId, ObservationKind,
PagePath, ProjectId, SessionId, Tier, WorkspaceId,
};
use rusqlite::{Connection, params};
use tempfile::TempDir;
@@ -113,6 +113,76 @@ mod tests {
}
}
#[tokio::test]
async fn cross_project_links_surface_in_graph_briefing_and_lint() {
let tmp = TempDir::new().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default")
.await
.unwrap();
let app = store
.writer
.get_or_create_project(ws, "app", None)
.await
.unwrap();
let infra = store
.writer
.get_or_create_project(ws, "infra", None)
.await
.unwrap();
// Target page in `infra`, then a page in `app` that depends on it
// plus a dangling link to a non-existent project.
store
.writer
.upsert_page(sample_page(ws, infra, "runbooks/02.md", "the runbook"))
.await
.unwrap();
let mut dep = sample_page(ws, app, "concepts/dep.md", "needs infra + a typo");
dep.links = vec![
LinkTarget {
workspace: None,
project: Some("infra".into()),
path: PagePath::new("runbooks/02.md").unwrap(),
},
LinkTarget {
workspace: None,
project: Some("nope".into()),
path: PagePath::new("ghost.md").unwrap(),
},
];
store.writer.upsert_page(dep).await.unwrap();
// Graph: exactly one resolved cross-project edge, app -> infra.
let edges = store.reader.cross_project_edges(None).await.unwrap();
assert_eq!(edges.len(), 1, "one resolved cross-project edge");
assert_eq!(edges[0].from_project, "app");
assert_eq!(edges[0].to_project, "infra");
// Briefing degree: app depends on 1 project; infra has 1 dependent.
let app_brief = store.reader.briefing_for_project(ws, app, 5).await.unwrap();
assert_eq!(app_brief.cross_project_dependencies, 1);
assert_eq!(app_brief.cross_project_dependents, 0);
let infra_brief = store
.reader
.briefing_for_project(ws, infra, 5)
.await
.unwrap();
assert_eq!(infra_brief.cross_project_dependents, 1);
// Lint: the dangling link to project `nope` is reported as unknown.
let dangling = store
.reader
.dangling_cross_project_links(ws, app)
.await
.unwrap();
assert_eq!(dangling.len(), 1, "only the unresolved `nope` link");
assert_eq!(dangling[0].project, "nope");
assert!(!dangling[0].project_exists);
}
#[tokio::test]
async fn open_and_upsert_page() {
let tmp = TempDir::new().unwrap();
@@ -414,6 +484,43 @@ mod tests {
);
}
#[tokio::test]
async fn search_is_accent_insensitive() {
// V13: an accent-free query matches accented stored text (PT-friendly).
let tmp = TempDir::new().unwrap();
let store = Store::open(tmp.path()).unwrap();
let ws = store
.writer
.get_or_create_workspace("default")
.await
.unwrap();
let proj = store
.writer
.get_or_create_project(ws, "scratch", None)
.await
.unwrap();
store
.writer
.upsert_page(sample_page(
ws,
proj,
"notes/decisao.md",
"a descrição da sessão e a consolidação dos commits",
))
.await
.unwrap();
let hits = store
.reader
.search_pages("descricao sessao".into(), 10)
.await
.unwrap();
assert!(
!hits.is_empty(),
"accent-free query must match accented stored text"
);
}
#[tokio::test]
async fn search_boolean_or_still_works() {
let tmp = TempDir::new().unwrap();
@@ -498,7 +605,7 @@ mod tests {
.await
.unwrap();
let mut source = sample_page(ws, proj, "source.md", "needle source content");
source.links = vec![PagePath::new("target.md").unwrap()];
source.links = vec![PagePath::new("target.md").unwrap().into()];
store.writer.upsert_page(source).await.unwrap();
let hits = store
+159 -18
View File
@@ -7,8 +7,8 @@
use std::collections::BTreeSet;
use ai_memory_core::{
AgentKind, HandoffId, NewHandoff, NewObservation, NewPage, NewSession, ObservationId,
ObservationKind, PageId, PagePath, ProjectId, SessionId, WorkspaceId,
AgentKind, HandoffId, LinkTarget, NewHandoff, NewObservation, NewPage, NewSession,
ObservationId, ObservationKind, PageId, PagePath, ProjectId, SessionId, WorkspaceId,
};
/// Summary returned by [`reorg_sessions`] and exposed via
@@ -294,35 +294,84 @@ fn replace_links_in_tx(
)?;
let mut seen = BTreeSet::new();
for to_path in &page.links {
if !seen.insert(to_path.as_str().to_string()) {
for link in &page.links {
let key = (
link.workspace.clone(),
link.project.clone(),
link.path.as_str().to_string(),
);
if !seen.insert(key) {
continue;
}
let to_page_id = latest_page_id_for_path(tx, page, to_path)?;
let to_page_id = latest_page_id_for_link(tx, page, link)?;
let to_page_blob = to_page_id.as_ref().map(|id| &id.as_bytes()[..]);
tx.execute(
"INSERT INTO links (from_page_id, to_page_id, to_path, link_type) \
VALUES (?1, ?2, ?3, 'references')",
params![from_page_id.as_bytes(), to_page_blob, to_path.as_str()],
"INSERT INTO links \
(from_page_id, to_page_id, to_workspace, to_project, to_path, link_type) \
VALUES (?1, ?2, ?3, ?4, ?5, 'references')",
params![
from_page_id.as_bytes(),
to_page_blob,
link.workspace,
link.project,
link.path.as_str(),
],
)?;
}
Ok(())
}
fn latest_page_id_for_path(
/// Resolve a link target to the latest page id it points at, or `None` if the
/// target workspace / project / page does not exist yet (an unresolved forward
/// link). A bare link resolves within the source page's own project; a
/// `[[project:path]]` / `[[workspace/project:path]]` link resolves against the
/// named project (same workspace when only the project is given).
fn latest_page_id_for_link(
tx: &rusqlite::Transaction<'_>,
page: &NewPage,
to_path: &PagePath,
link: &LinkTarget,
) -> StoreResult<Option<PageId>> {
let (workspace_blob, project_blob): (Vec<u8>, Vec<u8>) = match &link.project {
None => (
page.workspace_id.as_bytes().to_vec(),
page.project_id.as_bytes().to_vec(),
),
Some(project_name) => {
let workspace_blob: Vec<u8> = match &link.workspace {
None => page.workspace_id.as_bytes().to_vec(),
Some(workspace_name) => {
let found: Option<Vec<u8>> = tx
.query_row(
"SELECT id FROM workspaces WHERE name = ?1",
params![workspace_name],
|row| row.get(0),
)
.optional()?;
match found {
Some(id) => id,
None => return Ok(None),
}
}
};
let project_blob: Option<Vec<u8>> = tx
.query_row(
"SELECT id FROM projects WHERE workspace_id = ?1 AND name = ?2",
params![workspace_blob, project_name],
|row| row.get(0),
)
.optional()?;
match project_blob {
Some(id) => (workspace_blob, id),
None => return Ok(None),
}
}
};
let bytes: Option<Vec<u8>> = tx
.query_row(
"SELECT id FROM pages \
WHERE workspace_id = ?1 AND project_id = ?2 AND path = ?3 AND is_latest = 1",
params![
page.workspace_id.as_bytes(),
page.project_id.as_bytes(),
to_path.as_str(),
],
params![workspace_blob, project_blob, link.path.as_str()],
|row| row.get(0),
)
.optional()?;
@@ -336,10 +385,13 @@ fn refresh_incoming_links_for_path(
page: &NewPage,
latest_page_id: &PageId,
) -> StoreResult<()> {
// (1) Bare (same-project) links: from_page lives in this page's project and
// the target carries no scope. Repoints all matches (not only unresolved):
// a new page version changes the latest id, so resolved links must follow.
tx.execute(
"UPDATE links \
SET to_page_id = ?1 \
WHERE to_path = ?2 \
WHERE to_project IS NULL AND to_path = ?2 \
AND EXISTS ( \
SELECT 1 FROM pages from_page \
WHERE from_page.id = links.from_page_id \
@@ -353,6 +405,48 @@ fn refresh_incoming_links_for_path(
page.project_id.as_bytes(),
],
)?;
// (2) Cross-project links naming this page's project by name. `to_workspace`
// may be explicit (cross-workspace) or NULL (same workspace as the source).
let project_name: Option<String> = tx
.query_row(
"SELECT name FROM projects WHERE id = ?1",
params![page.project_id.as_bytes()],
|row| row.get(0),
)
.optional()?;
let workspace_name: Option<String> = tx
.query_row(
"SELECT name FROM workspaces WHERE id = ?1",
params![page.workspace_id.as_bytes()],
|row| row.get(0),
)
.optional()?;
if let (Some(project_name), Some(workspace_name)) = (project_name, workspace_name) {
tx.execute(
"UPDATE links \
SET to_page_id = ?1 \
WHERE to_project = ?2 AND to_path = ?3 \
AND ( \
to_workspace = ?4 \
OR ( \
to_workspace IS NULL \
AND EXISTS ( \
SELECT 1 FROM pages from_page \
WHERE from_page.id = links.from_page_id \
AND from_page.workspace_id = ?5 \
) \
) \
)",
params![
latest_page_id.as_bytes(),
project_name,
page.path.as_str(),
workspace_name,
page.workspace_id.as_bytes(),
],
)?;
}
Ok(())
}
@@ -939,7 +1033,7 @@ mod tests {
//! deserve direct coverage so a regression surfaces with a
//! one-line diff instead of a cascading e2e failure.
use super::*;
use ai_memory_core::{NewHandoff, NewPage, NewSession, PagePath, Tier};
use ai_memory_core::{LinkTarget, NewHandoff, NewPage, NewSession, PagePath, Tier};
use rusqlite::Connection;
use tempfile::TempDir;
@@ -1098,7 +1192,7 @@ mod tests {
fn upsert_page_persists_and_resolves_links() {
let (_tmp, mut conn, ws, proj) = fresh_db();
let mut source = page(ws, proj, "concepts/source.md", "see target");
source.links = vec![PagePath::new("decisions/target.md").unwrap()];
source.links = vec![PagePath::new("decisions/target.md").unwrap().into()];
let source_id = upsert_page(&mut conn, &source).unwrap();
let unresolved: i64 = conn
@@ -1127,6 +1221,53 @@ mod tests {
assert_eq!(resolved.as_deref(), Some(&target_id.as_bytes()[..]));
}
/// A `[[infra:runbooks/02.md]]` link from one project resolves to a page
/// in a sibling project once that page exists — the cross-project edge.
#[test]
fn upsert_page_resolves_cross_project_link() {
let (_tmp, mut conn, ws, scratch) = fresh_db();
let infra = get_or_create_project(&mut conn, &ws, "infra", None).unwrap();
let mut source = page(ws, scratch, "concepts/dep.md", "depends on infra runbook");
source.links = vec![LinkTarget {
workspace: None,
project: Some("infra".into()),
path: PagePath::new("runbooks/02.md").unwrap(),
}];
let source_id = upsert_page(&mut conn, &source).unwrap();
// Persisted with the scope, unresolved until the target project's page exists.
let (to_project, resolved): (Option<String>, Option<Vec<u8>>) = conn
.query_row(
"SELECT to_project, to_page_id FROM links \
WHERE from_page_id = ?1 AND to_path = ?2",
params![source_id.as_bytes(), "runbooks/02.md"],
|r| Ok((r.get(0)?, r.get(1)?)),
)
.unwrap();
assert_eq!(to_project.as_deref(), Some("infra"));
assert!(
resolved.is_none(),
"cross-project link is unresolved before the target exists"
);
// Create the target in `infra` → the forward link repoints across projects.
let target_id =
upsert_page(&mut conn, &page(ws, infra, "runbooks/02.md", "the runbook")).unwrap();
let resolved: Option<Vec<u8>> = conn
.query_row(
"SELECT to_page_id FROM links WHERE from_page_id = ?1 AND to_path = ?2",
params![source_id.as_bytes(), "runbooks/02.md"],
|r| r.get(0),
)
.unwrap();
assert_eq!(
resolved.as_deref(),
Some(&target_id.as_bytes()[..]),
"link must resolve across projects once the target lands"
);
}
/// Handoff state machine: insert → Open; accept_handoff → Accepted
/// with accepted_by stamped. Calling accept again must be safe
/// (idempotent at the DB level) because hooks fire-and-forget.
+214 -4
View File
@@ -172,6 +172,13 @@ pub struct BriefingSnapshot {
pub slots: Vec<BriefingPage>,
/// Top-N most-recently-updated `is_latest = 1` pages.
pub recent_pages: Vec<BriefingPage>,
/// Distinct other projects whose pages link INTO this project (who
/// depends on us). Project-scoped briefings only; `0` for
/// workspace/global snapshots.
pub cross_project_dependents: u64,
/// Distinct other projects this project's pages link OUT to (what we
/// depend on). Project-scoped briefings only; `0` otherwise.
pub cross_project_dependencies: u64,
}
/// Trimmed page view for the briefing — path, title, kind, updated_at
@@ -265,6 +272,43 @@ pub struct PageMeta {
pub supersedes: Option<String>,
}
/// One resolved cross-project edge (a link whose endpoints live in
/// different projects). The `/api/v1/graph` endpoint returns these; the UI
/// builds nodes from the endpoints and can aggregate to a project graph.
#[derive(Debug, Clone, Serialize)]
pub struct CrossProjectEdge {
/// Source page workspace.
pub from_workspace: String,
/// Source page project.
pub from_project: String,
/// Source page path.
pub from_path: String,
/// Target page workspace.
pub to_workspace: String,
/// Target page project.
pub to_project: String,
/// Target page path.
pub to_path: String,
}
/// An unresolved cross-project link — a declared dependency on another
/// project's page that does not resolve. Surfaced by `memory_lint`.
#[derive(Debug, Clone, Serialize)]
pub struct DanglingCrossLink {
/// Path of the page that authored the link (in the queried project).
pub from_path: String,
/// Target workspace name (`None` = the source page's own workspace).
pub workspace: Option<String>,
/// Target project name.
pub project: String,
/// Target page path within that project.
pub path: String,
/// Whether the named target project exists at all. `false` →
/// likely a typo / wrong name; `true` → the page is missing or was
/// renamed/deleted in an existing project (a broken dependency).
pub project_exists: bool,
}
/// A page related to another through the link graph — used by the
/// page-view "references / referenced by" panel. Body is omitted; just
/// enough to render a clickable row.
@@ -276,6 +320,11 @@ pub struct RelatedPage {
pub title: String,
/// Semantic kind of the related page.
pub kind: String,
/// Workspace the related page lives in. Lets a backlink from another
/// project be labelled / navigated (the cross-project dependency signal).
pub workspace: String,
/// Project the related page lives in.
pub project: String,
}
/// Resolved outgoing links and incoming back-links for one page.
@@ -1411,6 +1460,8 @@ impl ReaderPool {
rules,
slots,
recent_pages,
cross_project_dependents: 0,
cross_project_dependencies: 0,
})
})
.await
@@ -1542,6 +1593,8 @@ impl ReaderPool {
.into_iter()
.collect::<Result<Vec<_>, _>>()?;
let (cross_project_dependents, cross_project_dependencies) =
cross_project_degree(conn, workspace_id, project_id)?;
Ok(BriefingSnapshot {
counts,
activity_7d,
@@ -1551,6 +1604,8 @@ impl ReaderPool {
rules,
slots,
recent_pages,
cross_project_dependents,
cross_project_dependencies,
})
})
.await
@@ -1768,6 +1823,8 @@ impl ReaderPool {
rules,
slots,
recent_pages,
cross_project_dependents: 0,
cross_project_dependencies: 0,
})
})
.await
@@ -2134,11 +2191,14 @@ impl ReaderPool {
WHEN pg.path LIKE 'gotchas/%' THEN 'gotcha' \
ELSE 'fact' \
END \
) \
), \
ws.name, pr.name \
FROM links l \
JOIN pages pg ON pg.id = l.to_page_id \
JOIN projects pr ON pr.id = pg.project_id \
JOIN workspaces ws ON ws.id = pg.workspace_id \
WHERE l.from_page_id = ?1 AND pg.is_latest = 1 \
ORDER BY pg.path";
ORDER BY ws.name, pr.name, pg.path";
let incoming = "SELECT DISTINCT pg.path, pg.title, \
COALESCE( \
json_extract(pg.frontmatter_json, '$.kind'), \
@@ -2148,11 +2208,14 @@ impl ReaderPool {
WHEN pg.path LIKE 'gotchas/%' THEN 'gotcha' \
ELSE 'fact' \
END \
) \
), \
ws.name, pr.name \
FROM links l \
JOIN pages pg ON pg.id = l.from_page_id \
JOIN projects pr ON pr.id = pg.project_id \
JOIN workspaces ws ON ws.id = pg.workspace_id \
WHERE l.to_page_id = ?1 AND pg.is_latest = 1 \
ORDER BY pg.path";
ORDER BY ws.name, pr.name, pg.path";
let collect = |sql: &str| -> StoreResult<Vec<RelatedPage>> {
let mut stmt = conn.prepare(sql)?;
@@ -2161,6 +2224,8 @@ impl ReaderPool {
path: row.get(0)?,
title: row.get(1)?,
kind: row.get(2)?,
workspace: row.get(3)?,
project: row.get(4)?,
})
})?;
let mut out = Vec::new();
@@ -2178,6 +2243,111 @@ impl ReaderPool {
.await
}
/// List unresolved cross-project links authored by pages in this project
/// — declared dependencies on another project's page that don't resolve.
/// Each row says whether the named target project exists, so the lint can
/// tell a typo'd project name from a missing / renamed target page.
///
/// # Errors
/// Propagates any SQL or pool error.
pub async fn dangling_cross_project_links(
&self,
workspace_id: WorkspaceId,
project_id: ProjectId,
) -> StoreResult<Vec<DanglingCrossLink>> {
self.with_conn(move |conn| {
let mut stmt = conn.prepare(
"SELECT fp.path, l.to_workspace, l.to_project, l.to_path, \
EXISTS ( \
SELECT 1 FROM projects pr \
JOIN workspaces ws ON ws.id = pr.workspace_id \
WHERE pr.name = l.to_project \
AND ws.name = COALESCE( \
l.to_workspace, \
(SELECT name FROM workspaces WHERE id = ?1) \
) \
) AS project_exists \
FROM links l \
JOIN pages fp ON fp.id = l.from_page_id \
AND fp.workspace_id = ?1 AND fp.project_id = ?2 AND fp.is_latest = 1 \
WHERE l.to_page_id IS NULL AND l.to_project IS NOT NULL \
ORDER BY fp.path, l.to_project, l.to_path",
)?;
let rows = stmt.query_map(
params![workspace_id.as_bytes(), project_id.as_bytes()],
|row| {
let exists: i64 = row.get(4)?;
Ok(DanglingCrossLink {
from_path: row.get(0)?,
workspace: row.get(1)?,
project: row.get(2)?,
path: row.get(3)?,
project_exists: exists != 0,
})
},
)?;
let mut out = Vec::new();
for r in rows {
out.push(r?);
}
Ok(out)
})
.await
}
/// Resolved cross-project edges (links whose endpoints are in different
/// projects). When `scope` is `Some((ws, proj))`, only edges that touch
/// that project (as source or target) are returned; `None` returns the
/// whole cross-project graph. Powers `/api/v1/graph`.
///
/// # Errors
/// Propagates any SQL or pool error.
pub async fn cross_project_edges(
&self,
scope: Option<(WorkspaceId, ProjectId)>,
) -> StoreResult<Vec<CrossProjectEdge>> {
self.with_conn(move |conn| {
let base = "SELECT fw.name, fpr.name, fp.path, tw.name, tpr.name, tp.path \
FROM links l \
JOIN pages fp ON fp.id = l.from_page_id AND fp.is_latest = 1 \
JOIN pages tp ON tp.id = l.to_page_id AND tp.is_latest = 1 \
JOIN projects fpr ON fpr.id = fp.project_id \
JOIN workspaces fw ON fw.id = fp.workspace_id \
JOIN projects tpr ON tpr.id = tp.project_id \
JOIN workspaces tw ON tw.id = tp.workspace_id \
WHERE fp.project_id != tp.project_id";
let map_row = |row: &rusqlite::Row<'_>| {
Ok(CrossProjectEdge {
from_workspace: row.get(0)?,
from_project: row.get(1)?,
from_path: row.get(2)?,
to_workspace: row.get(3)?,
to_project: row.get(4)?,
to_path: row.get(5)?,
})
};
let mut out = Vec::new();
if let Some((_ws, proj)) = scope {
let sql =
format!("{base} AND (fp.project_id = ?1 OR tp.project_id = ?1) ORDER BY fw.name, fpr.name, fp.path");
let mut stmt = conn.prepare(&sql)?;
let rows = stmt.query_map(params![proj.as_bytes()], map_row)?;
for r in rows {
out.push(r?);
}
} else {
let sql = format!("{base} ORDER BY fw.name, fpr.name, fp.path");
let mut stmt = conn.prepare(&sql)?;
let rows = stmt.query_map([], map_row)?;
for r in rows {
out.push(r?);
}
}
Ok(out)
})
.await
}
/// Return one row per (workspace, project) with page-count and
/// last-updated aggregates. Used by the web UI project-list view.
///
@@ -2987,6 +3157,46 @@ fn count_project(
Ok(u64::try_from(n.unwrap_or(0)).unwrap_or(0))
}
/// Cross-project link degree for a project: `(dependents, dependencies)`.
/// `dependents` = distinct other projects whose pages link into this one;
/// `dependencies` = distinct other projects this one links out to. Counts
/// resolved links only (`to_page_id` is set); project ids are globally
/// unique, so a bare `!= project_id` excludes self across all workspaces.
fn cross_project_degree(
conn: &Connection,
workspace_id: WorkspaceId,
project_id: ProjectId,
) -> StoreResult<(u64, u64)> {
let dependents: Option<i64> = conn
.query_row(
"SELECT COUNT(DISTINCT fp.project_id) \
FROM links l \
JOIN pages tp ON tp.id = l.to_page_id \
AND tp.workspace_id = ?1 AND tp.project_id = ?2 AND tp.is_latest = 1 \
JOIN pages fp ON fp.id = l.from_page_id AND fp.is_latest = 1 \
WHERE fp.project_id != ?2",
params![workspace_id.as_bytes(), project_id.as_bytes()],
|row| row.get(0),
)
.optional()?;
let dependencies: Option<i64> = conn
.query_row(
"SELECT COUNT(DISTINCT tp.project_id) \
FROM links l \
JOIN pages fp ON fp.id = l.from_page_id \
AND fp.workspace_id = ?1 AND fp.project_id = ?2 AND fp.is_latest = 1 \
JOIN pages tp ON tp.id = l.to_page_id AND tp.is_latest = 1 \
WHERE tp.project_id != ?2",
params![workspace_id.as_bytes(), project_id.as_bytes()],
|row| row.get(0),
)
.optional()?;
Ok((
u64::try_from(dependents.unwrap_or(0)).unwrap_or(0),
u64::try_from(dependencies.unwrap_or(0)).unwrap_or(0),
))
}
fn count_workspace(conn: &Connection, sql: &str, workspace_id: WorkspaceId) -> StoreResult<u64> {
let n: Option<i64> = conn
.query_row(sql, params![workspace_id.as_bytes()], |row| row.get(0))
+18
View File
@@ -54,6 +54,7 @@ pub(crate) fn build(state: Arc<WebState>) -> Router {
"/workspaces/{workspace}/projects/{project}/overview",
axum::routing::get(project_overview_handler),
)
.route("/graph", axum::routing::get(graph_handler))
.with_state(state)
}
@@ -91,6 +92,23 @@ async fn workspaces_handler(State(state): State<Arc<WebState>>) -> Result<Respon
))
}
/// Cross-project dependency graph: every resolved link whose endpoints are
/// in different projects, each carrying both endpoints' workspace/project/
/// path. The UI builds nodes from the endpoints (and may aggregate to a
/// project-level dependency graph). Global for now; project scoping is a
/// follow-up query param.
async fn graph_handler(State(state): State<Arc<WebState>>) -> Result<Response, Response> {
let edges = state
.reader
.cross_project_edges(None)
.await
.map_err(internal_error)?;
Ok(with_cache(
Json(serde_json::json!({ "edges": edges })).into_response(),
LIST_CACHE_MAX_AGE,
))
}
async fn projects_handler(
State(state): State<Arc<WebState>>,
Query(query): Query<ProjectListQuery>,
+110 -15
View File
@@ -8,7 +8,7 @@
use std::collections::BTreeSet;
use ai_memory_core::PagePath;
use ai_memory_core::{LinkTarget, PagePath};
use serde::{Deserialize, Serialize};
use crate::error::WikiResult;
@@ -102,15 +102,21 @@ pub fn derive_title(frontmatter: &serde_json::Value, body: &str, path: &PagePath
stem.strip_suffix(".md").unwrap_or(stem).to_string()
}
/// A normalised link key: `(workspace, project, path)`. `workspace` and
/// `project` are `None` for a link that resolves within the source page's
/// own project. Collected in a `BTreeSet` so output is deduped + stable.
type LinkKey = (Option<String>, Option<String>, String);
/// Extract internal wiki links from a markdown body.
///
/// Supports `[[wiki links]]`, `[[wiki links|labels]]`, and ordinary
/// markdown links such as `[label](../decisions/foo.md#anchor)`. External
/// URLs, anchors, images, and non-markdown assets are ignored. Returned
/// paths are normalised to wiki-root-relative [`PagePath`] values.
/// Supports `[[wiki links]]`, `[[wiki links|labels]]`, cross-project
/// `[[project:path]]` / `[[workspace/project:path]]` wikilinks, and
/// ordinary markdown links such as `[label](../decisions/foo.md#anchor)`.
/// External URLs, anchors, images, and non-markdown assets are ignored.
/// Returned values are normalised to wiki-root-relative [`LinkTarget`]s.
#[must_use]
pub fn extract_links(body: &str, page_path: &PagePath) -> Vec<PagePath> {
let mut out = BTreeSet::new();
pub fn extract_links(body: &str, page_path: &PagePath) -> Vec<LinkTarget> {
let mut out: BTreeSet<LinkKey> = BTreeSet::new();
let mut in_fence = false;
for line in body.lines() {
@@ -127,11 +133,54 @@ pub fn extract_links(body: &str, page_path: &PagePath) -> Vec<PagePath> {
}
out.into_iter()
.filter_map(|path| PagePath::new(path).ok())
.filter_map(|(workspace, project, path)| {
PagePath::new(path).ok().map(|path| LinkTarget {
workspace,
project,
path,
})
})
.collect()
}
fn extract_wikilinks(line: &str, page_path: &PagePath, out: &mut BTreeSet<String>) {
/// Split an optional `[workspace/]project:` scope qualifier off the front
/// of a wikilink target. Returns `(workspace, project, path_part)`. URL and
/// scheme-prefixed targets carry no scope (the `:` belongs to the scheme);
/// [`normalize_link_target`] rejects those downstream.
fn split_scope(target: &str) -> LinkKey {
let lower = target.to_ascii_lowercase();
if target.contains("://")
|| lower.starts_with("mailto:")
|| lower.starts_with("data:")
|| lower.starts_with("javascript:")
|| lower.starts_with("tel:")
{
return (None, None, target.to_string());
}
if let Some((scope, rest)) = target.split_once(':') {
let scope = scope.trim();
let scope_ok = !scope.is_empty()
&& scope
.chars()
.all(|c| c.is_alphanumeric() || matches!(c, '-' | '_' | '/' | '.'));
if scope_ok {
let (workspace, project) = match scope.split_once('/') {
Some((ws, proj)) => (Some(ws.trim().to_string()), proj.trim()),
None => (None, scope),
};
if !project.is_empty() {
return (
workspace,
Some(project.to_string()),
rest.trim().to_string(),
);
}
}
}
(None, None, target.to_string())
}
fn extract_wikilinks(line: &str, page_path: &PagePath, out: &mut BTreeSet<LinkKey>) {
let mut rest = line;
while let Some(start) = rest.find("[[") {
let after_start = &rest[start + 2..];
@@ -139,14 +188,18 @@ fn extract_wikilinks(line: &str, page_path: &PagePath, out: &mut BTreeSet<String
break;
};
let raw = &after_start[..end];
if let Some(path) = normalize_link_target(raw, page_path, true) {
out.insert(path);
// Strip the `|label` first, then peel any cross-project scope so the
// remaining path normalises the same way a bare wikilink does.
let unlabelled = raw.split_once('|').map_or(raw, |(target, _)| target).trim();
let (workspace, project, path_part) = split_scope(unlabelled);
if let Some(path) = normalize_link_target(&path_part, page_path, true) {
out.insert((workspace, project, path));
}
rest = &after_start[end + 2..];
}
}
fn extract_markdown_links(line: &str, page_path: &PagePath, out: &mut BTreeSet<String>) {
fn extract_markdown_links(line: &str, page_path: &PagePath, out: &mut BTreeSet<LinkKey>) {
let mut start_at = 0;
while let Some(rel_start) = line[start_at..].find('[') {
let start = start_at + rel_start;
@@ -170,7 +223,7 @@ fn extract_markdown_links(line: &str, page_path: &PagePath, out: &mut BTreeSet<S
let target_end = target_start + rel_end;
let raw = &line[target_start..target_end];
if let Some(path) = normalize_link_target(raw, page_path, false) {
out.insert(path);
out.insert((None, None, path));
}
start_at = target_end + 1;
}
@@ -250,6 +303,48 @@ fn resolve_relative(page_path: &PagePath, target: &str, root_relative: bool) ->
mod tests {
use super::*;
fn page() -> PagePath {
PagePath::new("notes/here.md").unwrap()
}
#[test]
fn extract_links_bare_wikilink_is_local() {
let links = extract_links("see [[decisions/0001.md]] and [[other]]", &page());
assert!(links.iter().all(|l| !l.is_cross_project()));
assert!(
links
.iter()
.any(|l| l.path.as_str() == "decisions/0001.md" && l.project.is_none())
);
// bare name gets `.md` appended, still local
assert!(links.iter().any(|l| l.path.as_str() == "other.md"));
}
#[test]
fn extract_links_cross_project_wikilink() {
let links = extract_links("dep on [[infra:runbooks/02.md]]", &page());
let l = links.iter().find(|l| l.is_cross_project()).expect("xproj");
assert_eq!(l.workspace, None);
assert_eq!(l.project.as_deref(), Some("infra"));
assert_eq!(l.path.as_str(), "runbooks/02.md");
}
#[test]
fn extract_links_cross_workspace_wikilink_with_label() {
let links = extract_links("[[zommehq/zomme:decisions/adr-1.md|the ADR]]", &page());
let l = links.iter().find(|l| l.is_cross_project()).expect("xws");
assert_eq!(l.workspace.as_deref(), Some("zommehq"));
assert_eq!(l.project.as_deref(), Some("zomme"));
assert_eq!(l.path.as_str(), "decisions/adr-1.md");
}
#[test]
fn extract_links_url_wikilink_is_not_a_scope() {
// `https://...` must not be parsed as project "https".
let links = extract_links("[[https://example.com]] [[mailto:a@b.com]]", &page());
assert!(links.is_empty(), "URLs/schemes are not links: {links:?}");
}
#[test]
fn parses_frontmatter_and_body() {
let src = "---\ntitle: Hello\ntags:\n - a\n - b\n---\nThe body.\n";
@@ -350,7 +445,7 @@ mod tests {
[gotcha](../gotchas/hooks.md#details). Also \
[external](https://example.com) and ![image](../img/logo.png).";
let links = extract_links(body, &path);
let paths: Vec<&str> = links.iter().map(PagePath::as_str).collect();
let paths: Vec<&str> = links.iter().map(|l| l.path.as_str()).collect();
assert_eq!(
paths,
vec!["decisions/0001-single-sqlite-file.md", "gotchas/hooks.md"]
@@ -362,7 +457,7 @@ mod tests {
let path = PagePath::new("notes/a.md").unwrap();
let body = "```\n[[notes/ignored]]\n```\n[[notes/kept]]\n";
let links = extract_links(body, &path);
let paths: Vec<&str> = links.iter().map(PagePath::as_str).collect();
let paths: Vec<&str> = links.iter().map(|l| l.path.as_str()).collect();
assert_eq!(paths, vec!["notes/kept.md"]);
}
}
+28 -2
View File
@@ -160,7 +160,7 @@ backpressure, or single-writer SQLite actor.
| `pages_fts` | FTS5 virtual table over `(title, body)`, auto-synced by triggers. |
| `sessions`, `observations` | Hook capture, full audit log. |
| `observations_fts` | FTS5 virtual table over raw observation `(title, body)`, used only as bounded fallback. |
| `links` | Wikilink / markdown cross-references with `to_page_id` nullable for unresolved forward links. |
| `links` | Wikilink / markdown cross-references. `to_page_id` (a global PageId) is nullable for unresolved forward links. `to_workspace` / `to_project` carry a cross-project scope (NULL = the source page's own project). |
| `handoffs` | Typed cross-agent handoff records (open / accepted / expired). |
| `page_embeddings` | Optional vector rows for latest pages, with `(provider, model, dim)` denormalised so hybrid search can ignore stale vectors after an embedding config change and report missing-embedding diagnostics. |
| `audit_log` | Every mutation, addressable by `at DESC`. |
@@ -185,6 +185,32 @@ context, identity, rules, or user preferences; consolidation should not
rewrite an existing invariant slot unless new observations directly
contradict specific existing content.
## Cross-project links
Pages normally link within their own project (`[[decisions/0001.md]]`,
`[label](../gotchas/x.md)`). A wikilink can also name another project so
that dependencies between projects become explicit edges in the graph:
* `[[project:path.md]]` — a sibling project in the same workspace.
* `[[workspace/project:path.md]]` — a project in another workspace.
The parser (`ai-memory-wiki::extract_links`) yields a `LinkTarget
{ workspace, project, path }`; the store resolves it against the named
project's latest page and records the scope in `links.to_workspace` /
`links.to_project` (NULL = the source's own project, the common case).
Resolution is deferred-safe: a link to a page that does not exist yet
stays `to_page_id = NULL` and is repointed by
`refresh_incoming_links_for_path` when that page later lands — across
projects, not only within one.
Because `to_page_id` is a global id and `ReaderPool::page_links` joins by
id without a project filter, a resolved cross-project link surfaces as a
backlink on its target for free; `RelatedPage` carries the source's
`workspace` / `project` so the dependency is labelled and navigable. This
is what turns the per-project wikis into one dependency graph (see also
the `memory_lint` dangling-ref check, the briefing dependents counts, and
the `/api/v1/graph` endpoint).
## Crate layout
```
@@ -207,7 +233,7 @@ invariants below.
| Tool | Hint | Purpose |
|---|---|---|
| `memory_query` | read-only | FTS5 + graph RRF + optional vector RRF search, with raw fallback. Bumps access counters for page hits. |
| `memory_query` | read-only | FTS5 + graph RRF + optional vector RRF search, with raw fallback. Bumps access counters for page hits. Defaults to the current project; `scopes` searches named sibling projects; `global=true` searches every project at once (each hit annotated with its workspace + project). |
| `memory_recent` | read-only | Most-recently-updated `is_latest=1` pages. |
| `memory_read_page` | read-only | Fetch the FULL body of a single wiki page by `path` or by top FTS5 hit for a `query`. Use when an agent needs more than the 24-word snippets from `memory_query`. |
| `memory_status` | read-only | Counts, paths, version. |