Merge main into release/2.5 (forward-merge: 2.4.2 + #995 #985 #984 #978 #991)

# Conflicts:
#	CHANGELOG.md
#	crates/ai-memory-hooks/src/router.rs
#	docs/security-boundaries.md
This commit is contained in:
AkitaOnRails
2026-09-30 00:04:01 -03:00
10 changed files with 1140 additions and 96 deletions
+242 -8
View File
@@ -1,10 +1,3 @@
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
### Added
@@ -279,6 +272,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
(`docs/lifecycle-ops.md#backup`) is unchanged. (#950)
### Changed
- `ai-memory purge-session` without `--confirm` now previews what a confirmed
purge would delete before refusing, the same way `purge-project` does (#945):
@@ -351,6 +345,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
provider. (#981)
### Fixed
- A Kiro v3 resume that falls back to the default session store drops
`KIRO_HOME` from the child, but auto-wire still installed hooks and MCP
@@ -689,6 +684,244 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
its list structure. Some strict OKF v0.2 validators read §11.3 as
rejecting it. (#979)
- `install-hooks --apply --as-user <user> --auth-token <key>` no longer fails
with `--as-user '<user>' requires --auth-token` when the token was supplied.
The guard was handed the *rendered* credential, which is deliberately `None`
on the #552 secure path (the token is persisted under the data dir for the
hooks to read), so the recommended multi-user install command in
`docs/users.md` and `install-hooks --help` bailed on every native install
while reporting the token as absent. It now validates the resolved token.
`install-hooks` without `--as-user` was unaffected, and `--as-user` with no
token anywhere still bails. (#993)
- Wiki link extraction and the wikilink export now follow CommonMark for code
fences and link destinations: a fence closes only on the same glyph and at
least the opening length (so a ```` ``` ```` line inside `~~~` or a four-tick
fence no longer ends it), and a backtick line whose info string holds a
backtick is text, not a fence that hides the rest of the page from the link
index. `[doc](notes/foo_(1).md)` keeps its balanced parentheses, `<...>`
destinations and trailing titles are parsed (a `)` or link inside a title is
no longer read as part of the link), and a destination scan is bounded so a
line of unclosed `[a](` cannot stall a page write. (#985)
- `[routing] mid_session = "sticky"` was silently disabled for every
marker-covered install. `workspace` is a required `.ai-memory.toml` key, so
the host hook forwards `&workspace=…` on every event under any marker's
tree — including events that never left the session's own workspace — and
`overrides_permit_sticky` (`ai-memory-hooks::router`) disqualified
stickiness on that presence alone, before the project-provenance logic
(`project_src=repo-root` vs. `marker`) ever ran. A mid-session `cd` into a
sibling checkout therefore always rescoped under a marker, exactly as if
`sticky` were unset, while the identical scenario outside any marker's tree
(`workspace_override` naturally `None`) worked as documented — the entire
difference between the two cases this bug reports as "identical except for
being under `$HOME`". `find_session_scope` now runs before the sticky-permit
gate, and a `workspace_override` is resolved once against the session's own
workspace (`workspace_override_is_rescope`, via the existing no-create
`lookup_existing_workspace` — no new lookup mechanism): the SAME workspace
is not a rescope and falls through to the existing project-provenance logic
unchanged; a genuinely DIFFERENT or unresolvable workspace still fails
closed and disqualifies sticky, exactly as before. (#984)
- Auto-improve no longer rejects every proposal from a model that spells the
full-page edit mode as `"full"`. `edit_mode` had the same shape #458 fixed
for `operation`: a free-form string validated by exact match, no schema
constraint, and a system prompt that says "Full-page proposals" without the
literal value. `gpt-oss-20b` via LM Studio answered `"full"` for every
candidate, so runs finished with zero accepted proposals and only
`unsupported_edit_mode` rejections. The schema now advertises
`["full_page", "patch"]`, and normalisation folds `full`, `full-page` and
`Full Page` into `full_page` (with a warning) for providers without
constrained decoding. Unknown modes still fail validation. (#991)
## [2.4.2] - 2026-09-29
### Changed
- Documented Cheaper Inference as an endpoint for the existing `openai-compat`
provider. (#981)
- `docs/llm-providers.md` now has a dedicated OpenRouter subsection and a
matching row in the recommended-defaults table. The wiring
(`openai-compat` + `AI_MEMORY_LLM_BASE_URL=https://openrouter.ai/api/v1`)
and the `HTTP-Referer` / `X-Title` app-attribution headers were already
shipped, and `docker/.env.production.example` already ships an OpenRouter
default, but the provider-facing doc mentioned OpenRouter only inside the
generic `openai-compat` row. The new subsection covers a full working
env-file, the `openai-compat` embedder path (with a note to verify any
provider's `/v1/embeddings` endpoint before relying on it), a pointer to
`docs/llm-provider-comparison.md` for model selection, and a security /
gotchas block (auth-token requirement for non-loopback binds, env-file
hygiene, cheap-model consolidation drift, `:free`-tier shared-pool
limits, reasoning-model incompatibility). (#949)
- `docs/backup.md` documents the remote-git-mirror backup pattern for a
single-user install: what to include, what to exclude (derived SQLite index,
models cache, logs, secrets), how to schedule with a `systemd --user` timer,
how to restore, and the security posture per `SECURITY.md`, including a
"what ends up in your wiki" section that names the exposure (sanitized
prompts, tool I/O, page bodies) and lists encrypted-archive alternatives
(`age`, `restic`, `borg`, `git-crypt`) for cases where a private mirror
repo is not enough. A worked example ships under `docs/examples/backup/`
(snapshot script, `.service` and `.timer` unit files, `.gitignore` for
the mirror repo). Pointers added from `docs/deploy.md#backups`,
`docs/airgapped-install.md`, and the README docs table. The on-box
`ai-memory backup --to <tarball>` command
(`docs/lifecycle-ops.md#backup`) is unchanged. (#950)
### Fixed
- `observations.title` is now sanitized before it is truncated, not after.
`title_hint` used to be cut to 80 chars in `ai-memory-hooks::payload`
*before* the sanitizer ever ran, so a secret straddling that cutoff was
often left as a fragment too short to match a built-in or `[sanitize]
extra_patterns` rule — landing in the title, its FTS index, and every
surface that renders titles (session pages, briefings, handoffs, search)
unredacted, even though the same observation's body was correctly scrubbed
first. `title_hint` extraction now keeps the full first line untruncated;
`Sanitized::new` scrubs the title and only then applies the 80-char display
cap (`ai_memory_core::sanitize::truncate_for_title`), mirroring the order
the body already used. (#982)
- The `/web` page view keeps a leading H1 that is not the page title. It
dropped the body's first H1 whatever it said, as a duplicate of the title
in the header, but a frontmatter `title:` outranks the H1 and a setext H1
never names the page, so a heading like `# Token refresh after sleep`
under `title: Auth decisions` vanished from the rendered page. An H1 that
repeats the title is still dropped. (#967)
- `memory_query` now returns `global_scope_hits` (standing `_global` user/team
preferences) for a single-project query whose project is named explicitly
with `workspace`+`project`, not only when scope is omitted. The routing
doctrine tells static MCP clients to pass `workspace`+`project` on every call,
which set the old gate's "no named scope" condition to false, so those clients
never received global preferences despite the documented contract. The union
now keys on single-project resolution (`scopes` empty); only an explicit
multi-`scopes` set opts out, and `global=true`/`as_of` are unaffected. The
reserved-scope union is still keyed strictly to `_global` and never leaks
another project's pages. (#930)
- A page file rewritten under `wiki/<ws>/<project>/` (the OKF import path)
now gets its new version embedded the same way a brand-new file does.
The watcher's `reindex_page` upserted the new version and stopped —
embedding only ever ran on the `write_page` API path — so a rewrite left
hybrid search silently degraded to FTS-only ranking for that page until
someone ran `ai-memory embed` by hand. (#958)
- `[capture] ignore_paths` now covers shell commands. Shell tools (`Bash`,
`shell`, `execute_bash`, `terminal`, …) were classified as non-file and always
kept, so `cat docs/adr/*.md` stored the ignored file's full text in the
observation body. The native `ai-memory hook` now splits the command line
lexically and drops the event when an argument, resolved from the event's
`cwd`, matches an ignored pattern or is a glob that can reach one. Variables,
command substitution and commands that name no path are not followed. The
marker-file reference also documents excluding large tool results that Claude
Code saves and re-reads from `~/.claude/projects/**/tool-results/**`. (#946)
- The generated OpenCode, OMP, Pi and OpenClaw integrations now apply the same
lexical shell-command `ignore_paths` matching as the native hook, so a `bash`
call such as `cat docs/adr/0001.md` is dropped there too instead of being
captured. Both matchers now also recognize OpenClaw's and Devin's `exec` shell
tool, resolve relative arguments from a shell tool's `workdir` (OpenCode
`bash`, OpenClaw `exec`, Codex `shell`) instead of the event cwd, and treat
`dir/**` as covering `dir` itself when `dir` holds a glob (`docs/a?r/**`), as
the generated plugins already did. Refresh or reinstall generated plugins to
pick it up. (#948)
- Shell-command `ignore_paths` matching no longer joins an argument vector
before splitting it, which broke a path with spaces (`["cat", "private
notes/x.md"]`) apart and let one element's stray quote hide the elements
after it; each element now counts whole and is split on its own. An invalid
`.ai-memory.toml` now makes a shell command metadata-only, like a file tool,
instead of keeping its command and output, including one whose command is
missing or unparseable (`web_search` runs nothing and is still kept), and the
server now does the same when it cannot parse the client's capture marker. A
long `bash -lc "<script>"` element is read only as words, not also as one
path, so it can no longer exhaust the match budget and drop an innocuous
event. Applies to the native hook and the generated plugins; the server
accepts the new metadata-only shell form, so upgrade it together with them
(an older server drops such an event). (#973)
- `ai-memory-importer omc-wiki` now reads the frontmatter of a page saved
with CRLF line endings or a UTF-8 BOM, as a wiki checked out on Windows
with `core.autocrlf=true` is. It missed the fence, so the page's kind,
tier, tags and pin were dropped and the YAML block was imported as the
top of the body; the fence check now matches the wiki's own parser.
(#970)
- Grok Build CLI tool observations are no longer stored with an empty body.
Grok posts Claude Code's snake_case tool fields (`tool_name` / `tool_input` /
`tool_use_id`), but it was missing from both `closed_tool_agent` and the
tool-metadata agent match, so every `PostToolUse` body extraction returned
nothing while the observation itself was still captured. Grok now shares the
Claude Code tool mapping, so tool family, outcome and output land in the
body. (#931)
- `ai-memory serve --web-ui-dir` no longer panics at startup when the
custom SPA's `index.html` starts with a UTF-8 BOM, or has any other
non-ASCII text before `<head>`. The `<base href>` injection scanned the
page a byte at a time and sliced inside the multi-byte character
("byte index 1 is not a char boundary"); it now steps a whole
character. (#969)
- `ai-memory bootstrap` on a repository small enough for one chunk no longer
asks the provider for 64K output tokens. The output cap was keyed on the
number of chunks, so the only chunk of a small repo got the one-shot cap
meant for `--chunk-input-tokens 0`, and every such run failed on a
64K-context model. Under chunking (the default) every call now asks for up
to 16K; only `--chunk-input-tokens 0` keeps 64K. The `--max-input-tokens`
help no longer claims its 150K default leaves room for 64K of output in a
200K window. (#928)
- `ai-memory bootstrap` now leaves headroom for its own token estimate. It
counts bytes ÷ 4, which undercounts non-English text and source code (about
40% on Portuguese mixed with code, as measured for consolidation), and it
filled `--max-input-tokens` and `--chunk-input-tokens` to the last estimated
token, so a chunk sized to fit a model's window could overflow it on input
alone. Prunes and chunks now fill 80% of each budget by the estimate, the
same default consolidation uses; a run may plan more chunks than before.
(#937)
- The `bootstrap.md` manifest no longer shows bare `---` separators when a
chunk returns no rationale. Empty rationales are dropped before the
per-chunk ones are joined, and a run where no chunk returned one says so.
(#939)
- Multi-page consolidation (`memory_consolidate` with `multi_page=true`) no
longer overwrites a pinned page. The batch's page paths are chosen by the
model, and an update that named an existing pinned page replaced its body
and wrote the new version unpinned, despite pinned pages being documented as
immutable to automation. Such updates are now skipped with a warning; the
rest of the batch is written. `_slots/` pages, which are pinned
automatically, keep their state/invariant rules. (#934)
- The `/api/v1` single-page route's `ETag` now covers the whole JSON it
returns. It hashed only the markdown body and author, so pinning a page,
a frontmatter edit, or a new backlink changed the response without
changing the tag, and a client revalidating with `If-None-Match` got
`304` and kept the stale page. (#971)
- Session consolidation no longer writes a page title that already exists
in the project. A colliding session title gets a deterministic
`(session <8-char-id>)` suffix (stable for the same session, distinct
across sessions) and a matching leading H1 is retitled with it. The
consolidator prompt tells the model to name THIS session rather than a
generic harness-run phrase and not to reuse listed titles; that wording
is compact enough that the advertised 6000-token input floor still
projects observation bodies instead of dropping them. (#926)
- The web page view now links a `[[wikilink]]` on a line indented four
spaces that is not code. A nested list item written with four spaces
(`- Decisions:` then ` - see [[decisions/auth]]`) or a paragraph's
continuation line showed the wikilink as literal text, although the
engine indexed it and listed the page in the target's backlinks. The
preprocessor now skips exactly the code blocks and inline code the
renderer's parser reads as code. (#955)
- A wikilink or markdown link written inside an inline code span is no
longer indexed as a link. The engine skipped only fenced blocks, so a
page showing the syntax as code (`` `[[other-project:notes/x]]` ``) got a
lint `broken_link` finding for a dependency it does not have, and a
local example listed the page in the target's backlinks, while the web
page rendered neither as a link. A link whose label is code
(`` [`foo`](foo.md) ``) is still indexed. (#968)
- `export-okf`'s generated `index.md` no longer has a prose sentence outside
its list structure. Some strict OKF v0.2 validators read §11.3 as
rejecting it. (#979)
- `export-okf` now backfills a `title` (derived, same as `derive_title`) and
a `description` (from `summary`, else `abstract`) on an exported page when
missing, so a generic OKF consumer sees both §4.1-recommended keys. The
backfill only ever changes the bundle's copy, never the on-disk wiki file.
(#979)
- `export-okf` now rewrites a page's local `[[wikilink]]`s to bundle-relative
standard Markdown links, since a generic OKF consumer has no idea what
`[[decisions/b.md]]` means. A cross-project or cross-workspace wikilink has
no Markdown equivalent and ships untouched, as literal `[[...]]` text.
(#979)
- `sources[].author` in conformed frontmatter is now `process:<agent>`
(e.g. `process:claude-code`) instead of a bare agent name, matching the
OKF actor grammar's `process:<id>` form for automated processes (#979).
This changes the default `sources[].author` value written for every page
from now on; already-written pages are not retroactively rewritten.
## [2.4.1] - 2026-09-25
### Changed
@@ -7237,7 +7470,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- Consolidator used server startup default project instead of the
session's actual project.
[Unreleased]: https://github.com/akitaonrails/ai-memory/compare/v2.4.1...HEAD
[Unreleased]: https://github.com/akitaonrails/ai-memory/compare/v2.4.2...HEAD
[2.4.2]: https://github.com/akitaonrails/ai-memory/compare/v2.4.1...v2.4.2
[2.4.1]: https://github.com/akitaonrails/ai-memory/compare/v2.4.0...v2.4.1
[2.4.0]: https://github.com/akitaonrails/ai-memory/releases/tag/v2.4.0
[2.3.2]: https://github.com/akitaonrails/ai-memory/releases/tag/v2.3.2
Generated
+12 -12
View File
@@ -33,7 +33,7 @@ dependencies = [
[[package]]
name = "ai-memory-cli"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -84,7 +84,7 @@ dependencies = [
[[package]]
name = "ai-memory-consolidate"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"ai-memory-llm",
@@ -111,7 +111,7 @@ dependencies = [
[[package]]
name = "ai-memory-core"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"icu_normalizer",
"jiff",
@@ -127,7 +127,7 @@ dependencies = [
[[package]]
name = "ai-memory-eval"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -150,7 +150,7 @@ dependencies = [
[[package]]
name = "ai-memory-hooks"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -175,7 +175,7 @@ dependencies = [
[[package]]
name = "ai-memory-llm"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"anyhow",
@@ -204,7 +204,7 @@ dependencies = [
[[package]]
name = "ai-memory-mcp"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-consolidate",
"ai-memory-core",
@@ -239,7 +239,7 @@ dependencies = [
[[package]]
name = "ai-memory-store"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"anyhow",
@@ -265,11 +265,11 @@ dependencies = [
[[package]]
name = "ai-memory-test-support"
version = "2.4.1"
version = "2.4.2"
[[package]]
name = "ai-memory-web"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"ai-memory-store",
@@ -292,7 +292,7 @@ dependencies = [
[[package]]
name = "ai-memory-wiki"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"ai-memory-llm",
@@ -324,7 +324,7 @@ dependencies = [
[[package]]
name = "ai-memory-workstream"
version = "2.4.1"
version = "2.4.2"
dependencies = [
"ai-memory-core",
"anyhow",
+1 -1
View File
@@ -35,7 +35,7 @@ default-members = [
]
[workspace.package]
version = "2.4.1"
version = "2.4.2"
edition = "2024"
rust-version = "1.95"
license = "MIT"
@@ -437,7 +437,13 @@ pub fn run(config: &Config, mut args: InstallHooksArgs) -> Result<()> {
// to. Mismatch between `--as-user` and the actual token's owner is
// the operator's concern; we don't reach back to the server to
// verify (keeps install-hooks offline-capable).
validate_as_user(args.as_user.as_deref(), auth)?;
//
// Validate against the resolved token, not the rendered one. `auth` is
// deliberately `None` on the #552 secure path — the token was persisted
// under the data dir for the hooks to read — so passing it here made
// `--apply --as-user X --auth-token T` bail on the very combination
// docs/users.md recommends, with a message saying the token was absent.
validate_as_user(args.as_user.as_deref(), auth_token_owned.as_deref())?;
if let Some(user) = args.as_user.as_deref().filter(|s| !s.trim().is_empty()) {
eprintln!("[ai-memory] hooks installing for user: {user}");
}
@@ -8418,6 +8424,53 @@ model = "gpt-5"
);
}
/// #993 regression guard: `--apply --as-user X --auth-token T` is the
/// command docs/users.md tells operators to run, and it failed on every
/// native install. The guard was handed the rendered credential, which is
/// `None` precisely when the token was persisted (#552), so it reported the
/// token as missing while the token sat under the data dir. The unit tests
/// on `validate_as_user` itself could not see it — the wiring was wrong,
/// not the function.
#[test]
fn apply_with_as_user_succeeds_when_the_token_is_persisted() {
let home = TempDir::new().unwrap();
let cfg_dir = TempDir::new().unwrap();
let settings = cfg_dir.path().join("settings.json");
std::fs::write(&settings, "{}").unwrap();
let config = crate::config::Config::load(None, Some(home.path().to_path_buf())).unwrap();
let args = InstallHooksArgs {
agent: AgentChoice::ClaudeCode,
apply: true,
server_url: Some("http://127.0.0.1:49374".to_string()),
auth_token: Some("ALICE-KEY-993".to_string()),
as_user: Some("alice".to_string()),
config_file: Some(settings.clone()),
hooks_dir: Some(
std::path::PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../../hooks"),
),
..default_hook_args()
};
// The whole point of the fix: persisting the token must not turn a
// present `--auth-token` into an "absent" one for the `--as-user` guard.
run(&config, args).expect("install-hooks --apply --as-user with a token must succeed");
// And the attribution it promised is real: the token is on disk for the
// hooks, and still absent from the agent's own config.
assert_eq!(
crate::config::read_hook_auth_token(&config.data_dir).as_deref(),
Some("ALICE-KEY-993"),
"alice's key must be persisted where the hook reads it"
);
let rendered = std::fs::read_to_string(&settings).unwrap();
assert!(
!rendered.contains("ALICE-KEY-993"),
"the bearer must not reach the agent config: {rendered}"
);
}
/// F5 (docker-wrapper audit), the regression guard: when the bearer
/// *cannot* be persisted under the data dir, `--apply` must still install
/// working hooks by embedding the credential inline instead of aborting the
@@ -320,7 +320,12 @@ pub struct AutoImproveProposal {
#[serde(default, alias = "body", alias = "markdown", alias = "content")]
pub body_markdown: String,
/// `full_page` (default) or `patch`.
///
/// Advertised as an enum for the same reason as `operation`; still a
/// `String` so an unconstrained provider's `"full"` reaches
/// [`normalize_edit_mode`] instead of failing to deserialise.
#[serde(default = "default_edit_mode")]
#[schemars(extend("enum" = ["full_page", "patch"]))]
pub edit_mode: String,
/// Patch edits for existing _rules/ or procedures/ pages.
#[serde(default)]
@@ -383,6 +388,26 @@ fn default_edit_mode() -> String {
"full_page".into()
}
/// Map the ways a model spells the two supported edit modes onto their
/// canonical form.
///
/// Same narrow policy as [`normalize_operation`]: the system prompt says
/// "Full-page proposals", and models answer `"full"`. Unknown values are left
/// untouched so they still fail validation.
fn normalize_edit_mode(raw: &str) -> String {
let squashed: String = raw
.trim()
.to_ascii_lowercase()
.chars()
.filter(|c| !matches!(c, ' ' | '_' | '-'))
.collect();
match squashed.as_str() {
"" | "fullpage" | "full" => default_edit_mode(),
"patch" => "patch".into(),
_ => raw.to_string(),
}
}
/// A candidate the reviewer or validator rejected.
#[derive(Debug, Clone, Serialize, Deserialize, JsonSchema)]
pub struct AutoImproveRejectedCandidate {
@@ -1522,8 +1547,18 @@ pub(crate) fn validate_response(
fn normalize_proposal(proposal: &mut AutoImproveProposal, warnings: &mut Vec<String>) {
normalize_kind(proposal, warnings);
if proposal.edit_mode.trim().is_empty() {
proposal.edit_mode = default_edit_mode();
// Server-owned: set when a patch is materialized. The schema still shows
// it to the model, and a model-supplied non-hash value on a full-page
// proposal would reach staging and fail the whole run on `hex_to_sha256`.
proposal.expected_base_body_sha256 = None;
let original_edit_mode = proposal.edit_mode.clone();
proposal.edit_mode = normalize_edit_mode(&original_edit_mode);
if !original_edit_mode.trim().is_empty() && proposal.edit_mode != original_edit_mode {
warnings.push(format!(
"proposal {} edit_mode normalized from {:?} to {:?}",
proposal.path, original_edit_mode, proposal.edit_mode
));
}
if proposal.edit_mode == "patch" {
return;
@@ -3359,6 +3394,66 @@ mod tests {
);
}
/// Same shape as #458 on the neighbouring field: `gpt-oss-20b` via LM
/// Studio answered `"edit_mode": "full"` for every candidate, so every run
/// ended with zero accepted proposals and only `unsupported_edit_mode`
/// rejections.
#[test]
fn a_proposal_saying_full_is_accepted_as_full_page() {
let mut candidate = proposal("gotchas/thing.md", "gotcha", 0.91);
candidate.edit_mode = "full".into();
let raw = AutoImproveLlmResponse {
summary: "ok".into(),
proposals: vec![candidate],
rejected_candidates: Vec::new(),
};
let (accepted, rejected, warnings) =
validate_response(raw, &cfg(), &ExistingPageIndex::default());
assert!(rejected.is_empty(), "got {rejected:?}");
assert_eq!(accepted.len(), 1);
assert_eq!(accepted[0].edit_mode, "full_page");
assert!(
warnings
.iter()
.any(|w| w.contains("edit_mode normalized from \"full\"")),
"normalisation should be reported, got {warnings:?}"
);
let mut unknown = proposal("gotchas/thing.md", "gotcha", 0.91);
unknown.edit_mode = "rewrite".into();
let raw = AutoImproveLlmResponse {
summary: "ok".into(),
proposals: vec![unknown],
rejected_candidates: Vec::new(),
};
let (accepted, rejected, _) = validate_response(raw, &cfg(), &ExistingPageIndex::default());
assert!(accepted.is_empty());
assert_eq!(rejected[0].reason, "unsupported_edit_mode");
}
/// Once `"full"` stopped being rejected, the same `gpt-oss-20b` runs
/// failed at staging with `invalid expected_base_body_sha256: expected 64
/// hex chars`: the model filled in a field only the server can compute.
#[test]
fn a_model_supplied_base_sha_is_dropped_from_full_page_proposals() {
for supplied in ["", "null", "abc123", "N/A"] {
let mut candidate = proposal("gotchas/thing.md", "gotcha", 0.91);
candidate.expected_base_body_sha256 = Some(supplied.into());
let raw = AutoImproveLlmResponse {
summary: "ok".into(),
proposals: vec![candidate],
rejected_candidates: Vec::new(),
};
let (accepted, rejected, _) =
validate_response(raw, &cfg(), &ExistingPageIndex::default());
assert!(rejected.is_empty(), "{supplied:?}: got {rejected:?}");
assert_eq!(
accepted[0].expected_base_body_sha256, None,
"{supplied:?} must not survive to staging"
);
}
}
#[test]
fn patch_to_missing_or_non_context_target_rejects() {
let raw = AutoImproveLlmResponse {
@@ -3768,3 +3863,64 @@ mod operation_normalization_tests {
assert_eq!(parsed.operation, "create");
}
}
#[cfg(test)]
mod edit_mode_normalization_tests {
use super::*;
#[test]
fn the_spellings_models_actually_emit_are_accepted() {
for raw in [
"full_page",
"full",
"Full",
"full-page",
"Full Page",
"FULLPAGE",
"",
" ",
] {
assert_eq!(
normalize_edit_mode(raw),
"full_page",
"{raw:?} means full page and must normalise"
);
}
for raw in ["patch", "PATCH", " Patch "] {
assert_eq!(normalize_edit_mode(raw), "patch", "{raw:?} must normalise");
}
}
#[test]
fn unknown_edit_modes_are_left_to_fail() {
for raw in ["rewrite", "replace", "append", "diff", "delete"] {
assert_eq!(
normalize_edit_mode(raw),
raw,
"{raw:?} is not a supported mode and must keep failing validation"
);
}
}
#[test]
fn schema_constrains_edit_mode_to_the_supported_values() {
let schema = schemars::schema_for!(AutoImproveProposal);
let value = serde_json::to_value(&schema).expect("schema serialises");
let variants = value
.pointer("/properties/edit_mode/enum")
.and_then(|e| e.as_array())
.expect("edit_mode carries an enum constraint");
assert_eq!(
variants,
&vec![serde_json::json!("full_page"), serde_json::json!("patch")]
);
}
#[test]
fn a_non_canonical_edit_mode_still_deserialises() {
let parsed: AutoImproveProposal =
serde_json::from_value(serde_json::json!({ "edit_mode": "full" }))
.expect("must not fail to parse; normalisation happens before validation");
assert_eq!(parsed.edit_mode, "full");
}
}
+276 -45
View File
@@ -19,7 +19,10 @@ use ai_memory_core::{
NewSession, ObservationKind, ProjectId, Sanitized, Sanitizer, SessionId, WorkspaceId,
WorkstreamEvent, WorkstreamEventKind,
};
use ai_memory_store::{HookSessionAdmission, IngestObservationOutcome, StoreError, WriterHandle};
use ai_memory_store::{
HookSessionAdmission, IngestObservationOutcome, StoreError, WriterHandle,
lookup_existing_workspace,
};
use ai_memory_wiki::{AdmissionContext, AdmissionOp, Wiki};
use axum::Json;
use axum::Router;
@@ -2540,6 +2543,39 @@ fn sticky_out_of_tree_under_repo_root(
&& meaningful_session_anchor(session_cwd, home_dir).is_some()
}
/// Whether `workspace_override`, resolved against the session's own
/// workspace, is a genuine rescope (issue #976).
///
/// `workspace` is a *required* key in every `.ai-memory.toml` (see
/// `docs/marker-file.md`), so the host-side hook forwards `&workspace=…` on
/// every event under any marker's tree — including events that never left the
/// session's own workspace. Treating that presence alone as a rescope (the
/// pre-#976 behavior) silently disqualified sticky routing for every
/// marker-covered install, because the marker's own workspace rides along
/// with every event whether or not the agent actually moved.
///
/// `None` never disqualifies. `Some(name)` that resolves, via the existing
/// no-create scope lookup (never a hand-rolled chain — see the "Scope
/// resolution" rule in AGENTS.md), to the SAME workspace as the session is
/// not a rescope either: the event just re-declared where it already is.
/// Different, or unresolvable (a typo, a workspace that doesn't exist yet),
/// fails closed and is treated as a genuine rescope — matching
/// `ProjectSource::parse`'s fail-closed stance in this same file: a typo must
/// never silently downgrade a real rescope into a sticky no-op.
async fn workspace_override_is_rescope(
reader: &ai_memory_store::ReaderPool,
workspace_override: Option<&str>,
session_workspace: WorkspaceId,
) -> bool {
let Some(name) = workspace_override else {
return false;
};
match lookup_existing_workspace(reader, name).await {
Ok(resolved) => resolved != session_workspace,
Err(_) => true,
}
}
/// Whether this event's declared overrides leave room for session-sticky
/// attribution at all (issue #394's `sticky` knob).
///
@@ -2552,14 +2588,17 @@ fn sticky_out_of_tree_under_repo_root(
/// too old to tag its override reports `Unspecified` and keeps today's
/// behavior, so `sticky` degrades safely rather than silently capturing
/// deliberate rescopes.
///
/// `workspace_is_rescope` is [`workspace_override_is_rescope`]'s verdict,
/// resolved by the caller before this function runs (it needs the session's
/// own workspace and a DB read that this function has no access to).
fn overrides_permit_sticky(
workspace_override: Option<&str>,
workspace_is_rescope: bool,
project_override: Option<&str>,
project_source: ProjectSource,
routing: MidSessionRouting,
) -> bool {
// A workspace override only ever comes from a marker file.
if workspace_override.is_some() {
if workspace_is_rescope {
return false;
}
if project_override.is_none() {
@@ -2797,28 +2836,39 @@ async fn process_authorized(
// - Under `[routing] mid_session = "sticky"` the session also overrules a
// host-derived `repo-root` override, closing the cross-repo `cd` case;
// marker-declared scopes still win. See `overrides_permit_sticky`.
let sticky_scope = if moved_from_cwd.is_none()
&& overrides_permit_sticky(
env.workspace_override.as_deref(),
env.project_override.as_deref(),
env.project_source,
state.mid_session_routing,
) {
state
.reader
.find_session_scope(session_id)
.await?
.filter(|(_, _, session_cwd)| {
sticky_cwd_admits(
session_cwd.as_deref(),
env.cwd.as_deref(),
state.home_dir.as_deref(),
env.project_strategy,
state.mid_session_routing,
)
})
} else {
None
//
// `find_session_scope` runs unconditionally (issue #976) — the
// `workspace_override.is_some()` early return used to gate it, so a
// marker's mandatory `workspace` key disqualified stickiness before the
// session's own workspace was ever consulted. Resolving the session's
// scope first lets `workspace_override_is_rescope` compare the two:
// "still the session's own workspace" is not a rescope, only a genuinely
// different (or unresolvable) one is. A session with a pending native
// move (`moved_from_cwd`) never sticks — it is being rebound.
let session_scope = state.reader.find_session_scope(session_id).await?;
let sticky_scope = match session_scope {
Some((session_ws, session_proj, session_cwd)) if moved_from_cwd.is_none() => {
let workspace_is_rescope = workspace_override_is_rescope(
&state.reader,
env.workspace_override.as_deref(),
session_ws,
)
.await;
let permits = overrides_permit_sticky(
workspace_is_rescope,
env.project_override.as_deref(),
env.project_source,
state.mid_session_routing,
) && sticky_cwd_admits(
session_cwd.as_deref(),
env.cwd.as_deref(),
state.home_dir.as_deref(),
env.project_strategy,
state.mid_session_routing,
);
permits.then_some((session_ws, session_proj, session_cwd))
}
_ => None,
};
let publishable_scope = sticky_scope.is_some()
|| has_publishable_scope_hint(env.cwd.as_deref(), env.project_override.as_deref());
@@ -5210,42 +5260,104 @@ mod tests {
// The override gate for `[routing] mid_session` (#394). The invariant
// under test: a marker-declared project is a deliberate rescope and wins
// in BOTH modes; only a host-derived repo-root name may yield, and only
// under `sticky`.
// under `sticky`. The workspace side of the gate is exercised in terms of
// `workspace_is_rescope` directly here (the caller's already-resolved
// verdict) — `workspace_override_is_rescope`'s own DB-backed resolution
// (same-workspace vs. different vs. unresolvable) is covered separately
// below (issue #976) and in the `cross_repo_cd`-based integration tests.
#[test]
fn override_gate_distinguishes_marker_from_derived_project() {
use MidSessionRouting::{FollowCwd, Sticky};
use ProjectSource::{Marker, RepoRoot, Unspecified};
// (workspace, project, source, routing, expected)
// (workspace_is_rescope, project, source, routing, expected)
let cases = [
// No override at all: both modes may stick (pre-#394 behavior).
(None, None, Unspecified, FollowCwd, true),
(None, None, Unspecified, Sticky, true),
// No workspace rescope, no project override: both modes may
// stick (pre-#394 behavior).
(false, None, Unspecified, FollowCwd, true),
(false, None, Unspecified, Sticky, true),
// Marker-declared project: never yields, in either mode.
(None, Some("acme"), Marker, FollowCwd, false),
(None, Some("acme"), Marker, Sticky, false),
(false, Some("acme"), Marker, FollowCwd, false),
(false, Some("acme"), Marker, Sticky, false),
// Host-derived repo-root name: yields only under `sticky`. This
// is the cross-repo `cd` case the knob exists for.
(None, Some("acme"), RepoRoot, FollowCwd, false),
(None, Some("acme"), RepoRoot, Sticky, true),
(false, Some("acme"), RepoRoot, FollowCwd, false),
(false, Some("acme"), RepoRoot, Sticky, true),
// An untagged override from an older client stays authoritative,
// so `sticky` degrades safely instead of capturing a rescope.
(None, Some("acme"), Unspecified, Sticky, false),
// A marker workspace is itself a deliberate scope declaration.
(Some("oss"), None, Unspecified, Sticky, false),
(Some("oss"), Some("acme"), RepoRoot, Sticky, false),
(false, Some("acme"), Unspecified, Sticky, false),
// A genuine workspace rescope (a marker naming a DIFFERENT
// workspace than the session's own, or an unresolvable one)
// disqualifies sticky outright, regardless of project override.
(true, None, Unspecified, Sticky, false),
(true, Some("acme"), RepoRoot, Sticky, false),
];
for (ws, project, source, routing, expected) in cases {
for (workspace_is_rescope, project, source, routing, expected) in cases {
assert_eq!(
overrides_permit_sticky(ws, project, source, routing),
overrides_permit_sticky(workspace_is_rescope, project, source, routing),
expected,
"ws={ws:?} project={project:?} source={} routing={}",
"workspace_is_rescope={workspace_is_rescope:?} project={project:?} source={} routing={}",
source.as_str(),
routing.as_str(),
);
}
}
// Issue #976: `workspace_override_is_rescope`'s own resolution, in
// isolation from the rest of the sticky gate. `workspace` is a required
// marker key, so every marker-covered install forwards it on every
// event — the same workspace the session is already in, most of the
// time. Only a genuinely different (or unresolvable) workspace name may
// disqualify stickiness.
#[tokio::test]
async fn workspace_override_is_rescope_only_on_a_genuine_workspace_change() {
let tmp = TempDir::new().unwrap();
let state = make_state(&tmp).await;
let other_ws = state
.writer
.get_or_create_workspace("other-workspace")
.await
.unwrap();
// No override at all: never a rescope.
assert!(
!workspace_override_is_rescope(&state.reader, None, state.workspace_id).await,
"a missing workspace override must never disqualify sticky"
);
// The marker re-declaring the session's OWN workspace: not a
// rescope — this is the exact #976 repro (every event under a
// marker's tree carries the workspace it's already in).
assert!(
!workspace_override_is_rescope(&state.reader, Some("default"), state.workspace_id)
.await,
"re-declaring the session's own workspace must not disqualify sticky"
);
// A marker naming a genuinely different, EXISTING workspace: a real
// rescope.
assert!(
workspace_override_is_rescope(
&state.reader,
Some("other-workspace"),
state.workspace_id
)
.await,
"a marker naming a different workspace must disqualify sticky"
);
let _ = other_ws;
// An unresolvable workspace name (typo, not-yet-created): fails
// closed, exactly like `ProjectSource::parse`'s stance on an unknown
// `project_src` value.
assert!(
workspace_override_is_rescope(
&state.reader,
Some("no-such-workspace"),
state.workspace_id
)
.await,
"an unresolvable workspace override must fail closed and disqualify sticky"
);
}
// `sticky` extends out-of-tree inheritance to every strategy, but must
// NOT relax the broad-anchor guard that keeps a stray `$HOME` session
// from becoming a catch-all (#103).
@@ -5538,10 +5650,22 @@ mod tests {
/// host hook tagging each `project` override by provenance, and report
/// where the second observation landed. The shared fixture behind the
/// cross-repo cases below (#394).
///
/// `workspace`, when set, is forwarded on BOTH events — mirroring every
/// real marker-covered install, where `workspace` is a required marker
/// key and rides along with every event under the marker's tree whether
/// or not the agent actually changed workspace (#976). Before #976 this
/// fixture never set `workspace` at all, which is exactly why the bug —
/// a marker's mandatory `workspace` unconditionally disqualifying sticky
/// — went untested: none of the cross-repo cases below ever exercised a
/// marker's workspace key riding alongside a repo-root-derived project,
/// which is exactly the combination every real marker-covered install
/// sends.
async fn cross_repo_cd(
state: &HookState,
sid: &str,
source: &str,
workspace: Option<&str>,
) -> (ProjectId, Option<ProjectId>) {
let fire = |event: &str, cwd: &str, project: &str, project_src: Option<&str>| {
HookEnvelope::from_query_and_body(
@@ -5549,6 +5673,7 @@ mod tests {
event: event.into(),
agent: Some("claude-code".into()),
cwd: Some(cwd.to_string()),
workspace: workspace.map(str::to_owned),
project: Some(project.to_string()),
project_src: project_src.map(str::to_owned),
project_strategy: Some("repo-root".into()),
@@ -5598,6 +5723,15 @@ mod tests {
// hop into a sibling checkout keeps the session's project: the override
// is host-derived (`project_src=repo-root`), so it carries no operator
// intent and yields to the session.
//
// Also issue #976's regression case: BOTH events carry `workspace=default`
// (matching `state.workspace_id`'s name, per `make_state`), exactly as a
// real marker-covered install forwards its required `workspace` key on
// every event under the marker's tree — the same workspace the session
// is already in. Before the #976 fix, `workspace_override.is_some()`
// alone disqualified sticky here, so this exact test would have failed
// (see `mid_session_out_of_tree_cwd_still_resolves_per_event_under_basename`-
// style per-event splitting) had the fixture ever forwarded `workspace`.
#[tokio::test]
async fn sticky_routing_keeps_the_session_project_across_a_sibling_checkout() {
let tmp = TempDir::new().unwrap();
@@ -5612,7 +5746,7 @@ mod tests {
.unwrap();
assert_eq!(repo_a, None, "fixture starts clean");
let (landed, repo_b) = cross_repo_cd(&state, sid, "repo-root").await;
let (landed, repo_b) = cross_repo_cd(&state, sid, "repo-root", Some("default")).await;
let repo_a = state
.reader
.find_project(state.workspace_id, "repo-a".to_string())
@@ -5643,7 +5777,7 @@ mod tests {
);
let sid = "88888888-8888-4888-8888-888888888888";
let (landed, repo_b) = cross_repo_cd(&state, sid, "repo-root").await;
let (landed, repo_b) = cross_repo_cd(&state, sid, "repo-root", Some("default")).await;
let repo_b = repo_b.expect("follow-cwd mints the visited checkout's project");
assert_eq!(
landed, repo_b,
@@ -5661,7 +5795,7 @@ mod tests {
state.mid_session_routing = MidSessionRouting::Sticky;
let sid = "99999999-9999-4999-8999-999999999999";
let (landed, repo_b) = cross_repo_cd(&state, sid, "marker").await;
let (landed, repo_b) = cross_repo_cd(&state, sid, "marker", Some("default")).await;
let repo_b = repo_b.expect("a marker-declared project must still be created");
assert_eq!(
landed, repo_b,
@@ -5669,6 +5803,103 @@ mod tests {
);
}
// Issue #976's opposite-direction case, left untested before this fix: a
// marker naming a genuinely DIFFERENT workspace than the session's own
// must still rescope, even under `sticky`. Only "same workspace, still
// here" may fall through to the project-provenance logic above — a real
// workspace change is not drift and must never be captured by
// stickiness.
#[tokio::test]
async fn sticky_routing_still_rescopes_when_marker_names_a_different_workspace_mid_session() {
let tmp = TempDir::new().unwrap();
let mut state = make_state(&tmp).await;
state.mid_session_routing = MidSessionRouting::Sticky;
// Pre-create the target workspace so the mid-session event's lookup
// resolves to a real, distinct WorkspaceId (the `Ok(resolved) !=
// session_workspace` branch) instead of accidentally only exercising
// the fail-closed "unresolvable name" branch, which would also
// disqualify stickiness but for the wrong reason and leave the
// resolved-and-different comparison unproven at this level.
state
.writer
.get_or_create_workspace("other-workspace")
.await
.unwrap();
let sid = "cdcdcdcd-cdcd-4cdc-8cdc-cdcdcdcdcdcd";
let fire = |event: &str, cwd: &str, project: &str, workspace: &str| {
HookEnvelope::from_query_and_body(
HookQuery {
event: event.into(),
agent: Some("claude-code".into()),
cwd: Some(cwd.to_string()),
workspace: Some(workspace.to_string()),
project: Some(project.to_string()),
project_src: Some("repo-root".into()),
project_strategy: Some("repo-root".into()),
..Default::default()
},
serde_json::json!({
"session_id": sid,
"cwd": cwd,
"tool_name": "Bash",
}),
)
};
process(
&state,
fire("session-start", "/checkouts/repo-a", "repo-a", "default"),
None,
Vec::new(),
)
.await
.unwrap();
process(
&state,
fire(
"post-tool-use",
"/checkouts/repo-b",
"repo-b",
"other-workspace",
),
None,
Vec::new(),
)
.await
.unwrap();
let session_id: SessionId = sid.parse().unwrap();
let observations = state
.reader
.observations_for_session(session_id)
.await
.unwrap();
assert_eq!(observations.len(), 2);
let other_ws = state
.reader
.find_workspace("other-workspace".to_string())
.await
.unwrap()
.expect(
"a genuinely different workspace override must still resolve/create its own workspace",
);
let repo_b_in_other_ws = state
.reader
.find_project(other_ws, "repo-b".to_string())
.await
.unwrap()
.expect("the rescoped event's project must exist in the NEW workspace");
let second = observations.last().unwrap();
assert_eq!(
second.project_id, repo_b_in_other_ws,
"a marker naming a genuinely different workspace must still rescope, even under sticky"
);
assert_ne!(
second.workspace_id, state.workspace_id,
"the rescoped event must not stay in the session's original workspace"
);
}
// A client too old to tag its override reports no provenance. `sticky`
// must treat that as authoritative rather than assume it was derived,
// so a mixed-version fleet cannot silently capture deliberate rescopes.
+35 -2
View File
@@ -85,8 +85,12 @@ fn scope_relative_link<'a>(dest: CowStr<'a>, workspace: &str, project: &str) ->
if path_part.is_empty() || path_part.contains("..") {
return dest; // don't rewrite traversal; safe_url/router will reject
}
let path_part = path_part.trim_end_matches('/');
if path_part.is_empty() {
let mut path_part = path_part.trim_end_matches('/');
while let Some(rest) = path_part.strip_prefix("./") {
path_part = rest;
}
// `.//x` would leave a leading `/` and an empty first segment.
if path_part.is_empty() || path_part == "." || path_part.starts_with('/') {
return dest;
}
CowStr::Boxed(
@@ -801,4 +805,33 @@ mod tests {
);
}
}
#[test]
fn relative_link_with_leading_dot_slash_resolves() {
let html = render(
"[doc](./notes/foo.md) and [pointy](<./notes/bar.md>)",
"default",
"scratch",
);
assert!(
html.contains(r#"href="w/default/scratch/p/notes/foo.md""#),
"leading ./ must be stripped: {html}"
);
// The parser hands back a `<…>` destination without its brackets, so
// only the `./` needs handling here.
assert!(
html.contains(r#"href="w/default/scratch/p/notes/bar.md""#),
"pointy brackets and leading ./ must resolve: {html}"
);
}
#[test]
fn relative_link_with_leading_dot_slash_never_yields_an_empty_segment() {
let html = render("[a](.//x.md) and [b](././y.md)", "default", "scratch");
assert!(!html.contains("p//"), "empty path segment: {html}");
assert!(
html.contains(r#"href="w/default/scratch/p/y.md""#),
"repeated ./ must be stripped: {html}"
);
}
}
+359 -24
View File
@@ -127,6 +127,43 @@ pub fn derive_title(frontmatter: &serde_json::Value, body: &str, path: &PagePath
/// own project. Collected in a `BTreeSet` so output is deduped + stable.
type LinkKey = (Option<String>, Option<String>, String);
/// Active code fence delimiter and its opening run length (CommonMark §4.5).
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
struct CodeFence {
glyph: char,
len: usize,
}
impl CodeFence {
/// Update the fence state based on `line`.
///
/// Returns `(updated_fence_state, line_is_code_or_fence)`.
fn step(current: Option<Self>, line: &str) -> (Option<Self>, bool) {
let in_fence = current.is_some();
let trimmed = line.trim_start();
let glyph = match trimmed.chars().next() {
Some(c @ ('`' | '~')) => c,
_ => return (current, in_fence),
};
let count = trimmed.chars().take_while(|&c| c == glyph).count();
if count < 3 {
return (current, in_fence);
}
let info = &trimmed[count..];
match current {
// A backtick fence's info string cannot contain a backtick, so
// ```` ```code``` text ```` is a paragraph with an inline span,
// not a fence that would swallow the rest of the page.
None if glyph == '`' && info.contains('`') => (None, false),
None => (Some(CodeFence { glyph, len: count }), true),
Some(fence) if fence.glyph == glyph && count >= fence.len && info.trim().is_empty() => {
(None, true)
}
Some(fence) => (Some(fence), true),
}
}
}
/// Extract internal wiki links from a markdown body.
///
/// Supports `[[wiki links]]`, `[[wiki links|labels]]`, cross-project
@@ -139,15 +176,12 @@ type LinkKey = (Option<String>, Option<String>, String);
#[must_use]
pub fn extract_links(body: &str, page_path: &PagePath) -> Vec<LinkTarget> {
let mut out: BTreeSet<LinkKey> = BTreeSet::new();
let mut in_fence = false;
let mut fence: Option<CodeFence> = None;
for line in body.lines() {
let trimmed = line.trim_start();
if trimmed.starts_with("```") || trimmed.starts_with("~~~") {
in_fence = !in_fence;
continue;
}
if in_fence {
let (next_fence, is_fence_or_code) = CodeFence::step(fence, line);
fence = next_fence;
if is_fence_or_code {
continue;
}
let line = blank_inline_code(line);
@@ -397,17 +431,12 @@ fn split_scope(target: &str) -> LinkKey {
#[must_use]
pub fn rewrite_local_wikilinks(body: &str, page_path: &PagePath, own_project: &str) -> String {
let mut out = String::with_capacity(body.len() + 64);
let mut in_fence = false;
let mut fence: Option<CodeFence> = None;
for raw_line in body.split_inclusive('\n') {
let (line, terminator) = split_line_terminator(raw_line);
let trimmed = line.trim_start();
if trimmed.starts_with("```") || trimmed.starts_with("~~~") {
in_fence = !in_fence;
out.push_str(line);
out.push_str(terminator);
continue;
}
if in_fence {
let (next_fence, is_fence_or_code) = CodeFence::step(fence, line);
fence = next_fence;
if is_fence_or_code {
out.push_str(line);
out.push_str(terminator);
continue;
@@ -465,12 +494,14 @@ fn rewrite_wikilinks_in_segment(
return;
};
let raw = &after_start[..end];
match resolve_local_wikilink(raw, page_path, own_project) {
Some((href, label)) => {
let resolved = resolve_local_wikilink(raw, page_path, own_project)
.and_then(|(href, label)| Some((markdown_destination(&href)?, label)));
match resolved {
Some((destination, label)) => {
out.push('[');
out.push_str(&escape_markdown_link_label(&label));
out.push_str("](");
out.push_str(&href);
out.push_str(&destination);
out.push(')');
}
None => {
@@ -541,6 +572,38 @@ fn relative_href(page_dir: &[&str], target: &str) -> String {
parts.join("/")
}
/// `href` as a Markdown link destination, or `None` when it has none.
///
/// A bare destination cannot hold a space or control character, cannot
/// start with `<`, and needs balanced parentheses; anything else goes
/// inside `<…>` (CommonMark §4.7). A `<` or `>` in a destination that
/// needs the brackets can only be written with a backslash escape that
/// [`extract_links`] does not read back, so that target keeps its literal
/// `[[wikilink]]` rather than export a link that indexes elsewhere.
fn markdown_destination(href: &str) -> Option<String> {
let mut depth = 0usize;
let mut bare = !href.starts_with('<');
for c in href.chars() {
match c {
'(' => depth += 1,
')' => match depth.checked_sub(1) {
Some(open) => depth = open,
None => bare = false,
},
' ' => bare = false,
c if c.is_ascii_control() => bare = false,
_ => {}
}
}
if bare && depth == 0 {
return Some(href.to_string());
}
if href.contains(['<', '>']) {
return None;
}
Some(format!("<{href}>"))
}
/// Escape characters that would prematurely close a Markdown link label
/// (`[<label>](<href>)`): brackets, the backslash, and parentheses — an
/// unescaped `)` inside the label closes the link early.
@@ -590,18 +653,108 @@ fn extract_markdown_links(line: &str, page_path: &PagePath, out: &mut BTreeSet<L
continue;
}
let target_start = close + 2;
let Some(rel_end) = line[target_start..].find(')') else {
break;
let Some((raw, target_end)) = parse_link_destination(&line[target_start..]) else {
start_at = close + 1;
continue;
};
let target_end = target_start + rel_end;
let raw = &line[target_start..target_end];
if let Some(path) = normalize_link_target(raw, page_path, false) {
out.insert((None, None, path));
}
start_at = target_end + 1;
start_at = target_start + target_end + 1;
}
}
/// Longest destination (title included) worth scanning. A page path cannot
/// exceed the filesystem's limits, and an unbounded scan made a line of
/// `[a](` repeated quadratic on the page-write path.
const MAX_LINK_DESTINATION_BYTES: usize = 512;
/// How many whitespace runs in a bare destination are tried as the start of a
/// title before the rest is read as part of the destination.
const MAX_TITLE_PROBES: usize = 8;
/// Parse a CommonMark link destination immediately following the opening `(`
/// of `[label](<destination>)` or `[label](destination)` (CommonMark §4.7),
/// stepping over an optional title.
///
/// Returns `(destination_str, closing_paren_byte_offset)`.
///
/// Two deliberate departures from the spec: a bare destination may hold
/// whitespace when no title follows it (`[x](my page.md)`), which is what
/// earlier versions indexed and what the wikilink export used to emit; and
/// the scan gives up after [`MAX_LINK_DESTINATION_BYTES`].
fn parse_link_destination(rest: &str) -> Option<(&str, usize)> {
let rest = &rest[..rest.floor_char_boundary(MAX_LINK_DESTINATION_BYTES)];
let trimmed = rest.trim_start();
let leading = rest.len() - trimmed.len();
if let Some(after_lt) = trimmed.strip_prefix('<') {
let rel_gt = after_lt.find('>')?;
let after_gt = &after_lt[rel_gt + 1..];
let rel_paren = closing_paren_after_destination(after_gt)?;
Some((&after_lt[..rel_gt], leading + 1 + rel_gt + 1 + rel_paren))
} else {
let mut depth = 0usize;
let mut probes = 0;
let mut prev_ws = false;
for (i, c) in trimmed.char_indices() {
let ws = c.is_ascii_whitespace();
match c {
'(' => depth += 1,
')' if depth == 0 => return Some((trimmed[..i].trim_end(), leading + i)),
')' => depth -= 1,
_ if ws && !prev_ws && depth == 0 && probes < MAX_TITLE_PROBES => {
if let Some(close) = closing_paren_after_destination(&trimmed[i..]) {
return Some((&trimmed[..i], leading + i + close));
}
probes += 1;
}
_ => {}
}
prev_ws = ws;
}
None
}
}
/// Byte offset in `tail` of the `)` that closes a link whose destination
/// ended just before `tail`: optional whitespace, then optionally a title
/// (`"…"`, `'…'` or `(…)`, separated from the destination by whitespace),
/// then the `)`. `None` when `tail` is anything else, so a `)` or a link
/// inside a title is never taken for the end of the link.
fn closing_paren_after_destination(tail: &str) -> Option<usize> {
let body = tail.trim_start();
let lead = tail.len() - body.len();
let mut chars = body.char_indices();
let (_, first) = chars.next()?;
if first == ')' {
return Some(lead);
}
let closer = match first {
_ if lead == 0 => return None,
'"' => '"',
'\'' => '\'',
'(' => ')',
_ => return None,
};
let mut escaped = false;
for (i, c) in chars {
if escaped {
escaped = false;
} else if c == '\\' {
escaped = true;
} else if c == closer {
let after = &body[i + 1..];
let rest = after.trim_start();
return rest
.starts_with(')')
.then(|| lead + i + 1 + (after.len() - rest.len()));
} else if first == '(' && c == '(' {
return None; // an unescaped `(` cannot appear in a `(…)` title
}
}
None
}
fn normalize_link_target(raw: &str, page_path: &PagePath, wikilink: bool) -> Option<String> {
let target = raw
.split_once('|')
@@ -1109,4 +1262,186 @@ mod tests {
let _ = rewrite_local_wikilinks(body, &path, "proj");
}
}
#[test]
fn code_fence_respects_glyph_and_length() {
let md = "~~~\n[[a/b.md]]\n```\n[[c/d.md]]\n~~~\nafter [[e/f.md]]\n";
let path = PagePath::new("concepts/a.md").unwrap();
let links = extract_links(md, &path);
assert_eq!(links.len(), 1, "only link outside fence extracted");
assert_eq!(links[0].path.as_str(), "e/f.md");
let rewritten = rewrite_local_wikilinks(md, &path, "proj");
assert!(rewritten.contains("[[a/b.md]]"), "a/b remains literal");
assert!(rewritten.contains("[[c/d.md]]"), "c/d remains literal");
assert!(
rewritten.contains("[e/f.md](../e/f.md)"),
"post-fence wikilink rewritten: {rewritten}"
);
// 4 backticks cannot be closed by 3 backticks
let md4 = "````\n[[inside4.md]]\n```\n[[still_inside.md]]\n````\nafter [[outside.md]]\n";
let links4 = extract_links(md4, &path);
assert_eq!(links4.len(), 1);
assert_eq!(links4[0].path.as_str(), "outside.md");
}
#[test]
fn extract_links_parses_balanced_parentheses_and_pointy_destinations() {
let root = PagePath::new("here.md").unwrap();
let md = "See [doc](notes/foo_(1).md), [space](<notes/bar (2).md>), and [title](notes/baz.md \"a title\").\n";
let links = extract_links(md, &root);
assert_eq!(links.len(), 3, "{links:?}");
assert!(links.iter().any(|l| l.path.as_str() == "notes/foo_(1).md"));
assert!(links.iter().any(|l| l.path.as_str() == "notes/bar (2).md"));
assert!(links.iter().any(|l| l.path.as_str() == "notes/baz.md"));
}
#[test]
fn rewrite_local_wikilinks_encloses_destinations_with_spaces_in_pointy_brackets() {
let path = PagePath::new("concepts/a.md").unwrap();
let rewritten = rewrite_local_wikilinks("See [[decisions/my decision.md]].", &path, "proj");
assert_eq!(
rewritten,
"See [decisions/my decision.md](<../decisions/my decision.md>)."
);
// Without spaces, no pointy brackets needed
let plain = rewrite_local_wikilinks("See [[decisions/b.md]].", &path, "proj");
assert_eq!(plain, "See [decisions/b.md](../decisions/b.md).");
}
/// Link paths `md` yields when it sits at the wiki root, sorted.
fn linked(md: &str) -> Vec<String> {
let root = PagePath::new("here.md").unwrap();
extract_links(md, &root)
.into_iter()
.map(|l| l.path.as_str().to_string())
.collect()
}
#[test]
fn extract_links_keeps_whitespace_in_bare_destinations() {
// Earlier scans indexed these and the wikilink export used to emit
// the first form; cutting at the space aimed them at `decisions/my.md`.
assert_eq!(
linked("[d](decisions/my decision.md) [e](notes/a b/c d.md)"),
["decisions/my decision.md", "notes/a b/c d.md"]
);
assert_eq!(
linked("[f](notes/my file (1).md)"),
["notes/my file (1).md"]
);
assert_eq!(
linked("[g](notes/my page.md \"a title\")"),
["notes/my page.md"]
);
assert_eq!(linked("[h](b.md )"), ["b.md"]);
}
#[test]
fn extract_links_steps_over_titles_holding_parens_and_links() {
// A `)` or a link inside a title is not the end of the link, nor a link.
assert_eq!(linked("[a](<b.md> \"t (c) [d](e.md)\")"), ["b.md"]);
assert_eq!(linked("[a](c.md \"x :) [d](e.md)\")"), ["c.md"]);
assert_eq!(linked("[a](f.md \"x (y\")"), ["f.md"]);
assert_eq!(
linked("[a](g.md 'single') [b](h.md (paren title))"),
["g.md", "h.md"]
);
}
#[test]
fn extract_links_rejects_pointy_destination_followed_by_junk() {
assert!(linked("[a](<b.md>zzz)").is_empty());
// A title must be separated from the destination by whitespace.
assert!(linked("[a](<b.md>\"t\")").is_empty());
assert_eq!(linked("[a](<b.md> \"t\")"), ["b.md"]);
}
#[test]
fn extract_links_stays_linear_on_unclosed_destinations() {
// Each `](` used to rescan the rest of the line for a `)` that never
// comes: 256 KiB of `[a](` took seconds in release, minutes in debug,
// on the page-write path. The bound keeps it to a fraction of a second.
let root = PagePath::new("here.md").unwrap();
let started = std::time::Instant::now();
for unit in ["[a](", "[a](<x>"] {
let line = unit.repeat(256 * 1024 / unit.len());
assert!(extract_links(&line, &root).is_empty());
}
assert!(
started.elapsed() < std::time::Duration::from_secs(10),
"destination scan is no longer linear: {:?}",
started.elapsed()
);
}
#[test]
fn extract_links_bounds_destination_length_without_splitting_characters() {
// Straddle the scan window with 1- and 2-byte characters: cutting
// inside a character must not panic.
for n in 250..262 {
let _ = linked(&format!("[a]({}.md)", "é".repeat(n)));
let _ = linked(&format!("[a](a{}.md)", "é".repeat(n)));
}
assert!(linked(&format!("[a]({}.md)", "x".repeat(600))).is_empty());
assert_eq!(linked(&format!("[a]({}.md)", "x".repeat(200))).len(), 1);
}
#[test]
fn backtick_fence_info_string_cannot_contain_backticks() {
// Not a fence: a paragraph opening with an inline code span.
let md =
"```code``` intro [[a/b.md]]\n[[c/d.md]]\n```\n[[e/f.md]]\n```\nafter [[g/h.md]]\n";
assert_eq!(linked(md), ["a/b.md", "c/d.md", "g/h.md"]);
let path = PagePath::new("concepts/a.md").unwrap();
let rewritten = rewrite_local_wikilinks(md, &path, "proj");
assert!(rewritten.contains("[a/b.md](../a/b.md)"), "{rewritten}");
assert!(rewritten.contains("[[e/f.md]]"), "{rewritten}");
// A tilde fence may carry backticks in its info string.
assert!(linked("~~~ a`b\n[[x.md]]\n~~~\n").is_empty());
}
#[test]
fn rewrite_local_wikilinks_wraps_each_destination_a_bare_link_cannot_hold() {
let path = PagePath::new("concepts/a.md").unwrap();
for target in [
"notes/foo_(1).md",
"notes/a).md",
"notes/a(b.md",
"notes/a\tb.md",
"notes/a b).md",
"notes/(x) y.md",
] {
let rewritten = rewrite_local_wikilinks(&format!("See [[{target}]]."), &path, "proj");
assert_eq!(
extract_links(&rewritten, &path)
.iter()
.map(|l| l.path.as_str())
.collect::<Vec<_>>(),
[target],
"{rewritten:?} must index back to the page it was written from"
);
}
// Balanced parentheses need no brackets; unbalanced ones do.
let balanced = rewrite_local_wikilinks("[[notes/foo_(1).md]]", &path, "proj");
assert!(balanced.ends_with("](../notes/foo_(1).md)"), "{balanced}");
let unbalanced = rewrite_local_wikilinks("[[notes/a).md]]", &path, "proj");
assert!(unbalanced.ends_with("](<../notes/a).md>)"), "{unbalanced}");
}
#[test]
fn rewrite_local_wikilinks_keeps_targets_it_cannot_express_as_a_link() {
// `<` or `>` in a destination that needs the brackets would need a
// backslash escape the extractor does not read back.
let path = PagePath::new("concepts/a.md").unwrap();
let body = "See [[notes/a>b c.md]].";
assert_eq!(rewrite_local_wikilinks(body, &path, "proj"), body);
// Without whitespace or parens a bare destination holds them fine.
let bare = rewrite_local_wikilinks("[[notes/a>b.md]]", &path, "proj");
assert!(bare.ends_with("](../notes/a>b.md)"), "{bare}");
}
}
+2
View File
@@ -249,6 +249,8 @@ retain current behavior.
Recognized shell tools (`Bash`, `shell`, `exec`, `execute_bash`, `terminal`, …)
have no path field, so the command line is split into words lexically, the way
a POSIX shell quotes and separates them, without expanding or running anything.
Codex shell calls, including its `exec_command` path, reach hooks as `Bash`
with the command in `tool_input.command`.
A command given as an argument vector keeps each element as one word (a path
with spaces stays whole, up to 256 characters) and also splits each element on
its own, so a `bash -lc "<script>"` script is read like any command line. Each
+1 -1
View File
@@ -42,7 +42,7 @@ boundary not yet built.
| 8b | Messaging: cancel-own-only | `ai-memory-store/src/ops.rs` `cancel_messages` (AND-gated on `from_*` sender coordinate) | `agent_messages.rs` — whole-outbox **and** a foreign **specific-id** cancel refused | STRONG |
| 8c | Messaging: pop-exactly-once | `ai-memory-store/src/ops.rs` `pop_message_in_transaction` atomic CAS `WHERE state='pending'` | `agent_messages.rs` — sequential **and** `tokio::join!` concurrent double-pop yields exactly one `Some` | STRONG |
| 8d | Messaging: inferred-scope read is diagnosed (#854) | `scope.rs` `is_inferred`; `server.rs` `inferred_scope_hint` on empty pop/list | `agent_messages_briefing.rs` `no_scope_pop_that_misses_the_mail_is_diagnosed_not_a_silent_null` | STRONG |
| 9 | Scope resolution fail-closed | `ai-memory-store/src/scope.rs` no-create `lookup_existing_*`; create only via `create_explicit_scope` | `scope.rs` no-auto-create + `unscoped_write_with_unresolvable_coordinate_errors` | STRONG |
| 9 | Scope resolution fail-closed | `ai-memory-store/src/scope.rs` no-create `lookup_existing_*`; create only via `create_explicit_scope`. Also `ai-memory-hooks/src/router.rs` `workspace_override_is_rescope` — the hook router's session-sticky gate resolves a marker's `workspace` override via this same no-create `lookup_existing_workspace` (never a hand-rolled lookup) and fails closed (treats it as a rescope, disqualifying sticky) on a different or unresolvable workspace; only an override that resolves to the session's OWN workspace is not a rescope (#976) | `scope.rs` no-auto-create + `unscoped_write_with_unresolvable_coordinate_errors`; `router.rs` `workspace_override_is_rescope_only_on_a_genuine_workspace_change` (same/different/unresolvable workspace cases), `sticky_routing_keeps_the_session_project_across_a_sibling_checkout` (regression: same workspace forwarded on every event + `project_src=repo-root` from a nested checkout stays sticky), `sticky_routing_still_rescopes_when_marker_names_a_different_workspace_mid_session` (opposite direction: a genuinely different workspace still rescopes even under `sticky`) | STRONG |
| 10 | Destructive-op live-process refusal + confirm flags (invariant #9); `purge-project`'s `dry_run` always wins over `confirm` so a preview request can never become destructive | `ai-memory-cli/src/commands/process_guard.rs` `sibling_processes` + confirm flags in `reset`/`restore`/`reindex`/`uninstall --purge-data`/`purge_project`; `admin.rs` `handle_purge_project` checks `req.dry_run` unconditionally, before `req.confirm`, mirroring `reclaim-ledger-versions` | `admin_purge.rs` confirm→400; `removal.rs` — injected live-sibling makes each destructive command bail before touching the data dir; `admin_purge.rs` `purge_project_confirm_true_and_dry_run_true_still_only_previews` — `{"confirm": true, "dry_run": true}` deletes nothing (the regression this row now also covers); `ops.rs` `purge_project_dry_run_changes_nothing` — every project-scoped row count, the `purged_scopes` tombstone, and the `audit_log` row are identical before/after a `PurgeMode::Preview` run; `ops.rs` `purge_project_counts_collateral_damage_in_another_project` and `admin_purge.rs` `purge_project_dry_run_and_confirmed_purge_both_report_collateral_damage_in_another_project` — a purge of project P collaterally deletes an observation, and orphans a handoff's session reference, in a *different* project Q (via `sessions` cascading out of P), and both the preview and the confirmed report count it | STRONG |
| 10b | Session purge is scope+owner-bound (no cross-session/project over-delete); `purge-session`'s `dry_run` always wins over `confirm`, same as `purge-project` | `ai-memory-store/src/ops.rs` `purge_session` — selection scoped to `(workspace_id, project_id)` and keyed on this session's own `summary_page_id` **or** `path='sessions/<sid>.md'` + `json_extract(frontmatter_json,'$.session_id')=<sid>` (frontmatter owner, not the recursive latest-chain); `in_scope==0 → NotFound` fail-closed; whole op in one transaction; `admin.rs` `handle_purge_session` checks `req.dry_run` unconditionally, before `req.confirm` | `ops.rs` `purge_session_leaves_a_sibling_session_in_the_same_project_intact`, `…refuses_a_session_from_another_project_and_deletes_nothing`, `…refuses_a_session_from_another_workspace`, `…does_not_delete_an_identically_pathed_page_in_another_project`, `…removes_older_summary_versions_without_deleting_prior_manual_page` (#862); `admin_purge.rs` `purge_session_confirm_true_and_dry_run_true_still_only_previews` — `{"confirm": true, "dry_run": true}` deletes nothing; `ops.rs` `purge_session_dry_run_changes_nothing` — every table a session purge touches, the `purged_sessions` tombstone, and the `audit_log` row are identical before/after a `PurgeMode::Preview` run; `ops.rs` `purge_session_counts_collateral_damage_in_another_project` and `admin_purge.rs` `purge_session_dry_run_and_confirmed_purge_both_report_collateral_damage_in_another_project` — purging session S collaterally deletes an observation, and orphans a handoff's session reference, in a *different* project (via `S`'s own id cascading), and both the preview and the confirmed report count it; `admin_purge.rs` `purge_session_dry_run_refuses_a_session_outside_its_named_scope` — the same session previewed under a different real project (same workspace) or a different real workspace whose project shares the name both 404 with no counts in the body, with the session's own scope as the 200 control | STRONG |
| 10c | Ledger reclaim is content-gated, not filename-gated (invariant #16) | `ai-memory-store/src/ops.rs` `reclaim_ledger_versions` — a candidate row must pass `ai_memory_core::log_ledger::body_opens_with_log_ledger` on its own `body`, and only `is_latest=0 AND superseded_at IS NULL` rows are eligible, so decay-owned rows stay with `forget-sweep`; confirm gate in `admin.rs` `handle_reclaim_ledger_versions` | `reclaim_ledger_versions.rs` `a_prose_page_wearing_the_ledger_name_keeps_its_whole_version_chain`, `an_ordinary_page_chain_in_another_project_is_untouched`, `a_decay_tombstone_stays_with_the_sweep_that_owns_it`; `admin_reclaim_ledger_versions.rs` `deleting_without_confirm_is_refused_and_changes_nothing` (with dry-run and confirmed controls) | STRONG |