mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
main
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8e42800a61 |
fix(sanitize): redact JSON secrets and Basic auth, strip control sequences
Three gaps at the privacy boundary, all reproduced against a running server by
posting hook payloads and reading the stored observations back out of SQLite.
1. A secret written in JSON was stored verbatim. The auth word rule ends in a
value class that excludes the quote, so `{"db_password":"..."}` never
matched while `db_password: ...` in YAML did. JSON is the shape most
captured tool payloads arrive in, which made this the common case rather
than the corner one. The rule now allows the quote on both sides of the
separator, accepts `=` as well as `:`, and allows a leading underscore in
the key so npm style `_authToken=` matches.
2. `Authorization: Basic <base64>` was stored verbatim although the rule names
`authorization`. The scheme word sits where the value class expected the
secret, and `Basic` is five characters, under the eight character floor, so
the match never started. `Bearer` looked covered only because a separate
pattern matches that keyword itself. The rule now allows an optional scheme
word before the value.
3. Terminal escape sequences, NUL and bidi overrides went into pages and back
out to the terminal. Stored text is replayed by `read-page` and `search`,
where an OSC sequence rewrites the window title and a bidi override
reverses what the reader sees, and a NUL makes the markdown file binary so
`grep` skips it and git stops diffing it. They are stripped rather than
redacted since they are not secrets, whole sequences at a time so a colour
code leaves no `[31m` litter behind, before the redaction passes so an
escape inside a secret cannot split it out of reach of every pattern. Tabs,
newlines and carriage returns are kept.
The value floor that keeps ordinary prose intact is unchanged, and a test
pins it: `Access-Control-Allow-Credentials: true`, `Idempotency-Key: abc` and
a sentence mentioning a password all survive byte identical.
|
||
|
|
3859cfe767 |
fix(wiki): refuse page paths that collide on a case-folding filesystem
Two page paths that differ only in case (`concepts/alpha.md` against `concepts/Alpha.md`) or in Unicode normalization (NFC `café.md` against its NFD spelling) are one file on macOS (APFS) and Windows (NTFS). The store kept a row for each while the second write overwrote the first page's file, so a read of either path returned the survivor's body, and the watcher then reindexed that file and superseded the overwritten row, removing the original content from disk and index alike. No error anywhere. Creates now fail closed: `upsert_page_in_tx` looks for a live page in the same scope whose portable key matches and returns `PagePathCollides` naming both paths. The file write is already rollback snapshotted, so the refused write leaves the existing page's bytes intact. The check runs on creates only. A supersede targets a path that already has its own row, so it cannot introduce a pair that did not exist before, and keeping the scan off that path avoids repeating it for every consolidation pass over the same page. The rule is platform independent for the same reason `ensure_portable` is: a wiki authored on Linux is rsynced and git synced to macOS and Windows, so what it accepts cannot depend on the filesystem underneath. `reindex` therefore skips such a pair (it can exist in a tree authored on a case sensitive filesystem) instead of failing the whole rebuild, counts it in the summary, and logs both paths. `portable_page_key` lowercases before composing: `İ` lowercases to `i` plus U+0307, which only composes back to a single scalar after the fold. `icu_normalizer` is already in the dependency tree via idna, so this pulls in no new code. |
||
|
|
dd7ad3fb96 |
fix(wiki): auto-commits stage what the wiki wrote, with the repository kept open
Since #665 a commit no longer re-hashes every page, but a session end still walked the whole tree, reopened the repository per commit, and dropped the commit when another session was writing a file mid-read. Every write into the tree now goes through the git adapter's own methods, which report the path; the crate's `clippy.toml` refuses the raw calls. A commit stages the reported paths and falls back to a full walk when nothing was reported, after a failed commit, when a reported path is outside the root, and once every ten minutes; a walk that stages an unreported write logs and counts it. The repository stays open between commits, the index file is written after a walk and every fifty commits, and a read that collides with an outside writer is retried instead of failing the commit. Measured on the LongMemEval harness (120 questions, Windows): 11 min 28 s against an estimated 35 min on main, with 0 dropped commits against 457. |
||
|
|
2504a2fb12 | docs(changelog): reference #665 on the wiki commit entry | ||
|
|
eabb9beba5 |
fix(wiki): stop re-hashing the whole tree on every commit
Since the #594 guard, every wiki commit cleared the git index and read and hashed every file in the working tree, so a session end cost the size of the wiki and grew with it. Measured with the LongMemEval harness on a Windows box, 30 questions ingested through the hook path with four in flight, upstream/main and this branch side by side: 12 min 54 s against 3 min 18 s, the upstream rate decaying from 5 questions a minute to 2 while this branch never drops below 3. Retrieval results identical. Staging now goes through libgit2's stat cache, so a commit costs what changed. The #594 case (an index entry naming a blob absent from the object store, which crash-looped the migration at boot) stays covered as the recovery path: when the tree write fails, the index is cleared and re-hashed once and the commit retried. It is reproduced directly by a test that deletes the blob from .git/objects, which the earlier test could not fabricate. Commits on one repository are serialized behind a lock. Two session ends at once used to collide on libgit2's index lock; the loser failed with "the index is locked" and its snapshot was dropped with a warning. A clean commit no longer rewrites the index file. Tests pin that after a hundred committed pages a one-file change stages one path, and that a same-size rewrite whose mtime is no newer than the index is still committed with the new content (libgit2's racy-entry handling, forced by giving the file the index's own mtime). |
||
|
|
a720a5316b |
fix(purge-session): remove the session's wiki page file after the DB purge (#653)
purge-session deleted the session's SQLite rows but left the live sessions/<id>.md on disk, and the watcher's reconcile pass reindexed it back on the next run. After the transaction commits, the admin endpoint now removes the returned page paths under the scoped project root, reports the result as files_deleted / files_failed alongside the logical removed_paths, and marks async purge_session observers with partial_failure when a file could not be removed. The CLI prints the on-disk result and warns about failures. Along the way: - the store returns removed_paths as typed PagePath, so the handler never re-parses strings - Wiki::remove_page_file is admission-free and skips the scope round-trip, matching remove_project_dir / remove_workspace_dir - one Wiki::dispatch_purge replaces the three identical per-kind copies - tests cover file removal, sibling-project isolation, and the partial-failure webhook payload, sharing a capture-hook helper with the delete-workspace test |
||
|
|
3efb5b0b89 |
perf(dev): a self-contained build and a two-tier test loop on every platform
The edit-to-result loop was ~380s for the workspace on macOS and needed an environment variable on every command. This makes `cargo t` the whole story on macOS, Linux, and Windows, with numbers measured along the way. Build - `[profile.dev]` keeps only line tables (full debuginfo put ~190 MB of DWARF in each test binary and made the build linker-bound); dependencies build at opt-level 1 with no debuginfo; proc macros and build scripts at opt-level 3, since they are run once per dependent crate. - Test binaries: 78 to 11 in the everyday loop (13 under `--workspace`). Each one is a link and, on macOS (Gatekeeper) and Windows (Defender), a first-run malware scan of the whole file, paid serially by nextest's list phase before the first test starts. Integration tests now live in `tests/suite/` and compile into their crate's own test harness (`mod.rs`, included from `src/lib.rs` under `#[cfg(test)]`, with `extern crate self` so they keep addressing the public API by crate name). Only the CLI keeps a separate `suite` target, because its tests run the built executable. The evals harness leaves `default-members`, so a bare `cargo t` skips its two binaries while `--workspace` (CI, the hook, `cargo tf`) still builds them. A repo-layout test fails on an undeclared suite file, a stray top-level `tests/*.rs`, or a `mod.rs` that `lib.rs` never includes. - `ai-memory-cli` gains a lib target; `main.rs` is a shim. 806 tests that lived in the bin are reachable, and `--lib` runs skip the 127 MB binary. - The web crate's vendored `static/tailwind.css` is the default on every build, so nothing needs `TAILWIND_SKIP=1` any more: every release, Docker, and CI path already used the vendored file, and the download branch only ever ran for developers who forgot the flag (and then rewrote the source tree as a side effect). `TAILWIND_BUILD=1 cargo build -p ai-memory-web` regenerates it explicitly. CI runs that on Linux and fails if the committed file is stale, a check that did not exist before; the committed file reproduces byte for byte today. - `tokenizers` aligned on one version instead of the 0.21 pin plus the 0.22 candle pulled in. Test tiers - `.config/nextest.toml`: the `default` profile skips any test whose module path has a segment starting with `slow` or `stress` (`packaging::slow::*` drives real wrapper scripts and fake container engines at 10-20s each; `stress_*` modules hammer concurrency), reports every failure in one run, and marks anything over 5s in its summary so a new slow test is visible the day it lands. `full` runs everything. `ci` keeps its retries and writes JUnit. - `.cargo/config.toml` holds two aliases and nothing else: `cargo t` (default members) and `cargo tf` (`--workspace -P full`). `cargo t -p <crate>` builds just that crate. Neither passes `--all-targets`: there are no examples or benches, and it only added harnesses for two `test = false` targets. - `scripts/install-git-hooks.sh` installs an opt-in pre-push hook that runs the full tier, touching only its own marked block. Two independent things run the skipped tier: that hook, and CI, which uses `cargo test` and never reads the nextest config. Slow tests fixed rather than tiered - `project_observations` in the consolidator trimmed an over-budget projection one observation at a time, re-rendering the whole text and re-scoring every remaining candidate after each removal. Each score scans the body, so 256 observations of 4k chars cost ~65k body scans per prompt: 14s in production consolidation, exactly as in the unit test. Scores and per-block sizes are now computed once and the prune subtracts; output is unchanged and pinned by the existing tests. 13.9s to 0.18s. - Windows takes ~2s to refuse a loopback connect, so every hook test that posted to a closed port paid 2s per request. `dead_http_endpoint()` in the new `ai-memory-test-support` crate accepts and closes instead, with a fallback to the closed port where binding is denied. devin hook tests: 4.2s to 0.15s each. - The store unit fixture opened a file-backed SQLite with the default rollback journal and synchronous=FULL, so ~120 parallel fixtures fsynced every transaction. journal_mode=MEMORY + synchronous=OFF: 242s to 89s of test time, p90 1.6s to 0.5s. - Windows-only tests resolve `powershell.exe` or `pwsh.exe` once per process and the auto-improve eval fixtures are `.ps1` scripts instead of cmd.exe batch files; a post-bind settle sleep is gone; the two unpinned multi-thread tokio tests pin `worker_threads = 4`. The four copies of the PowerShell resolver and the mcp suite's duplicated `post`/`get` helpers are now one each. Not done, with the numbers in AGENTS.md: nextest vs in-process libtest is a wash per crate and a rout for the workspace (20s vs 309s); the `local-embeddings` default feature costs ~50s of cold build and ~27 MB per binary but under a second per relink, so it stays a product default. Measured: workspace loop ~380s to ~150s on macOS; on a 32-thread Windows box the warm everyday run is 20s of test time across 2919 tests in 11 binaries, and the rebuild after a core edit is 13s of cargo with lld plus the first-run scans. |
||
|
|
4f455b1f51 |
fix(tests): stop fixture git commits failing when the developer signs commits
Tests that need a real repository build a throwaway one in a temp dir and commit into it with an inline fake identity (-c user.email=t@example.com). Every other config key still resolves normally, so a global commit.gpgsign=true makes git try to sign as that fake identity, find no key for it, and abort with "gpg failed to sign the data". CI cannot catch this: runners start with no global git config, so signing is off there and all of these tests pass. It only reproduces on a developer machine, where it looks like a broken test rather than a machine-configuration problem. Add --no-gpg-sign to the eight fixture commit sites. It overrides commit.gpgsign for a single command and is inert where signing was already off. Also switch the router.rs git2 fixture from repo.signature() to a fixed identity, so fixture commits are not authored by whoever ran the suite (a determinism fix; libgit2 does not sign commits), and drop a commit plus two git config calls in bootstrap.rs that nothing depended on - MainRepoRoot resolves via Repository::discover() and never reads HEAD. |
||
|
|
47c9cf2e71 |
fix(sanitize): redact opaque auth-bearing HTTP headers
An HTTP header carrying a credential was stored verbatim unless it used the `Bearer` keyword or an `UPPER_SNAKE_TOKEN=` shape. The bearer rule requires the literal keyword, and the generic env-var rule requires `[A-Z][A-Z0-9_]*_TOKEN`, which never matches a kebab-case header name. `X-Amz-Security-Token` (AWS SigV4), `X-Api-Key`, `Private-Token` (GitLab) and `Ocp-Apim-Subscription-Key` (Azure) all fall in that gap. Tool output echoing a `curl` invocation is a common way they reach capture, so this contradicted the pattern list's stated policy of preferring a stray false positive over a leaked credential. A `key` or `token` suffix alone does not imply a secret. `Idempotency-Key`, `Continuation-Token` and storage partition keys use it for values that carry no credential and stay useful when reading captured output, so the new rule requires an auth word to qualify that suffix. Unambiguous words such as `password`, `secret` and `authorization` stand alone. The header name is matched with a flat character class, so the rule has no nested quantifiers. The value floor keeps short literals such as CORS `Access-Control-Allow-Credentials: true` intact; it is not what protects an already-redacted value, since `[REDACTED]` starts with `[` and the value character class excludes it. Tests: 10 auth-header shapes redact, covering AWS, Azure, Google, GitLab and RapidAPI; 6 ordinary headers and 5 non-secret key/token headers survive byte-identical; double redaction keeps the header name; the flat name class is pinned against an adversarial hyphen run. |