Forward-merge of the nine PRs that landed on main (the 2.0.4 batch: #642 auth
stale-bearer, #646/#640 LoginLimiter, #644 Cursor attribution, #650/#647
reindex manifest, #638 MCP routing, #652 CI docs, #645 dev-loop/build) into the
2.1 feature train, so release/2.1 carries every fix before 2.1.0 is cut.
Conflicts resolved:
- crates/ai-memory-wiki/src/wiki.rs: 2.1's per-page write lock (page_locks,
#607) and main's manifested_scopes memo (#650) are independent additions to
the same struct/imports/constructor — kept both; imports merged to
{HashMap, HashSet}.
- CHANGELOG.md: [Unreleased] now carries 2.1's ### Added features above main's
### Changed + ### Fixed (the 2.0.4 fixes), Keep-a-Changelog order, single
[2.0.3] section preserved.
- crates/ai-memory-llm/tests/extra_headers_on_the_wire.rs (2.1's #606 test)
relocated into tests/suite/ and declared in mod.rs to satisfy #645's
one-test-binary-per-crate harness convention (caught by the repo_layout guard).
fmt, clippy -D warnings, llm harness, and the repo_layout guard all green.
The edit-to-result loop was ~380s for the workspace on macOS and needed an
environment variable on every command. This makes `cargo t` the whole story
on macOS, Linux, and Windows, with numbers measured along the way.
Build
- `[profile.dev]` keeps only line tables (full debuginfo put ~190 MB of
DWARF in each test binary and made the build linker-bound); dependencies
build at opt-level 1 with no debuginfo; proc macros and build scripts at
opt-level 3, since they are run once per dependent crate.
- Test binaries: 78 to 11 in the everyday loop (13 under `--workspace`).
Each one is a link and, on macOS (Gatekeeper) and Windows (Defender), a
first-run malware scan of the whole file, paid serially by nextest's list
phase before the first test starts. Integration tests now live in
`tests/suite/` and compile into their crate's own test harness (`mod.rs`,
included from `src/lib.rs` under `#[cfg(test)]`, with `extern crate self`
so they keep addressing the public API by crate name). Only the CLI keeps
a separate `suite` target, because its tests run the built executable.
The evals harness leaves `default-members`, so a bare `cargo t` skips its
two binaries while `--workspace` (CI, the hook, `cargo tf`) still builds
them. A repo-layout test fails on an undeclared suite file, a stray
top-level `tests/*.rs`, or a `mod.rs` that `lib.rs` never includes.
- `ai-memory-cli` gains a lib target; `main.rs` is a shim. 806 tests that
lived in the bin are reachable, and `--lib` runs skip the 127 MB binary.
- The web crate's vendored `static/tailwind.css` is the default on every
build, so nothing needs `TAILWIND_SKIP=1` any more: every release, Docker,
and CI path already used the vendored file, and the download branch only
ever ran for developers who forgot the flag (and then rewrote the source
tree as a side effect). `TAILWIND_BUILD=1 cargo build -p ai-memory-web`
regenerates it explicitly. CI runs that on Linux and fails if the
committed file is stale, a check that did not exist before; the committed
file reproduces byte for byte today.
- `tokenizers` aligned on one version instead of the 0.21 pin plus the 0.22
candle pulled in.
Test tiers
- `.config/nextest.toml`: the `default` profile skips any test whose module
path has a segment starting with `slow` or `stress` (`packaging::slow::*`
drives real wrapper scripts and fake container engines at 10-20s each;
`stress_*` modules hammer concurrency), reports every failure in one run,
and marks anything over 5s in its summary so a new slow test is visible
the day it lands. `full` runs everything. `ci` keeps its retries and
writes JUnit.
- `.cargo/config.toml` holds two aliases and nothing else: `cargo t` (default
members) and `cargo tf` (`--workspace -P full`). `cargo t -p <crate>`
builds just that crate. Neither passes `--all-targets`: there are no
examples or benches, and it only added harnesses for two `test = false`
targets.
- `scripts/install-git-hooks.sh` installs an opt-in pre-push hook that runs
the full tier, touching only its own marked block. Two independent things
run the skipped tier: that hook, and CI, which uses `cargo test` and never
reads the nextest config.
Slow tests fixed rather than tiered
- `project_observations` in the consolidator trimmed an over-budget
projection one observation at a time, re-rendering the whole text and
re-scoring every remaining candidate after each removal. Each score scans
the body, so 256 observations of 4k chars cost ~65k body scans per prompt:
14s in production consolidation, exactly as in the unit test. Scores and
per-block sizes are now computed once and the prune subtracts; output is
unchanged and pinned by the existing tests. 13.9s to 0.18s.
- Windows takes ~2s to refuse a loopback connect, so every hook test that
posted to a closed port paid 2s per request. `dead_http_endpoint()` in the
new `ai-memory-test-support` crate accepts and closes instead, with a
fallback to the closed port where binding is denied. devin hook tests:
4.2s to 0.15s each.
- The store unit fixture opened a file-backed SQLite with the default
rollback journal and synchronous=FULL, so ~120 parallel fixtures fsynced
every transaction. journal_mode=MEMORY + synchronous=OFF: 242s to 89s of
test time, p90 1.6s to 0.5s.
- Windows-only tests resolve `powershell.exe` or `pwsh.exe` once per process
and the auto-improve eval fixtures are `.ps1` scripts instead of cmd.exe
batch files; a post-bind settle sleep is gone; the two unpinned
multi-thread tokio tests pin `worker_threads = 4`. The four copies of the
PowerShell resolver and the mcp suite's duplicated `post`/`get` helpers are
now one each.
Not done, with the numbers in AGENTS.md: nextest vs in-process libtest is a
wash per crate and a rout for the workspace (20s vs 309s); the
`local-embeddings` default feature costs ~50s of cold build and ~27 MB per
binary but under a second per relink, so it stays a product default.
Measured: workspace loop ~380s to ~150s on macOS; on a 32-thread Windows box
the warm everyday run is 20s of test time across 2919 tests in 11 binaries,
and the rebuild after a core edit is 13s of cargo with lld plus the
first-run scans.
Reconciles enrell's #622 onto release/2.1. Additive new-harness feature
(opencode2): new CLI enum arms, wire aliases folded into the existing
AgentKind::OpenCode (no new variant, no migration), a read-only parameterized
SQLite transcript adapter for the beta session_v2/session_message tables, and
an ai-memory-opencode2.ts plugin derived from v1. Sequenced after #625: the
opencode2 plugin inherits #625's resolveToken() auth, so its test assertion is
updated from "Bearer ${TOKEN}" to "Bearer ${token}". CHANGELOG (Added),
mcp-install and support-matrix rows reconciled with the Codex-SessionEnd (#605)
and capture-assistant (#627) wording. Queued for 2.1.0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
Tests that need a real repository build a throwaway one in a temp dir and
commit into it with an inline fake identity (-c user.email=t@example.com).
Every other config key still resolves normally, so a global
commit.gpgsign=true makes git try to sign as that fake identity, find no
key for it, and abort with "gpg failed to sign the data".
CI cannot catch this: runners start with no global git config, so signing
is off there and all of these tests pass. It only reproduces on a
developer machine, where it looks like a broken test rather than a
machine-configuration problem.
Add --no-gpg-sign to the eight fixture commit sites. It overrides
commit.gpgsign for a single command and is inert where signing was
already off.
Also switch the router.rs git2 fixture from repo.signature() to a fixed
identity, so fixture commits are not authored by whoever ran the suite
(a determinism fix; libgit2 does not sign commits), and drop a commit
plus two git config calls in bootstrap.rs that nothing depended on -
MainRepoRoot resolves via Repository::discover() and never reads HEAD.
Workstream names were fixed at `run --new` time and had no correction path,
so a typo outlived the work it labelled and the only escape was starting a
new workstream and abandoning the ledger attached to the old one.
The rename selects by current name or by the stable id `workstreams` prints,
resolving the same (workspace, project, repository, worktree) identity `run`
selects with. Both selectors repeat the checkout predicate: `workstreams.id`
is globally unique, so without it a caller holding an id from another
checkout could retitle a workstream their request never named. An id outside
the resolved scope reads as absent rather than renamable.
It is metadata only. `workstream_events`, `managed_runs`, and
`workstream_native_sessions` all key on `workstreams.id`, so the rename is a
single-row update with nothing to cascade, and `selected_at` and
`updated_at` are deliberately left untouched — relabelling is not activity.
Neither the discovery listing order nor the workstream a bare `ai-memory
run` resumes moves as a side effect of fixing a typo.
The destination is validated exactly like a `--new` name and refused with a
named conflict when another workstream in the same checkout already holds
it, turning the UNIQUE constraint into a typed error instead of a bare
SQLite failure. Renaming a workstream to the name it already has writes
nothing and is not an error.
The new `/workstream/rename` route requires NormalWrite rather than the
NormalRead its sibling discovery read uses, and resolves scope through
`lookup_existing_scope` so a rename never creates the workspace or project
it names. The Docker shell wrapper routes the command through its native
host client alongside `run`, `show`, `continue`, and `workstreams`, since
repository identity is a host resource the helper container cannot see.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
List the managed workstreams that `run --workstream` can select from the
current checkout, so branching between lines of work no longer requires
remembering names or reading the database by hand.
The list resolves the same (workspace, project, repository, worktree)
identity `run` selects with, puts the current selection first, then orders
by recent activity. Rows carry the linked harnesses and the stable
workstream id that `workstream-search --workstream-id` accepts.
Discovery is a read of workstream metadata only: checkout paths, repository
and worktree fingerprints, and native session ids never travel back to the
client. The new `/workstream/recent` route requires NormalRead, resolves
scope through `lookup_existing_scope` (no-create, fails closed), and bounds
the limit server-side.
The Docker shell wrapper routes the command through its native host client
alongside `run`, `show`, and `continue`, since repository identity is a host
resource the helper container cannot see.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Add `ai-memory run kiro` (alias `kiro-cli`) for the AWS Kiro CLI (#356),
deliberately scoped to the binary's default v2 agent engine.
Cross-engine resume is prevented by construction, not detection: Kiro v3
sessions occupy a separate id space the v2 engine cannot resume (the
resume hint printed after a v3 session silently starts a fresh v2
session unless --v3 is added), and the v3 persisted-session format is
not publicly documented — so any `--v3`, `--mode`, or non-v2
`--agent-engine` selection turns the whole invocation into an unmanaged
passthrough with argv byte-identical. Headless `--no-interactive` runs
persist to Kiro's v1 SQLite store rather than the v2 session files this
adapter reads, so they pass through too, as do the one-shot
list/delete/model flags and every root utility subcommand except `chat`
(list verified against kiro-cli 2.16.0 --help-all).
Launch planning: session ids are server-assigned, so a fresh launch
injects nothing and the session is linked by the session-start hook or
discovered post-exit; a returning session appends `--resume-id <id>`
after user arguments (accepted at the root and on `chat`, verified on
the 2.16.0 binary). Explicit `-r/--resume`, `--resume-id`,
`--resume-picker`/`--list` selections always win. Kiro's `-v` is
verbose, not version, and stays managed.
Discovery/import: the interactive store is flat —
`$KIRO_HOME/sessions/cli/<uuid>.json` metadata + `<uuid>.jsonl` event
stream — and the adapter matches checkouts on the metadata `cwd`,
requiring the metadata session_id to agree with the file stem. The
parser imports the versioned v1 envelope (Prompt / AssistantMessage
with toolUse parts / ToolResults), ignores unknown record kinds for
forward compatibility, and annotates unknown envelope versions and
non-text parts as explicit losses. The stream is not verified
append-only, so it shares Kimi's rewrite tolerance: a prefix-hash
cursor that resets and replays on in-place rewrites, with stable
line-hash record ids deduplicating server-side.
`--yolo` maps to the official v2 flag `--trust-all-tools`, treats
`-a`/`--trust-tools` as already-satisfying (an explicit narrower trust
set is never widened), and maps nothing on non-v2 engines — v3 replaced
the flag with permissions.yaml — with a stderr notice.
Kiro joins the automatic bare-`run` pool (client and server side) and
the deterministic phase of the acceptance script, including a
byte-identical `--v3` passthrough check; the real-harness phase skips
kiro with an explanation because scripted headless turns cannot write
the store the adapter reads. Session-file shapes derive from public
kiro-cli 1.29.x references and are documented as pending revalidation
on a live logged-in install; the CLI argument contract was verified
against kiro-cli 2.16.0 locally.
Closes#356.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJvj7D7czyZzqk4pDXgLxY
Grok gains a managed adapter following the managed-harness contribution
protocol: wrapper-generated --session-id for fresh sessions, --resume for
linked ones, explicit native selectors always win (including the bare
--resume picker and --fork-session), utility subcommands pass through, and
wrapper --yolo maps onto the native --yolo/--always-approve alias pair.
The bounded workstream context packet is delivered through Grok's native
--rules flag (system-prompt append) because Grok discards SessionStart
stdout and its UserPromptSubmit hook is passive; acceptance happens only
after the child spawns, matching the Crush contract, and a natively
supplied --rules/--append-system-prompt wins with redelivery on a later
run. Transcript import reads chat_history.jsonl read-only with a
prefix-validated cursor and content-hash event ids so rewind-driven
journal rewrites cannot duplicate history; system prompts, encrypted
reasoning, and the injected <user_info> block are excluded as loss
annotations. Discovery matches checkouts through summary.json's recorded
info.cwd, never the URL-encoded bucket name, and honors GROK_HOME. Grok
stays out of the bare-mode automatic pool.
Verified against Grok Build CLI v0.2.111.
The default PosixNative/WindowsNative installs invoke the hook with the
script stem (--event user-prompt-submit), while the legacy shell path
posts user-prompt; the new delivery branch only matched the latter, so
on production installs it never fired and the trailing "{}" was
injected into the turn as literal context. Gate on
HookEvent::parse(...)==UserPrompt, which canonicalizes both tokens (and
the snake/native spellings). Other literal event comparisons in the
hook (session-start, session-end, stop, pre-compact, post-tool-use)
already match the installed kimi script stems.
Also fix the acceptance fixture assertion for a fresh kimi launch: the
fake harness echoes one empty line for a zero-argument invocation, so a
byte-empty argv log was impossible and the check could never pass;
assert the log carries no real argument instead.
Add kimi to the default harness pool, a fake kimi mode that honors
--session <id> and writes the native store layout ($KIMI_CODE_HOME/
sessions/<bucket>/<id>/state.json + agents/main/wire.jsonl) with the
round sentinel as context.append_message records, deterministic
fresh-launch/resume/import/dedup checks, and the real-phase kimi arm
(isolated KIMI_CODE_HOME, config.toml hook merge, -p invocation).