24 Commits
Author SHA1 Message Date
AkitaOnRails 97b8e4cbe1 Merge main into release/2.1: pick up the 2.0.4 fix batch
Forward-merge of the nine PRs that landed on main (the 2.0.4 batch: #642 auth
stale-bearer, #646/#640 LoginLimiter, #644 Cursor attribution, #650/#647
reindex manifest, #638 MCP routing, #652 CI docs, #645 dev-loop/build) into the
2.1 feature train, so release/2.1 carries every fix before 2.1.0 is cut.

Conflicts resolved:
- crates/ai-memory-wiki/src/wiki.rs: 2.1's per-page write lock (page_locks,
  #607) and main's manifested_scopes memo (#650) are independent additions to
  the same struct/imports/constructor — kept both; imports merged to
  {HashMap, HashSet}.
- CHANGELOG.md: [Unreleased] now carries 2.1's ### Added features above main's
  ### Changed + ### Fixed (the 2.0.4 fixes), Keep-a-Changelog order, single
  [2.0.3] section preserved.
- crates/ai-memory-llm/tests/extra_headers_on_the_wire.rs (2.1's #606 test)
  relocated into tests/suite/ and declared in mod.rs to satisfy #645's
  one-test-binary-per-crate harness convention (caught by the repo_layout guard).

fmt, clippy -D warnings, llm harness, and the repo_layout guard all green.
2026-09-05 12:55:11 -03:00
gb 3efb5b0b89 perf(dev): a self-contained build and a two-tier test loop on every platform
The edit-to-result loop was ~380s for the workspace on macOS and needed an
environment variable on every command. This makes `cargo t` the whole story
on macOS, Linux, and Windows, with numbers measured along the way.

Build
- `[profile.dev]` keeps only line tables (full debuginfo put ~190 MB of
  DWARF in each test binary and made the build linker-bound); dependencies
  build at opt-level 1 with no debuginfo; proc macros and build scripts at
  opt-level 3, since they are run once per dependent crate.
- Test binaries: 78 to 11 in the everyday loop (13 under `--workspace`).
  Each one is a link and, on macOS (Gatekeeper) and Windows (Defender), a
  first-run malware scan of the whole file, paid serially by nextest's list
  phase before the first test starts. Integration tests now live in
  `tests/suite/` and compile into their crate's own test harness (`mod.rs`,
  included from `src/lib.rs` under `#[cfg(test)]`, with `extern crate self`
  so they keep addressing the public API by crate name). Only the CLI keeps
  a separate `suite` target, because its tests run the built executable.
  The evals harness leaves `default-members`, so a bare `cargo t` skips its
  two binaries while `--workspace` (CI, the hook, `cargo tf`) still builds
  them. A repo-layout test fails on an undeclared suite file, a stray
  top-level `tests/*.rs`, or a `mod.rs` that `lib.rs` never includes.
- `ai-memory-cli` gains a lib target; `main.rs` is a shim. 806 tests that
  lived in the bin are reachable, and `--lib` runs skip the 127 MB binary.
- The web crate's vendored `static/tailwind.css` is the default on every
  build, so nothing needs `TAILWIND_SKIP=1` any more: every release, Docker,
  and CI path already used the vendored file, and the download branch only
  ever ran for developers who forgot the flag (and then rewrote the source
  tree as a side effect). `TAILWIND_BUILD=1 cargo build -p ai-memory-web`
  regenerates it explicitly. CI runs that on Linux and fails if the
  committed file is stale, a check that did not exist before; the committed
  file reproduces byte for byte today.
- `tokenizers` aligned on one version instead of the 0.21 pin plus the 0.22
  candle pulled in.

Test tiers
- `.config/nextest.toml`: the `default` profile skips any test whose module
  path has a segment starting with `slow` or `stress` (`packaging::slow::*`
  drives real wrapper scripts and fake container engines at 10-20s each;
  `stress_*` modules hammer concurrency), reports every failure in one run,
  and marks anything over 5s in its summary so a new slow test is visible
  the day it lands. `full` runs everything. `ci` keeps its retries and
  writes JUnit.
- `.cargo/config.toml` holds two aliases and nothing else: `cargo t` (default
  members) and `cargo tf` (`--workspace -P full`). `cargo t -p <crate>`
  builds just that crate. Neither passes `--all-targets`: there are no
  examples or benches, and it only added harnesses for two `test = false`
  targets.
- `scripts/install-git-hooks.sh` installs an opt-in pre-push hook that runs
  the full tier, touching only its own marked block. Two independent things
  run the skipped tier: that hook, and CI, which uses `cargo test` and never
  reads the nextest config.

Slow tests fixed rather than tiered
- `project_observations` in the consolidator trimmed an over-budget
  projection one observation at a time, re-rendering the whole text and
  re-scoring every remaining candidate after each removal. Each score scans
  the body, so 256 observations of 4k chars cost ~65k body scans per prompt:
  14s in production consolidation, exactly as in the unit test. Scores and
  per-block sizes are now computed once and the prune subtracts; output is
  unchanged and pinned by the existing tests. 13.9s to 0.18s.
- Windows takes ~2s to refuse a loopback connect, so every hook test that
  posted to a closed port paid 2s per request. `dead_http_endpoint()` in the
  new `ai-memory-test-support` crate accepts and closes instead, with a
  fallback to the closed port where binding is denied. devin hook tests:
  4.2s to 0.15s each.
- The store unit fixture opened a file-backed SQLite with the default
  rollback journal and synchronous=FULL, so ~120 parallel fixtures fsynced
  every transaction. journal_mode=MEMORY + synchronous=OFF: 242s to 89s of
  test time, p90 1.6s to 0.5s.
- Windows-only tests resolve `powershell.exe` or `pwsh.exe` once per process
  and the auto-improve eval fixtures are `.ps1` scripts instead of cmd.exe
  batch files; a post-bind settle sleep is gone; the two unpinned
  multi-thread tokio tests pin `worker_threads = 4`. The four copies of the
  PowerShell resolver and the mcp suite's duplicated `post`/`get` helpers are
  now one each.

Not done, with the numbers in AGENTS.md: nextest vs in-process libtest is a
wash per crate and a rout for the workspace (20s vs 309s); the
`local-embeddings` default feature costs ~50s of cold build and ~27 MB per
binary but under a second per relink, so it stays a product default.

Measured: workspace loop ~380s to ~150s on macOS; on a 32-thread Windows box
the warm everyday run is 20s of test time across 2919 tests in 11 binaries,
and the rebuild after a core edit is 13s of cargo with lld plus the
first-run scans.
2026-09-05 08:27:08 -04:00
AkitaOnRailsandClaude Opus 4.8 30ef4812f8 Merge PR #622: first-party OpenCode 2.0 beta support [2.1]
Reconciles enrell's #622 onto release/2.1. Additive new-harness feature
(opencode2): new CLI enum arms, wire aliases folded into the existing
AgentKind::OpenCode (no new variant, no migration), a read-only parameterized
SQLite transcript adapter for the beta session_v2/session_message tables, and
an ai-memory-opencode2.ts plugin derived from v1. Sequenced after #625: the
opencode2 plugin inherits #625's resolveToken() auth, so its test assertion is
updated from "Bearer ${TOKEN}" to "Bearer ${token}". CHANGELOG (Added),
mcp-install and support-matrix rows reconciled with the Codex-SessionEnd (#605)
and capture-assistant (#627) wording. Queued for 2.1.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-04 12:16:34 -03:00
enrell 33b5dd2c2d feat(opencode2): add first-party OpenCode 2.0 beta support
MCP (mcp.servers + oauth:false), dependency-free plugin shape
(ai-memory-opencode2.ts), managed run opencode2, v2 transcript
adapter (session_v2/session_message), acceptance case, docs,
changelog. Shares v1's config dir, session store, and agent kind;
no migration. Verified live against beta-18999.
2026-09-03 23:30:16 -03:00
gb 4f455b1f51 fix(tests): stop fixture git commits failing when the developer signs commits
Tests that need a real repository build a throwaway one in a temp dir and
commit into it with an inline fake identity (-c user.email=t@example.com).
Every other config key still resolves normally, so a global
commit.gpgsign=true makes git try to sign as that fake identity, find no
key for it, and abort with "gpg failed to sign the data".

CI cannot catch this: runners start with no global git config, so signing
is off there and all of these tests pass. It only reproduces on a
developer machine, where it looks like a broken test rather than a
machine-configuration problem.

Add --no-gpg-sign to the eight fixture commit sites. It overrides
commit.gpgsign for a single command and is inert where signing was
already off.

Also switch the router.rs git2 fixture from repo.signature() to a fixed
identity, so fixture commits are not authored by whoever ran the suite
(a determinism fix; libgit2 does not sign commits), and drop a commit
plus two git config calls in bootstrap.rs that nothing depended on -
MainRepoRoot resolves via Repository::discover() and never reads HEAD.
2026-09-03 17:06:30 -07:00
ff4f260267 feat(cli): add ai-memory rename-workstream for checkout-local renames (#538)
Workstream names were fixed at `run --new` time and had no correction path,
so a typo outlived the work it labelled and the only escape was starting a
new workstream and abandoning the ledger attached to the old one.

The rename selects by current name or by the stable id `workstreams` prints,
resolving the same (workspace, project, repository, worktree) identity `run`
selects with. Both selectors repeat the checkout predicate: `workstreams.id`
is globally unique, so without it a caller holding an id from another
checkout could retitle a workstream their request never named. An id outside
the resolved scope reads as absent rather than renamable.

It is metadata only. `workstream_events`, `managed_runs`, and
`workstream_native_sessions` all key on `workstreams.id`, so the rename is a
single-row update with nothing to cascade, and `selected_at` and
`updated_at` are deliberately left untouched — relabelling is not activity.
Neither the discovery listing order nor the workstream a bare `ai-memory
run` resumes moves as a side effect of fixing a typo.

The destination is validated exactly like a `--new` name and refused with a
named conflict when another workstream in the same checkout already holds
it, turning the UNIQUE constraint into a typed error instead of a bare
SQLite failure. Renaming a workstream to the name it already has writes
nothing and is not an error.

The new `/workstream/rename` route requires NormalWrite rather than the
NormalRead its sibling discovery read uses, and resolves scope through
`lookup_existing_scope` so a rename never creates the workspace or project
it names. The Docker shell wrapper routes the command through its native
host client alongside `run`, `show`, `continue`, and `workstreams`, since
repository identity is a host resource the helper container cannot see.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
2026-08-31 22:23:50 -03:00
7e02f09f3b feat(cli): add ai-memory workstreams checkout-local discovery (#499)
List the managed workstreams that `run --workstream` can select from the
current checkout, so branching between lines of work no longer requires
remembering names or reading the database by hand.

The list resolves the same (workspace, project, repository, worktree)
identity `run` selects with, puts the current selection first, then orders
by recent activity. Rows carry the linked harnesses and the stable
workstream id that `workstream-search --workstream-id` accepts.

Discovery is a read of workstream metadata only: checkout paths, repository
and worktree fingerprints, and native session ids never travel back to the
client. The new `/workstream/recent` route requires NormalRead, resolves
scope through `lookup_existing_scope` (no-create, fails closed), and bounds
the limit server-side.

The Docker shell wrapper routes the command through its native host client
alongside `run`, `show`, and `continue`, since repository identity is a host
resource the helper container cannot see.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
2026-08-28 15:07:38 -03:00
AkitaOnRails 813b51aed1 feat(workstream): support Command Code sessions (#373) 2026-08-07 15:29:27 -03:00
AkitaOnRails ad7243eebe fix(workstream): accept current Kimi session cwd (#382) 2026-08-07 14:51:06 -03:00
AkitaOnRails bfee1547d2 feat(workstream): support Kiro CLI v3 sessions (#356) 2026-08-07 14:32:47 -03:00
AkitaOnRails eb05a773c7 fix(workstream): harden explicit Kiro v2 support 2026-08-04 13:21:13 -03:00
Samir Hanna VerzaandClaude Fable 5 4429adc34c feat(workstream): add a v2-engine Kiro CLI managed-workstream adapter
Add `ai-memory run kiro` (alias `kiro-cli`) for the AWS Kiro CLI (#356),
deliberately scoped to the binary's default v2 agent engine.

Cross-engine resume is prevented by construction, not detection: Kiro v3
sessions occupy a separate id space the v2 engine cannot resume (the
resume hint printed after a v3 session silently starts a fresh v2
session unless --v3 is added), and the v3 persisted-session format is
not publicly documented — so any `--v3`, `--mode`, or non-v2
`--agent-engine` selection turns the whole invocation into an unmanaged
passthrough with argv byte-identical. Headless `--no-interactive` runs
persist to Kiro's v1 SQLite store rather than the v2 session files this
adapter reads, so they pass through too, as do the one-shot
list/delete/model flags and every root utility subcommand except `chat`
(list verified against kiro-cli 2.16.0 --help-all).

Launch planning: session ids are server-assigned, so a fresh launch
injects nothing and the session is linked by the session-start hook or
discovered post-exit; a returning session appends `--resume-id <id>`
after user arguments (accepted at the root and on `chat`, verified on
the 2.16.0 binary). Explicit `-r/--resume`, `--resume-id`,
`--resume-picker`/`--list` selections always win. Kiro's `-v` is
verbose, not version, and stays managed.

Discovery/import: the interactive store is flat —
`$KIRO_HOME/sessions/cli/<uuid>.json` metadata + `<uuid>.jsonl` event
stream — and the adapter matches checkouts on the metadata `cwd`,
requiring the metadata session_id to agree with the file stem. The
parser imports the versioned v1 envelope (Prompt / AssistantMessage
with toolUse parts / ToolResults), ignores unknown record kinds for
forward compatibility, and annotates unknown envelope versions and
non-text parts as explicit losses. The stream is not verified
append-only, so it shares Kimi's rewrite tolerance: a prefix-hash
cursor that resets and replays on in-place rewrites, with stable
line-hash record ids deduplicating server-side.

`--yolo` maps to the official v2 flag `--trust-all-tools`, treats
`-a`/`--trust-tools` as already-satisfying (an explicit narrower trust
set is never widened), and maps nothing on non-v2 engines — v3 replaced
the flag with permissions.yaml — with a stderr notice.

Kiro joins the automatic bare-`run` pool (client and server side) and
the deterministic phase of the acceptance script, including a
byte-identical `--v3` passthrough check; the real-harness phase skips
kiro with an explanation because scripted headless turns cannot write
the store the adapter reads. Session-file shapes derive from public
kiro-cli 1.29.x references and are documented as pending revalidation
on a live logged-in install; the CLI argument contract was verified
against kiro-cli 2.16.0 locally.

Closes #356.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LJvj7D7czyZzqk4pDXgLxY
2026-08-04 13:07:51 -03:00
AkitaOnRails fc0b30657a test(workstream): harden Antigravity managed support 2026-08-01 10:28:02 -03:00
AkitaOnRails 960dbff7a6 fix(test): make workstream acceptance deterministic 2026-07-25 15:50:38 -03:00
AkitaOnRails 504a895416 fix(run): recover orphaned native sessions 2026-07-25 14:10:41 -03:00
iagogfe 9c6a2ab384 feat(workstream): add Grok Build CLI managed harness
Grok gains a managed adapter following the managed-harness contribution
protocol: wrapper-generated --session-id for fresh sessions, --resume for
linked ones, explicit native selectors always win (including the bare
--resume picker and --fork-session), utility subcommands pass through, and
wrapper --yolo maps onto the native --yolo/--always-approve alias pair.

The bounded workstream context packet is delivered through Grok's native
--rules flag (system-prompt append) because Grok discards SessionStart
stdout and its UserPromptSubmit hook is passive; acceptance happens only
after the child spawns, matching the Crush contract, and a natively
supplied --rules/--append-system-prompt wins with redelivery on a later
run. Transcript import reads chat_history.jsonl read-only with a
prefix-validated cursor and content-hash event ids so rewind-driven
journal rewrites cannot duplicate history; system prompts, encrypted
reasoning, and the injected <user_info> block are excluded as loss
annotations. Discovery matches checkouts through summary.json's recorded
info.cwd, never the URL-encoded bucket name, and honors GROK_HOME. Grok
stays out of the bare-mode automatic pool.

Verified against Grok Build CLI v0.2.111.
2026-07-24 17:16:03 -03:00
AkitaOnRails af29f57a61 feat(workstream): support kimi-cli alias 2026-07-23 15:23:24 -03:00
AkitaOnRails 260af15c8d fix(workstream): harden Kimi managed continuity 2026-07-23 11:57:26 -03:00
Lucas Oliveira 07366605b7 fix(cli): match the installed user-prompt-submit event stem for kimi
The default PosixNative/WindowsNative installs invoke the hook with the
script stem (--event user-prompt-submit), while the legacy shell path
posts user-prompt; the new delivery branch only matched the latter, so
on production installs it never fired and the trailing "{}" was
injected into the turn as literal context. Gate on
HookEvent::parse(...)==UserPrompt, which canonicalizes both tokens (and
the snake/native spellings). Other literal event comparisons in the
hook (session-start, session-end, stop, pre-compact, post-tool-use)
already match the installed kimi script stems.

Also fix the acceptance fixture assertion for a fresh kimi launch: the
fake harness echoes one empty line for a zero-argument invocation, so a
byte-empty argv log was impossible and the check could never pass;
assert the log carries no real argument instead.
2026-07-22 21:03:44 -03:00
Lucas Oliveira 9ddb322dac test: add kimi fake-mode fixture to managed workstream acceptance
Add kimi to the default harness pool, a fake kimi mode that honors
--session <id> and writes the native store layout ($KIMI_CODE_HOME/
sessions/<bucket>/<id>/state.json + agents/main/wire.jsonl) with the
round sentinel as context.append_message records, deterministic
fresh-launch/resume/import/dedup checks, and the real-phase kimi arm
(isolated KIMI_CODE_HOME, config.toml hook merge, -p invocation).
2026-07-22 21:03:44 -03:00
AkitaOnRails b3f137b142 fix: release managed leases after launcher failures 2026-07-20 19:35:54 -03:00
AkitaOnRails 382291ed80 fix: keep managed utility runs sessionless 2026-07-20 18:30:36 -03:00
AkitaOnRails dc494a2086 feat: improve managed workstream handoffs 2026-07-20 18:01:39 -03:00
AkitaOnRails 5c4469df13 feat: add managed cross-harness workstreams 2026-07-20 15:44:05 -03:00