Commit Graph
10 Commits
Author SHA1 Message Date
AkitaOnRailsandClaude Opus 4.8 da13f091f9 docs(https-via-proxy): recipe for TLS-inspecting firewalls / MITM roots (#954)
ai-memory's reqwest client is built with rustls-tls-native-roots, so it trusts
the OS store and honors SSL_CERT_FILE/SSL_CERT_DIR. Document how to give the
container a combined CA bundle (public roots + the interception root) so
outbound LLM/embedding calls stop failing with UnknownIssuer behind a
corporate MITM appliance or inspecting antivirus. Cross-referenced from
deploy.md's provider-failures troubleshooting bullet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-29 12:15:43 -03:00
AkitaOnRailsandClaude Opus 4.8 3e6db6d4c2 docs(proxy): note bootstrap's long-held POST needs raised proxy idle timeouts
Follow-up to #614: bootstrap holds one POST open for the whole multi-chunk
run (20+ min) with no bytes flowing, so a proxy's default read/idle timeout
resets it and the run is lost. Documents the Caddy (transport read/write
timeout 0) and nginx (proxy_read/send_timeout) fixes the reporter verified,
in a dedicated section of https-via-proxy.md. Ordinary MCP/api requests are
short and unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm
2026-09-03 16:23:12 -03:00
15178b885e feat(admin): add password sessions and separated API credentials (#533)
* feat(auth): add browser sessions and API credentials

* fix(auth): preserve 1.x compatibility paths

* fix(changelog): preserve released sections

* docs(auth): tie mirror triggers to 2.0 cutover

* docs(store): correct mirror migration version

* docs(changelog): link admin console PR

* fix(migrations): renumber human_auth/api_credentials to V54/V55

Main gained V52__purged_sessions_tombstones and V53__page_embed_failures
after this branch was opened, so the merge produced four migration files
across two version numbers. Git saw no conflict - the filenames differ -
and the collision only surfaces at runtime:

    UNIQUE constraint failed: refinery_schema_history.version

on a fresh database, so the merged tree could not open a store at all.

Renumbered V52__human_auth -> V54 and V53__api_credentials -> V55, and
shifted the version numbers the tests pin: `run_to`/`open_to` targets, the
`schema_version` assertions, the rollback assertion, and the test names.
The pre-migration fixture point moves 51 -> 53 because main's V52/V53
create unrelated tables (purged_sessions, page_embed_failures) and touch
neither `users` nor any table these migrations rewrite.

Migration content is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

---------

Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 23:26:41 -03:00
7aa268f4c6 fix(mcp-bridge): give the streamable-HTTP client a TLS backend (#497)
* fix(mcp-bridge): give the streamable-HTTP client a TLS backend

ai-memory-cli enabled rmcp's transport-streamable-http-client-reqwest feature,
which pulls reqwest in through `__reqwest = ["dep:reqwest"]` and stops there.
That leaves reqwest 0.13 with no TLS feature at all:

    reqwest v0.13.4
    |-- reqwest feature "json"
    `-- reqwest feature "stream"

reqwest built without a TLS backend refuses any non-http scheme, so every
`ai-memory mcp-bridge` run against an https:// server URL failed at connect:

    error sending request for url (https://...)
      caused by: client error (Connect)
      caused by: invalid URL, scheme is not http

The bridge is the Claude Code stdio bridge to a *remote* MCP server and it
attaches a bearer token via `config.auth_header`, so the practical position was
that the only scheme which worked was the one that sends that token in
cleartext.

This is a second, independent copy of reqwest in the tree — the workspace
client is 0.12 and its root-store configuration never applied here.

rmcp exposes `reqwest = ["__reqwest", "reqwest?/rustls"]`, and on reqwest 0.13
the `rustls` feature pulls rustls-platform-verifier rather than a bundled root
set, so the bridge now verifies against the platform trust store — the same
behaviour the workspace client uses for every other outbound request.

* fix(mcp-bridge): use rustls-no-provider with ring instead of pinning aws-lc

rmcp's plain `reqwest` feature is `reqwest?/rustls`, which on reqwest 0.13
pins aws-lc-rs and adds a compiled C dependency to every build of an otherwise
pure-Rust workspace, for a bridge most users never enable.

rmcp 1.7 exposes `reqwest-tls-no-provider` (`reqwest?/rustls-no-provider`):
the platform certificate verifier without a pinned crypto provider. ring is
already compiled for this binary via the reqwest 0.12 the rest of the workspace
uses, so the bridge installs that as the process default before opening its
transport, and fails with a named error rather than deep inside a handshake if
no provider ends up installed.

Also documents in docs/https-via-proxy.md that the bridge now reaches https://
endpoints, since that is the page an operator hitting this reads first.

* chore(mcp-bridge): correct what the no-provider TLS route actually avoids

The comment said the plain `reqwest?/rustls` route "pins aws-lc-rs and
drags in a C toolchain and an Android JNI stack", which reads as though
this route avoids all three. It avoids the first two.

Both routes go through `rustls-platform-verifier`, whose dependencies
include `jni` and `rustls-platform-verifier-android`, so both put them in
the lockfile - verified against the merged lock, where they enter with
this change. They are target-gated and not compiled for our platforms,
but they are not a difference between the two options.

The crypto provider is the whole of the difference, and avoiding aws-lc-rs
is a good enough reason on its own. Left as written, the comment would
misdirect whoever next weighs a TLS route here into looking for an option
that does not exist.

Maintainer edit on top of @abhisheksharma2411's work; the fix itself is
theirs and unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

* chore(changelog): move the mcp-bridge TLS entry into [Unreleased]

This branch was opened before v1.38.0 was cut. Its entry was written under
`## [Unreleased]`, and merging main moved that heading down without moving
the entry with it - git reanchored the addition by context and landed it
inside the frozen `[1.33.1]` section, with no conflict to notice.

`scripts/check-changelog-frozen.sh` exists for exactly this and caught it.
Content unchanged; only the section.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MDbhmszrjG9s5MrPrTuNtm

---------

Co-authored-by: AkitaOnRails <fabioakita@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-31 21:54:24 -03:00
AkitaOnRails 1a5806e75c security: owner-guarded hook sessions, private file modes, fail-closed binds
Audit-driven hardening bundle:

- Hook session UUIDs reject cross-owner reuse atomically inside the writer
  transaction before ingest-key claim, observation, summary, handoff, or
  end-state mutation; guarded end operations revalidate the persisted
  tuple+owner; explicit root `finalize-session --all-owners` recovery is
  SessionEnd-only and Admin-gated; keyed replay serialization holds the
  ingest gate through downstream completion.
- Unauthenticated non-loopback HTTP binds now fail closed before serving;
  `--allow-insecure-no-auth` is the explicit dangerous override. Loopback
  default unchanged.
- New data dirs (0700) and config/SQLite/segment/backup files (0600) are
  created owner-only regardless of umask; existing installs untouched.
- `[auth].secure_cookie` marks the /web browser cookie Secure for HTTPS
  reverse-proxy deployments; SameSite is now Strict; proxy headers are
  never trusted to infer HTTPS.
- /api/v1 internal errors return a stable generic 500 body while the
  detailed cause is logged server-side; AuthSettings Debug redacts bearer,
  pepper, and proxy secrets.
- Multi-user per-actor autoscope contract aligned with globally unique
  durable SessionIds: distinct run IDs per user; cross-owner same-ID reuse
  is dropped before pointer publication.

Verified: 2378 workspace tests, clippy -D warnings, gitleaks, cargo
audit/deny, plus live-server auth/permission checks.
2026-08-15 18:26:14 -03:00
AkitaOnRails 9d7ec11371 fix: harden memory and release trust boundaries 2026-07-30 00:42:38 -03:00
AkitaOnRails ab5eb71815 fix(docs): align current guidance with runtime 2026-07-19 16:11:17 -03:00
AkitaOnRails fad242a4fd docs: catch up with the v0.8.x feature surface
Stale-docs audit across the operator-facing surface. None of the
underlying behaviour changed — these are doc fixes for behaviour that
already shipped over the last week of PRs (#60 move-project,
#65 base-path, #68 wikilinks, #70 openai-compat strict, plus the
audit-cleanup follow-ups in the 65682dc…0d32be1 range).

Operator docs
- `README.md`: status badge bumped from v0.2 to v0.8, project
  derivation paragraph rewritten to describe the CLI's main-repo-root
  resolution (worktrees share one project) vs the hook router's
  basename default, write-page example switched to the H1-in-body
  convention (passing `--title` still works but invites issue #67's
  JSON-escape footgun).
- `docs/install.md`: bootstrap `--project` default is now described as
  "derived from cwd" instead of the stale `"scratch"`. Added an
  "Optional serve flags" table covering `--base-path`, `--web-slug`,
  `--web-ui-dir`, `--cors-allow-origin`.
- `docs/deploy.md`: cross-link to the "Hosting under a subpath"
  section in `https-via-proxy.md`.

Admin / lifecycle docs
- `docs/admission-webhooks.md`: payload sample includes the new
  `partial_failure` field (purge-project only, skipped on the wire
  when false). Spelled out that admission now fires BEFORE the SQL
  destruction in both `/admin/purge-project` and the
  `/admin/move-project` copy-purge path, so a `Reject` webhook on
  `purge_project` leaves the source intact. Documented that
  copy-purge fires two distinct webhook events from one request.
- `docs/lifecycle-ops.md`: failure-mode table covers
  `WikiError::DestinationExists` (409) and the block-policy 409 with
  `conflicts` body. `--on-conflict` documented alongside the JSON
  `on_conflict` field for direct `/admin/move-project` callers.
  Dedup naming cites the `DEDUP_FROM_TOKEN` const so the doc
  round-trips with the source. Safety-matrix `move-project` row
  notes the reject-policy escape hatch.

Frontend / web docs
- `docs/frontend-api.md`: removed two stale "Known gaps" claims —
  Cache-Control + ETag and CORS both ship today; §5 documents the
  cache headers and §9 is a new CORS section. Added `/api/v1/graph`
  to the endpoint reference (§4.9). §6 covers the base-path
  normaliser's safety rules (dot-segment rejection, fall-back-to-root
  on unsafe chars, query-preserving trailing-slash redirect).
- `docs/usage.md`: wikilink description spells out the "literal in
  code" guarantee (fence-glyph tracking, indented code, inline code)
  and the external-scheme allowlist; cross-links to the base-path
  flag docs.
- `docs/https-via-proxy.md`: subpath section now mentions the
  dot-segment rejection and the silent root-fallback warning.

LLM / MCP docs
- `docs/llm-provider-comparison.md`: new section on
  `AI_MEMORY_LLM_COMPAT_STRICT`, including the narrowed parse-shape
  fallback (5xx / auth / transport errors propagate now), the
  two-call cost of fallback, and a per-engine recommendation table.
  "When to revisit" entry on strict-JSON-schema availability rewritten
  to reflect that the feature exists today.
- `docs/mcp-install.md`: verify-it-works tool list now includes
  `memory_read_page` and `memory_delete_page` (their omission would
  have made users think their install was broken).

863 workspace tests still pass; fmt + clippy clean. No code touched.
2026-06-02 17:16:28 -03:00
AkitaOnRails ac425c4202 fix(serve): harden base-path web hosting 2026-06-02 14:06:46 -03:00
AkitaOnRails 1c8bb1aa18 docs: HTTPS-via-proxy guide + compose templates + sharpened startup warning
Per Phase 2 decision to NOT terminate TLS in-binary: ai-memory
stays HTTP-on-loopback by default (zero change for existing users
on upgrade) and operators front it with a battle-tested reverse
proxy when they cross the deployment shapes that need TLS.

Three new files:

- docs/https-via-proxy.md — the deployment guide. Covers:
  - When you DON'T need TLS (loopback + stdio cases honestly
    don't, called out up front so we don't add ceremony where
    it doesn't earn its keep).
  - When you DO need TLS (multi-user, non-loopback bind, /web
    from another machine, public exposure).
  - Five deployment paths with copy-paste configs:
    - Caddy + public domain + Let's Encrypt
    - Caddy + internal CA (LAN-only)
    - Cloudflare Tunnel (no open ports, TLS at the edge)
    - External cert files (Caddy or nginx)
    - Native (non-Docker) Caddy
  - Each path documents what can go wrong + the trust-install
    step for the internal-CA case (the load-bearing manual
    step — skipping it produces security theatre).
  - Final section "Don't paper over the security gap" calls out
    three specific anti-patterns operators reach for when
    things "don't work" (disable allowed-hosts, --insecure on
    clients, --no-tls-verify on cloudflared).

- docker/compose.tls.caddy.yml — copy-paste compose template
  with Caddy front. Three variants documented inline (LE,
  internal CA, external cert files). Same service definition
  across all three; only the Caddyfile differs.

- docker/compose.tls.cloudflared.yml — Cloudflare Tunnel sidecar
  compose template. Walks through the one-time CF dashboard
  setup, the .env.production shape (incl. CLOUDFLARE_TUNNEL_TOKEN),
  and the client-config commands. No ports exposed; the tunnel is
  outbound-only.

Code change (small):

- crates/ai-memory-cli/src/commands/serve.rs: extend the
  existing non-loopback startup warning. Previously fired only
  when bind wasn't loopback AND no auth token was configured;
  now ALSO fires (with different message) when bind isn't
  loopback AND auth IS configured — to remind the operator
  that bearer tokens still travel cleartext on plain HTTP and
  point them at docs/https-via-proxy.md. One-shot at startup,
  not a refusal to serve (operators may already be behind a
  proxy we can't reliably detect; refusing to bind would break
  their flow). Single line of log output, links to the doc.

Cross-linking:

- README "Security" section gets a clear "Want HTTPS?"
  paragraph pointing at docs/https-via-proxy.md + the two
  compose templates. Spells out the loopback exception so the
  single-user happy-path user doesn't think they're missing
  something.
- README docs table updated: docs/deploy.md row reframed as
  "pointers to the TLS guide"; new row added for
  docs/https-via-proxy.md with a full description of what it
  covers.
- docs/deploy.md's brief "encrypted transport" section
  trimmed from inline two-paragraph table to pointers at the
  new dedicated doc + the compose templates. Avoids drift
  between two places documenting the same thing.

Why this shape (instead of in-binary TLS):

- Reverse proxies are the production-grade answer anyway.
- Caddy does Let's Encrypt better than we ever would; we don't
  have to chase rcgen / instant-acme / rustls CVEs.
- Scope discipline — TLS termination is infrastructure, not
  ai-memory's job. Bundling them would repeat cognee's
  LiteLLM mistake at a smaller scale.
- The four "thinking you're secure when you're not" failure
  modes all become the proxy's problem, not ours: silent ACME
  expiry, trust-store dance, cert reload races, fail-open
  fallback.

No behavior change for existing single-user installs.
720/720 tests pass.
2026-05-30 12:36:21 -03:00