Add a FutureInfra example to the openai-compat section of docs/install.md, list it in docs/llm-providers.md, and add an [Unreleased] CHANGELOG entry. Docs only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
24 KiB
LLM Providers
Provider configuration for consolidation and embeddings. Moved verbatim from the README front page.
ai-memory runs without an LLM: hooks still capture sessions, search uses
FTS5 + declared entities + graph neighbors, and summaries fall back to
rule-based output. Add an LLM provider
when you want LLM consolidation (on PreCompact, on demand via
memory_consolidate, or opt-in at session end with
AI_MEMORY_CONSOLIDATE_ON_SESSION_END), richer linting, and bootstrap.
Substantive session ends always write a rule-based summary page + handoff either
way. A session containing only SessionStart / SessionEnd boundaries is
closed without a page, handoff, or provider job. When that empty session had
accepted startup context, its session-bound handoff is returned to the open
pool for the next receiver instead of being lost.
When the session-end opt-in is enabled, provider work is durably queued after
those deterministic writes and handled by one bounded server worker, so hook
drain latency does not cancel it. Failed jobs retry with backoff and survive a
server restart. A resumed native session is ended again only after its
observation generation advances; the persisted generation watermark makes
duplicate SessionEnd delivery and system clock skew converge without repeated
provider work. The end watermark and automatic handoff commit atomically, and
an interrupted keyed replay finishes the wiki commit, queue insert, and key
completion without duplicating that handoff. On the next SessionStart, the
newest cwd-eligible automatic handoff wins; accepting it expires older eligible
automatic handoffs without consuming manual or sibling-directory work. A new
automatic handoff also expires prior open automatic handoffs from its exact
cwd, so repeated SessionEnds cannot accumulate there before a receiver starts.
To keep consolidation style project-specific, write
_prompts/consolidation.md in that project's wiki. Its body can express
preferences such as "prefer Portuguese titles" or "omit routine CI noise".
Automatic, single-page, and multi-page consolidation use the page; a manual
memory_consolidate call can pass instructions to override it once. ai-memory
sanitizes and caps the value at 2,000 characters, JSON-encodes it in the user
message, and treats it as untrusted advisory data. It cannot supply facts,
request tool use or disclosure, or override the consolidation schema and
faithfulness rules. TTL-expired preference pages are ignored. With no active
page or argument, no preference block is appended.
Recommended defaults:
| Provider | Default | Use when |
|---|---|---|
anthropic |
claude-haiku-4-5 |
Best default for consolidation quality and rule classification. |
anthropic-oauth |
claude-sonnet-4-6 |
Use a Claude Pro/Max subscription via claude setup-token, no API key. |
openai |
gpt-5.4-mini |
Cheaper and faster hosted option. |
openai-oauth |
gpt-5.5 |
ChatGPT Pro/Plus/Codex backend via ai-memory auth login openai-oauth; no Platform API key. |
codex |
gpt-5.6-luna |
Reuse the Codex CLI-owned auth.json; access-token refresh remains owned by codex app-server. |
copilot |
gpt-5.5 |
GitHub Copilot Chat backend via ai-memory auth login copilot or COPILOT_GITHUB_TOKEN; requires a Copilot subscription. |
gemini |
gemini-3.5-flash |
Google-hosted option with a generous free tier. |
opencode |
claude-sonnet-4-6 |
OpenCode Go or Zen via OPENCODE_API_KEY. Go is the default endpoint; AI_MEMORY_LLM_BASE_URL selects Zen. Set AI_MEMORY_LLM_MODEL to an id the chosen endpoint serves. |
openai-compat |
no default | OpenRouter, Atlas Cloud, OrcaRouter, Cheaper Inference, API Route, FutureInfra, Ollama, vLLM, LM Studio, and other compatible endpoints. |
openai-compat + AI_MEMORY_LLM_BASE_URL=https://openrouter.ai/api/v1 |
no default (recommended: anthropic/claude-haiku-4.5) |
Hosted access to a large model catalogue through one key. See OpenRouter below and the empirical comparison in llm-provider-comparison.md. |
openai-oauth stores a refresh token in <data_dir>/auth.json and talks to
the ChatGPT/Codex Responses backend, not api.openai.com. For Docker quick
starts, run ai-memory auth login openai-oauth with the wrapper so the token
lands in the same ai-memory-data volume as the server.
codex is independent from openai-oauth: it never copies credentials into
ai-memory's data directory and never reads a refresh token or ID token. It
reads only tokens.access_token and tokens.account_id from
$CODEX_HOME/auth.json, or from the platform home's .codex/auth.json when
CODEX_HOME is unset or empty. Each request reloads the file. A first 401 may
trigger one serialized codex app-server --stdio recovery followed by one
retry; further 401 responses fail with a reauthentication hint. Override the
binary with AI_MEMORY_CODEX_EXECUTABLE when codex is not on PATH.
export AI_MEMORY_LLM_PROVIDER=codex
export AI_MEMORY_LLM_MODEL=gpt-5.6-luna
export AI_MEMORY_LLM_REASONING_EFFORT=medium
ai-memory llm-test --provider codex --model gpt-5.6-luna --prompt "Reply with OK"
ai-memory llm-test --provider codex --model gpt-5.6-luna --structured --prompt "Return a short answer"
Codex credential storage mode file is supported. auto works only when its
effective credential is present in auth.json; keyring-only and ephemeral
credentials are not read and produce an actionable missing-auth-file error.
Docker is not configured automatically: the Codex executable, CODEX_HOME,
and its credential file must all exist in the same container/environment as
ai-memory.
opencode talks to OpenCode's gateway with the sk-... key from
opencode.ai/auth, read from OPENCODE_API_KEY only; it does not fall back
to LLM_API_KEY. It defaults to the Go endpoint,
https://opencode.ai/zen/go/v1: a subscription with a per-model monthly
allowance rather than per-token billing. Set
AI_MEMORY_LLM_BASE_URL=https://opencode.ai/zen/v1 (or llm_base_url in
config.toml) for Zen's pay-per-token catalogue. Model ids are per
catalogue and written plainly (mimo-v2.5, glm-5.3-flash), not in the
opencode-go/<model> form OpenCode's own client config uses. The built-in
default is claude-sonnet-4-6; Go ids such as mimo-v2.5 or glm-5.3-flash
come from AI_MEMORY_LLM_MODEL, so set it to an id the endpoint you chose
serves. gpt-5.6-luna is sent through the Responses endpoint;
every other model uses Chat Completions. The provider sends the
x-opencode-session correlation header OpenCode asks for, one id per logical
operation, and its own User-Agent; AI_MEMORY_LLM_HEADERS overrides
either. AI_MEMORY_LLM_REASONING_EFFORT and AI_MEMORY_LLM_TIMEOUT_SECS
apply as for every other provider. AI_MEMORY_LLM_PROVIDER / llm_provider
also accept the historical aliases opencode-zen and opencode_zen for the
same provider (llm-test --provider takes only opencode); the endpoint is
chosen by the base URL, not the alias.
export AI_MEMORY_LLM_PROVIDER=opencode
export AI_MEMORY_LLM_MODEL=mimo-v2.5
ai-memory llm-test --provider opencode --model mimo-v2.5 --prompt "Reply with OK"
OpenRouter is used through the generic openai-compat
provider — the same wiring as Atlas Cloud, OrcaRouter, Ollama, vLLM, and LM
Studio. There is no dedicated openrouter provider name, and setting
AI_MEMORY_LLM_PROVIDER=openrouter fails at startup with
AI_MEMORY_LLM_PROVIDER=openrouter is not one of anthropic|openai|gemini|openai-compat|....
Select it by pointing the compat base URL at OpenRouter and supplying its
sk-or-v1-... key through LLM_API_KEY:
A minimal working env-file (drop into ~/.config/ai-memory/openrouter.env
or the Docker --env-file argument):
LLM_API_KEY=sk-or-v1-...
AI_MEMORY_LLM_PROVIDER=openai-compat
AI_MEMORY_LLM_BASE_URL=https://openrouter.ai/api/v1
AI_MEMORY_LLM_MODEL=deepseek/deepseek-v4-flash
AI_MEMORY_LLM_COMPAT_STRICT=true
Then verify:
ai-memory llm-test --provider openai-compat \
--model deepseek/deepseek-v4-flash --prompt "Reply with OK"
AI_MEMORY_LLM_COMPAT_STRICT=true sends OpenAI-style
response_format=json_schema for every structured call. Modern hosted
gateways honour it; some model backends (DeepSeek among them) otherwise wrap
their JSON output in markdown fences that ai-memory's tolerant parser has to
strip, which occasionally drops a field. Setting it to true at the
env-file level is safer for consolidation reliability. Turn it off
(AI_MEMORY_LLM_COMPAT_STRICT=false) only for a specific model backend that
explicitly rejects the schema-constrained shape.
The docker/.env.production.example file in the repo ships with an OpenRouter
setup pre-filled (with moonshotai/kimi-k2.6 as the sample model — see the
gotchas below for why reasoning models are ineligible). Model ids follow
OpenRouter's provider/model convention. There is no built-in default:
AI_MEMORY_LLM_MODEL is required, as it is for every openai-compat endpoint.
Choosing a model
For LLM selection guidance, see
llm-provider-comparison.md — the project
maintains its own A/B testing there on consolidation quality, latency, and
cost. That comparison uses English consolidation fixtures and covers the
recommended default (anthropic/claude-haiku-4.5), the cheaper alternative
(openai/gpt-5.4-mini), and the reasoning-model ineligibility note; the
OpenRouter subsection here defers to it rather than duplicating rankings.
For workloads dominated by non-English content, community reports suggest
that a higher-recall multilingual embedder (for example baai/bge-m3 via
the openai-compat embedder) noticeably improves retrieval quality
relative to the built-in English-only local embedder. Configuration is in
the embedding paragraph below.
ai-memory automatically layers OpenRouter's app-attribution headers
(HTTP-Referer and X-Title) whenever the base URL points at
openrouter.ai, so ai-memory traffic is credited on OpenRouter's app
leaderboard. AI_MEMORY_LLM_HEADERS=HTTP-Referer=https://example.com,X-Title=my-app
overrides them per operator. AI_MEMORY_LLM_REASONING_EFFORT is honoured on
this path: OpenRouter hosts receive reasoning: { effort, exclude: true }.
OpenRouter's LLM/completions API is what ai-memory uses above. It does not
reliably expose an OpenAI-compatible /v1/embeddings endpoint — historically it
has not — so do not assume an OpenRouter base URL works for the embedder. If a
given deployment of OpenRouter does offer an embeddings endpoint, verify it
answers the standard /v1/embeddings shape (curl the endpoint directly)
before relying on it. When it does, configure the embedder through the
openai-compat path — not the openai path, which has an OpenAI-specific key
precedence chain:
export AI_MEMORY_EMBEDDING_PROVIDER=openai-compat
export AI_MEMORY_EMBEDDING_BASE_URL=https://openrouter.ai/api/v1
export AI_MEMORY_EMBEDDING_MODEL=baai/bge-m3
export AI_MEMORY_EMBEDDING_DIM=1024
# LLM_API_KEY is reused automatically when EMBEDDING_API_KEY is unset on
# the openai-compat embedder path; no second key is required.
AI_MEMORY_EMBEDDING_DIM is mandatory for openai-compat embeddings and must
match the model. For English-only content the built-in local
sentence-transformer (AI_MEMORY_EMBEDDING_PROVIDER=local, see
local-embeddings.md) is a valid alternative that
avoids hosted egress; for non-English content, a multilingual hosted model
(baai/bge-m3, intfloat/multilingual-e5-large, or similar) served by any
provider that exposes a real /v1/embeddings endpoint is the practical choice.
The two endpoints are still independent, so setting AI_MEMORY_LLM_BASE_URL
alone does not redirect embeddings — set AI_MEMORY_EMBEDDING_BASE_URL too.
Security and gotchas
- Bind address vs. port publish. When running inside a container, the
process-level bind (
--bind 0.0.0.0:49374or similar) does not decide who can reach the port — that is decided by the container's publish spec.-p 127.0.0.1:49374:49374restricts to loopback on the host;-p 49374:49374publishes on every host interface. If the port is ever reachable beyond loopback, generate a bearer token first:The server logs a loud startup warning about non-loopback exposure without a token; treat it as blocking.ai-memory generate-auth-token export AI_MEMORY_AUTH_TOKEN=<token> export AI_MEMORY_ALLOWED_HOSTS=<host,host,...>AI_MEMORY_ALLOWED_HOSTSguards against DNS-rebinding. TLS termination is expected to sit in front of the process — seedocs/https-via-proxy.md. - Secret hygiene. The env file that holds
LLM_API_KEY(and any provider-specific key you set) should bechmod 600and outside any directory a version-control tool sees. If you back up~/.config/ai-memory/(seedocs/backup.md), exclude the env file — the sample script and mirror.gitignoreshipped there already exclude*.envby default. - Cheap models can drift. LLM consolidation writes durable wiki
pages, so a model that occasionally fabricates a
kind: decisionpage for a session that made no decision leaves a page in the wiki that ai-memory will treat as first-class evidence on subsequent reads. This is a known cost of running a smaller / cheaper model on the consolidation prompt. Spot-check<data_dir>/wiki/<workspace>/<project>/decisions/during the first week on a new model for pages you do not recognize; if any appear, either raise the provider tier for consolidation only (keep the cheaper model for retrieval reranking) or unsetAI_MEMORY_LLM_PROVIDERto fall back to the deterministic rule-based path. :free-suffix models draw from a shared quota. Any model id ending in:freeon OpenRouter is served from a pool shared across all users of the free tier, not the caller's paid budget. When that pool is exhausted, requests fail with an upstream-shared-pool rate-limit error. Fine for evaluation, unreliable for production; prefer a model listed with multiple upstream providers on its OpenRouter page.- Reasoning-mode models are ineligible for consolidation.
Consolidation and lint use a strict-JSON output contract; reasoning
models route most of their token budget through internal reasoning
and either stall past the 300 s per-request ceiling or return empty
strings.
docs/install.md#llm-provider-tiershas the standing "what we don't recommend" note listing the pattern; the same applies to any OpenRouter model flagged as a reasoning/thinking variant. Prefer a non-reasoning peer forAI_MEMORY_LLM_MODEL.
anthropic-oauth hits the same /v1/messages endpoint as anthropic but
authenticates with an OAuth bearer token instead of an API key. Run
claude setup-token once, then set AI_MEMORY_LLM_PROVIDER=anthropic-oauth and
ANTHROPIC_OAUTH_TOKEN=<token> (or CLAUDE_CODE_OAUTH_TOKEN, which claude setup-token writes automatically). No ANTHROPIC_API_KEY is needed. The Docker
wrappers forward either token by name to short-lived helper commands such as
llm-test; configure the long-lived server container separately as shown in the
installation guide.
For both Anthropic providers, ai-memory omits temperature for Claude
4.7 and later models and Claude Mythos Preview because those models reject
non-default sampling parameters. llm-test sends the same representative 0.2
value as the normal pipeline before the provider applies that compatibility
rule.
⚠️ Unofficial and against Anthropic's usage policies — use at your own risk;
it may get your account rate-limited or banned. See
the warning in docs/install.md.
copilot stores a GitHub user token in the same auth file, exchanges it for a
short-lived Copilot API token via GitHub's /copilot_internal/v2/token, and
uses the Copilot Chat endpoint with vscode-chat integration headers. You can
also set COPILOT_GITHUB_TOKEN, GH_TOKEN, or GITHUB_TOKEN on the server.
Tip
For the OAuth/subscription backends, prefer a small, fast model via
AI_MEMORY_LLM_MODELwhere the backend lets you choose one. ai-memory's LLM work (consolidation, lint, explore) is summarisation, not hard reasoning, so a Haiku/mini-class model is plenty and is much easier on subscription rate limits. Save the high-effort thinking models for your coding agent. Per backend:
anthropic-oauth: setclaude-haiku-4-5.openai-oauth/codex: leave the provider default (gpt-5.5). The Codex/ChatGPT backend only accepts a small server-defined set of model ids and rejects others (e.g.gpt-5-mini) with a deterministic 400, so do not override the model here.copilot: a mini-class id such asgpt-5-minimay work, but Copilot's accepted model set is unverified — check before relying on it, and fall back to the default if the endpoint rejects your choice.
Tip
OpenAI-compatible structured output is schema-constrained by default. ai-memory sends each operation's JSON Schema through
response_format=json_schema, which recent Ollama, vLLM, LM Studio, and llama.cpp releases honour. It falls back to the tolerant parser when an endpoint explicitly rejects that field or returns a malformed shape. SetAI_MEMORY_LLM_COMPAT_STRICT=falseonly for an incompatible endpoint.
For small-context local models, configure both consolidation limits. The input target accounts for the complete rendered prompt, including bounded slot and current-page context plus the structured-output schema; the output limit is sent to the provider. Their sum must fit the model context window, with extra headroom because provider tokenizers differ:
[consolidation]
max_input_tokens = 6500
max_output_tokens = 1000
The equivalent environment variables are
AI_MEMORY_CONSOLIDATION__MAX_INPUT_TOKENS and
AI_MEMORY_CONSOLIDATION__MAX_OUTPUT_TOKENS. Provider failures during an
automatic PreCompact/PostCompaction checkpoint fall back to the deterministic
rule-based page; admission, storage, and scope errors still fail closed. The
validated minimums are 6,000 input and 1,000 output tokens.
Every chat provider bounds each completion request at 300 seconds
(AI_MEMORY_LLM_TIMEOUT_SECS to override, or llm_timeout_secs = 900 in
config.toml; the quick openai-oauth token refresh keeps the default ceiling).
The default tolerates a local engine cold-loading a large model; slow hosted
gateways whose long completions exceed the ceiling fail every request with
http: error sending request, so raise the value there instead of watching
consolidation exhaust its retries.
Reranking is optional and off by default. With an LLM provider configured,
AI_MEMORY_RERANKER=llm makes project and explicit-scope memory_query
calls over-fetch from the hybrid stage, fuse scopes, and make at most one LLM
call to reorder the best candidates. This can promote a relevant page that
RRF ranked below the requested cut, at the cost of LLM latency and usage. The
request sends the query plus at most 30 bounded page titles and search snippets
to the configured provider; all values are JSON-encoded and treated as
untrusted data. A timeout, provider error, or incomplete/invalid score set
preserves the normal order. global=true and supplemental global-preference
hits keep their existing non-RRF ranking. Concurrent provider calls are capped
at four; saturated queries keep their local ranking without waiting.
Embeddings are optional and separate from the LLM provider. Set
AI_MEMORY_EMBEDDING_PROVIDER=openai, voyage, google/gemini,
openai-compat, or copilot when you want vector retrieval in addition to
FTS5 + entity + graph-neighbor retrieval. openai-compat targets self-hosted
engines (Ollama, LM Studio, vLLM): it needs no API key and requires explicit
AI_MEMORY_EMBEDDING_BASE_URL, AI_MEMORY_EMBEDDING_MODEL, and
AI_MEMORY_EMBEDDING_DIM. The optional EMBEDDING_API_KEY credentials the
embedder alone and is checked before OPENAI_API_KEY and LLM_API_KEY, so
embeddings can run on a different provider than the LLM. Because the two
endpoints are independent, AI_MEMORY_LLM_BASE_URL redirects only the LLM;
set AI_MEMORY_EMBEDDING_BASE_URL as well or embedding traffic still goes to
the embedding provider's default endpoint. Both the FTS-only and
hybrid paths apply the same bounded page-authority adjustment after candidate
generation; embeddings improve relevance recall but do not decide which source
is canonical.
Asymmetric embedding models — ones trained with a different instruction for
queries than for the text being indexed — need that instruction added to the
text itself; the OpenAI-compatible /v1/embeddings wire format has no field
for it. AI_MEMORY_EMBEDDING_QUERY_PREFIX and
AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX (only applied by the openai and
openai-compat embedders) are prepended to query and document text
respectively, before the existing truncation, so a long body is still
truncated to the same overall input cap with the prefix included. Both are
empty by default — no behaviour change when unset — and are not trimmed, so
a publisher's trailing space is preserved exactly. For example, NVIDIA's
nvidia/Nemotron-3-Embed-1B-BF16
(2048-dim, served locally through vLLM/openai-compat) specifies
"query: " for queries and "passage: " for documents per its model card:
embedding_provider = "openai-compat"
embedding_model = "nvidia/Nemotron-3-Embed-1B-BF16"
embedding_base_url = "http://localhost:8000/v1"
embedding_dim = 2048
embedding_query_prefix = "query: "
embedding_document_prefix = "passage: "
Base E5 models (intfloat/e5-base-v2, e5-large-v2, multilingual E5, …)
use the same "query: " / "passage: " convention. Instruction-tuned E5
variants and Qwen3-Embedding instead need a task-instruction string on the
query side only — their documents are embedded plain, with no document
prefix — but the two use different exact spacing, confirmed against
each model card:
intfloat/e5-mistral-7b-instruct:embedding_query_prefix = "Instruct: {task description}\nQuery: "— a trailing space afterQuery:.Qwen/Qwen3-Embedding-0.6B(and the other Qwen3-Embedding sizes):embedding_query_prefix = "Instruct: {task description}\nQuery:"— no trailing space; the query text follows the colon directly.
Fill in your own task description for {task description}, leave
embedding_document_prefix unset for both, and don't copy one model's
exact string for the other — the trailing-space difference is
publisher-specified, not a typo.
Changing embedding_query_prefix alone never requires re-embedding: only
embedding_document_prefix is folded into the stored embedding identity
(a non-empty value changes what ai-memory treats as the current
{provider, model, dim} triple for that page's vectors), so a document
prefix change makes memory_query stop matching the old vectors and the
next ai-memory embed (scheduled or manual) pass re-embed them
automatically — no --force needed. --force remains useful to
re-embed everything immediately rather than waiting for scheduled
backfill, or to force a re-embed for a reason the triple alone doesn't
capture (e.g. touching up the vLLM server's tokenizer or quantization).
AI_MEMORY_EMBEDDING_PROVIDER=copilot reuses the same Copilot OAuth login as
the copilot LLM provider (ai-memory auth login copilot or
COPILOT_GITHUB_TOKEN/GITHUB_COPILOT_API_TOKEN) — no separate API key.
It defaults to model text-embedding-3-small, dim 1536, and calls Copilot's
/embeddings endpoint following the same OpenAI-compatible contract Copilot
documents for chat. That endpoint's exact shape is not covered by a live
test against Copilot here; treat it as needing a real-Copilot smoke test
before relying on it in production.
AI_MEMORY_EMBEDDING_PROVIDER=local needs no key and no server at all:
sentence embeddings run in-process (pure-Rust all-MiniLM-L6-v2,
384-dim), with the model fetched once into <data_dir>/models/ under
pinned checksums — see docs/local-embeddings.md.
See docs/install.md#llm-provider-tiers
for env vars and Ollama/OpenRouter/Atlas Cloud/OrcaRouter/Cheaper Inference/API Route/FutureInfra
examples, and
docs/llm-provider-comparison.md
for the empirical model comparison.