mirror of
https://github.com/akitaonrails/ai-memory.git
synced 2026-10-02 03:24:46 +08:00
fix(embedding): preserve prefix whitespace through env loading; regression tests
figment's Env provider parses each var's string with its loose-value
parser, whose bare/unquoted branch calls .trim() (figment 0.10.19's
src/value/parse.rs:78, verified against the vendored source) — so
AI_MEMORY_EMBEDDING_QUERY_PREFIX="query: " reached Config::embedding_query_prefix
as "query:", silently dropping the publisher-significant trailing space.
Config::load now overlays the raw (untrimmed) env value for the two prefix
keys via figment::providers::Serialized (which hands figment an
already-typed value, bypassing the string parser), merged after the
generic Env::prefixed pass so it still wins over a config.toml value.
Presence, not non-emptiness, is the signal: a variable set to "" is a
deliberate override clearing a config.toml-configured prefix, distinct
from the variable being absent. Both wrapper scripts (bin/ai-memory,
bin/ai-memory.ps1) now forward these two keys on presence for the same
reason; a non-empty override was already forwarded correctly, only the
empty-override case was silently dropped by the [ -n ] check every other
forwarded var correctly uses.
The overlay itself lives in overlay_embedding_prefixes(), a pure function
taking the env values as parameters rather than reading std::env::var
itself, so it stays directly unit-testable without mutating process
environment or the current directory — the same pattern
ai-memory-cli/src/commands/path_util.rs's agent_config_home and
ai-memory-hooks's drain_with_live_token already use. std::env::set_var is
unsafe under edition 2024 and forbidden workspace-wide, and
figment::Jail calls std::env::set_current_dir on the real process
internally, racing every other test in this crate's multi-threaded lib
test binary that relies on cwd. Four pure unit tests build a minimal
in-memory Figment and assert on the merged Config directly. One
additional test exercises the real Config::load end to end through a
TOML file and an explicit absolute path; since this process's own
environment is shared with every other test in the binary and could
already carry one of the two prefix vars from the test runner's shell,
that test re-execs this same test binary filtered to just itself as a
genuinely separate child process, with both vars removed via
Command::env_remove — real process isolation rather than an in-process
assumption about the ambient environment.
Plus: a wiremock transport test proving OpenAiEmbedder (not just
OpenAiCompatEmbedder) sends the configured prefix on the wire; a
regression test for the memory_query embed_query fix from the prior
commit, using a task-aware fixture embedder whose
embed/embed_document/embed_query methods return distinguishable vectors
so a regression back to calling the wrong one fails a direct equality
assertion rather than an inferred ranking change; fake-Docker argument
tests for the wrapper's presence-based forwarding across
unset/empty/whitespace-only/non-empty, with both env vars explicitly
removed from the child environment before each case so the "unset" case
cannot silently inherit an ambient export from the test runner.
bin/ai-memory.ps1 also notes, briefly, the PowerShell/.NET version an
operator needs for `$env:NAME = ''` to reach the wrapper as set-but-empty
rather than deleted — no pwsh runtime was available to exercise it live.
Also narrows the query/document-prefix docs (config.rs, docs/llm-providers.md,
CHANGELOG.md) with exact, separate templates: Nemotron-3-Embed
(nvidia/Nemotron-3-Embed-1B-BF16) and base E5 use "query: "/"passage: ";
e5-mistral-7b-instruct wants "Instruct: {task}\nQuery: " (trailing space);
Qwen3-Embedding wants "Instruct: {task}\nQuery:" (no trailing space) — both
leave documents plain.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
0fff203e8c
commit
dbf44907ac
+12
-9
@@ -12,15 +12,18 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
||||
`AI_MEMORY_EMBEDDING_QUERY_PREFIX` / `AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX`)
|
||||
for the `openai` and `openai-compat` embedders: an optional string
|
||||
prepended to query / document text before the existing truncation, so
|
||||
truncation still bounds the whole input. Asymmetric embedding models —
|
||||
Nemotron-3-Embed, the E5 family, Qwen3-Embedding — need a `"query: "` /
|
||||
`"passage: "` instruction their publisher specifies; the
|
||||
OpenAI-compatible `/v1/embeddings` wire format has no field for it.
|
||||
Empty by default — no behaviour change when unset, and not trimmed, so a
|
||||
publisher's trailing space is preserved. Changing a prefix does not
|
||||
change the stored `{provider, model, dim}` triple pages are keyed by;
|
||||
run `ai-memory embed --force` to re-embed after changing one. See
|
||||
`docs/llm-providers.md`.
|
||||
truncation still bounds the whole input. Asymmetric embedding models need
|
||||
a query-side instruction their publisher specifies; the OpenAI-compatible
|
||||
`/v1/embeddings` wire format has no field for it.
|
||||
`nvidia/Nemotron-3-Embed-1B-BF16` and base E5 models use a simple
|
||||
`"query: "` / `"passage: "` pair; instruction-tuned E5 variants and
|
||||
Qwen3-Embedding instead need a task-instruction string on the query side
|
||||
only (documents stay plain). Empty by default — no behaviour change when
|
||||
unset, and not trimmed (nor is a present-but-empty env-var override,
|
||||
which now clears a `config.toml` value), so a publisher's trailing space
|
||||
is preserved. Changing a prefix does not change the stored
|
||||
`{provider, model, dim}` triple pages are keyed by; run `ai-memory embed
|
||||
--force` to re-embed after changing one. See `docs/llm-providers.md`.
|
||||
|
||||
### Fixed
|
||||
- `memory_query`'s vector stream called the generic `Embedder::embed`
|
||||
|
||||
+12
-2
@@ -586,8 +586,6 @@ for var in \
|
||||
AI_MEMORY_EMBEDDING_MODEL \
|
||||
AI_MEMORY_EMBEDDING_BASE_URL \
|
||||
AI_MEMORY_EMBEDDING_DIM \
|
||||
AI_MEMORY_EMBEDDING_QUERY_PREFIX \
|
||||
AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX \
|
||||
AI_MEMORY_ALLOWED_HOSTS \
|
||||
AI_MEMORY_WORKSTREAM_ID \
|
||||
AI_MEMORY_HOOK_PLATFORM \
|
||||
@@ -614,6 +612,18 @@ do
|
||||
fi
|
||||
done
|
||||
|
||||
# These two are presence-based, not non-empty-based like the loop above: an
|
||||
# operator sets one to the empty string to deliberately clear a
|
||||
# config.toml-configured prefix without editing the file (see
|
||||
# Config::load's figment overlay), and a present-but-empty value must reach
|
||||
# the container for that to work — `[ -n ]` above would drop it, making the
|
||||
# wrapper indistinguishable from the var never having been set at all.
|
||||
for var in AI_MEMORY_EMBEDDING_QUERY_PREFIX AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX; do
|
||||
if [ -n "${!var+x}" ]; then
|
||||
ENV_ARGS+=(-e "${var}")
|
||||
fi
|
||||
done
|
||||
|
||||
# The wrapper itself runs the CLI inside a short-lived helper container, while
|
||||
# the README server runs in the long-lived ai-memory container and publishes
|
||||
# 127.0.0.1:49374 on the host. Inside a normal bridge-network helper,
|
||||
|
||||
+24
-2
@@ -150,8 +150,6 @@ foreach ($Name in @(
|
||||
"AI_MEMORY_EMBEDDING_MODEL",
|
||||
"AI_MEMORY_EMBEDDING_BASE_URL",
|
||||
"AI_MEMORY_EMBEDDING_DIM",
|
||||
"AI_MEMORY_EMBEDDING_QUERY_PREFIX",
|
||||
"AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX",
|
||||
"AI_MEMORY_ALLOWED_HOSTS",
|
||||
"AI_MEMORY_WORKSTREAM_ID",
|
||||
"CLAUDE_CONFIG_DIR",
|
||||
@@ -176,6 +174,30 @@ foreach ($Name in @(
|
||||
}
|
||||
}
|
||||
|
||||
# Presence-based, not non-empty-based like the loop above: an operator sets
|
||||
# one of these to the empty string to deliberately clear a
|
||||
# config.toml-configured prefix without editing the file (see
|
||||
# Config::load's figment overlay), and a present-but-empty value must reach
|
||||
# the container for that to work — `IsNullOrEmpty` above would drop it,
|
||||
# making the wrapper indistinguishable from the variable never having been
|
||||
# set at all. `GetEnvironmentVariable` returns `$null` only when the
|
||||
# variable is truly unset, and `""` when it is set-but-empty, so a `-ne
|
||||
# $null` check is exactly the presence test needed here.
|
||||
#
|
||||
# An operator's own `$env:NAME = ''` additionally needs PowerShell 7.5+
|
||||
# (first built on .NET 9) to leave a set-but-empty variable rather than
|
||||
# deleting it; not exercised on a real pwsh runtime.
|
||||
# https://learn.microsoft.com/en-us/dotnet/api/system.environment.setenvironmentvariable
|
||||
# https://learn.microsoft.com/en-us/powershell/scripting/whats-new/what-s-new-in-powershell-75
|
||||
foreach ($Name in @(
|
||||
"AI_MEMORY_EMBEDDING_QUERY_PREFIX",
|
||||
"AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX"
|
||||
)) {
|
||||
if ($null -ne [Environment]::GetEnvironmentVariable($Name)) {
|
||||
$DockerArgs += @("-e", $Name)
|
||||
}
|
||||
}
|
||||
|
||||
# Docker Desktop gives Windows no host networking for Linux containers, so a
|
||||
# thin-client command (status, search, bootstrap, ...) reaches the loopback-
|
||||
# published server from this helper container through Docker Desktop's host
|
||||
|
||||
@@ -444,22 +444,38 @@ pub struct Config {
|
||||
/// Optional prefix prepended to every embedding **query** before it is
|
||||
/// sent to the `openai` or `openai-compat` embedder, ahead of the
|
||||
/// existing truncation. Unset (the default) is a no-op — no behaviour
|
||||
/// change. Asymmetric self-hosted models — Nemotron-3-Embed, the E5
|
||||
/// family, Qwen3-Embedding — need a `"query: "` instruction their
|
||||
/// publisher specifies; the OpenAI-compatible `/v1/embeddings` wire
|
||||
/// format has no field for it, so the client prepends it instead. Not
|
||||
/// trimmed: a publisher's trailing space (e.g. `"query: "`) is
|
||||
/// significant and preserved verbatim. Ignored by `google` (which has
|
||||
/// its own built-in query/document asymmetry), `voyage`, `local`, and
|
||||
/// `copilot`. Changing this does not change the stored
|
||||
/// `{provider, model, dim}` triple pages are keyed by — run
|
||||
/// `ai-memory embed --force` to re-embed after changing it. See
|
||||
/// `docs/llm-providers.md`. Settable via
|
||||
/// `AI_MEMORY_EMBEDDING_QUERY_PREFIX`.
|
||||
/// change. Asymmetric self-hosted models need a query-side instruction
|
||||
/// their publisher specifies; the OpenAI-compatible `/v1/embeddings`
|
||||
/// wire format has no field for it, so the client prepends it instead.
|
||||
/// `nvidia/Nemotron-3-Embed-1B-BF16` and base E5 models
|
||||
/// (`intfloat/e5-base-v2`, multilingual E5, …) use a simple
|
||||
/// `"query: "` / `"passage: "` pair (documents get
|
||||
/// `embedding_document_prefix = "passage: "`). Instruction-tuned E5
|
||||
/// variants and Qwen3-Embedding instead need a full task-instruction
|
||||
/// string on the query side only, with **different exact spacing each**
|
||||
/// — leave `embedding_document_prefix` unset for both (their documents
|
||||
/// are plain text, no prefix):
|
||||
/// `e5-mistral-7b-instruct` wants
|
||||
/// `"Instruct: {task description}\nQuery: "` (a trailing space after
|
||||
/// `Query:`); Qwen3-Embedding wants
|
||||
/// `"Instruct: {task description}\nQuery:"` (no trailing space — the
|
||||
/// query text follows the colon directly). Not trimmed: a publisher's
|
||||
/// trailing space or embedded newline is significant and preserved
|
||||
/// verbatim. Ignored by `google` (which has its own built-in
|
||||
/// query/document asymmetry), `voyage`, `local`, and `copilot`.
|
||||
/// Changing only this key never requires re-embedding existing pages —
|
||||
/// the query side has no stored identity. See `docs/llm-providers.md`.
|
||||
/// Settable via `AI_MEMORY_EMBEDDING_QUERY_PREFIX` (figment's `Env`
|
||||
/// provider would otherwise trim a trailing space; `Config::load`
|
||||
/// overlays the raw env bytes for this key specifically).
|
||||
pub embedding_query_prefix: Option<String>,
|
||||
/// Document-side counterpart of `embedding_query_prefix` (e.g.
|
||||
/// `"passage: "` for Nemotron-3-Embed / E5). Settable via
|
||||
/// `AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX`.
|
||||
/// `"passage: "` for Nemotron-3-Embed / base E5; see that field's doc
|
||||
/// comment for which models this applies to). Changing this does not
|
||||
/// change the stored `{provider, model, dim}` triple pages are keyed
|
||||
/// by — run `ai-memory embed --force` to re-embed after changing it.
|
||||
/// Settable via `AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX` (same raw-env
|
||||
/// overlay as `embedding_query_prefix`).
|
||||
pub embedding_document_prefix: Option<String>,
|
||||
/// M8 retention-sweep parameters. The defaults give an ~80-day
|
||||
/// "survival floor" for unused episodic content (above the cold
|
||||
@@ -1338,6 +1354,19 @@ impl Config {
|
||||
figment = figment.merge(Toml::file(&resolved_config_path));
|
||||
}
|
||||
figment = figment.merge(Env::prefixed("AI_MEMORY_").split("__"));
|
||||
// The environment is read once, here, and passed down as data —
|
||||
// never inside `overlay_embedding_prefixes` itself — so that
|
||||
// function stays directly testable without mutating process env or
|
||||
// cwd (see its doc comment).
|
||||
figment = overlay_embedding_prefixes(
|
||||
figment,
|
||||
std::env::var("AI_MEMORY_EMBEDDING_QUERY_PREFIX")
|
||||
.ok()
|
||||
.as_deref(),
|
||||
std::env::var("AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX")
|
||||
.ok()
|
||||
.as_deref(),
|
||||
);
|
||||
|
||||
let mut config: Config = figment.extract().with_context(|| {
|
||||
format!(
|
||||
@@ -1904,7 +1933,11 @@ impl Config {
|
||||
None
|
||||
};
|
||||
// Not `non_empty`: that trims, and a publisher's trailing space
|
||||
// (e.g. Nemotron-3-Embed / E5's `"query: "`) is significant.
|
||||
// (e.g. Nemotron-3-Embed's `"query: "`) is significant. By the time
|
||||
// `Load` has run, `self.embedding_query_prefix` already holds the
|
||||
// exact configured bytes regardless of source (TOML or env) — see
|
||||
// `Config::load`'s env-prefix overlay, which corrects for
|
||||
// figment's `Env` provider trimming unquoted values.
|
||||
let query_prefix = self.embedding_query_prefix.clone().unwrap_or_default();
|
||||
let document_prefix = self.embedding_document_prefix.clone().unwrap_or_default();
|
||||
Ok(Some(EmbedderConfig {
|
||||
@@ -2122,6 +2155,51 @@ fn env_string(name: &str) -> Option<String> {
|
||||
})
|
||||
}
|
||||
|
||||
/// Overlay the two embedding-prefix keys onto `figment` with their raw,
|
||||
/// untrimmed values, whenever the corresponding parameter is `Some` (even
|
||||
/// `Some("")` — an operator clearing a `config.toml`-set prefix back to
|
||||
/// none via an empty env var; only `None`, the variable genuinely absent,
|
||||
/// leaves a `config.toml` value or the default untouched).
|
||||
///
|
||||
/// figment's `Env` provider parses each var's string as a loose value
|
||||
/// (`figment::value::parse::value`), and its bare/unquoted branch calls
|
||||
/// `.trim()` — so `AI_MEMORY_EMBEDDING_QUERY_PREFIX="query: "` would
|
||||
/// otherwise reach `embedding_query_prefix` as `"query:"`, silently
|
||||
/// dropping the publisher-significant trailing space (verified against
|
||||
/// figment 0.10.19's vendored source, `src/value/parse.rs:78`).
|
||||
/// [`Serialized`] values are handed to figment as already-typed data (via
|
||||
/// `serde::Serialize`), so they never pass through that string parser and
|
||||
/// so are never trimmed. Callers merge this after `Env::prefixed` so it
|
||||
/// wins over the (possibly trimmed) value that provider already set.
|
||||
///
|
||||
/// The values come in as parameters, already read by the caller, rather
|
||||
/// than this function reading `std::env::var` itself — the same pattern
|
||||
/// `ai-memory-cli/src/commands/path_util.rs`'s `agent_config_home` and
|
||||
/// `ai-memory-hooks`'s `drain_with_live_token` use, and for the same
|
||||
/// reason: it keeps this function directly unit-testable without
|
||||
/// mutating process environment or the current directory. Both are
|
||||
/// unsafe or actively harmful to do from a `#[test]` in this crate's
|
||||
/// multi-threaded lib test binary — `std::env::set_var` is `unsafe` under
|
||||
/// edition 2024 and forbidden workspace-wide, and even a "safe" wrapper
|
||||
/// such as `figment::Jail` still calls `std::env::set_current_dir` on the
|
||||
/// real process (verified against its vendored source,
|
||||
/// `src/jail.rs:141`), racing every other test in the binary that reads
|
||||
/// env or relies on cwd — e.g. `tests/suite/backfill_e2e.rs`'s
|
||||
/// `Command::current_dir` calls.
|
||||
fn overlay_embedding_prefixes(
|
||||
mut figment: Figment,
|
||||
query_prefix_env: Option<&str>,
|
||||
document_prefix_env: Option<&str>,
|
||||
) -> Figment {
|
||||
if let Some(v) = query_prefix_env {
|
||||
figment = figment.merge(Serialized::default("embedding_query_prefix", v));
|
||||
}
|
||||
if let Some(v) = document_prefix_env {
|
||||
figment = figment.merge(Serialized::default("embedding_document_prefix", v));
|
||||
}
|
||||
figment
|
||||
}
|
||||
|
||||
fn env_path(name: &str) -> Option<PathBuf> {
|
||||
env_string(name).map(PathBuf::from)
|
||||
}
|
||||
@@ -3253,7 +3331,7 @@ mod tests {
|
||||
// embedders see byte-identical behaviour to before this feature.
|
||||
let unset = Config {
|
||||
embedding_provider: Some("openai-compat".into()),
|
||||
embedding_model: Some("nvidia/nemotron-3-embed".into()),
|
||||
embedding_model: Some("nvidia/Nemotron-3-Embed-1B-BF16".into()),
|
||||
embedding_dim: Some(2048),
|
||||
embedding_base_url: Some("http://localhost:8000/v1".into()),
|
||||
..Config::default()
|
||||
@@ -3274,6 +3352,141 @@ mod tests {
|
||||
assert_eq!(embedder.document_prefix, "passage: ");
|
||||
}
|
||||
|
||||
/// Pure unit tests for `overlay_embedding_prefixes`: no process env or
|
||||
/// cwd mutation anywhere here (the workspace forbids `std::env::set_var`
|
||||
/// as `unsafe` under edition 2024, and `figment::Jail` calls
|
||||
/// `std::env::set_current_dir` on the real process internally, racing
|
||||
/// every other test in this multi-threaded lib test binary that
|
||||
/// relies on cwd, such as `tests/suite/backfill_e2e.rs`'s
|
||||
/// `Command::current_dir` calls). Each test builds its own minimal
|
||||
/// `Figment` in memory instead, exactly mirroring what `Config::load`
|
||||
/// does (`Serialized::defaults` as the base, optionally a lower-priority
|
||||
/// `Serialized` merge standing in for a `config.toml` value), and
|
||||
/// extracts a `Config` to assert on — the identical merge machinery the
|
||||
/// real loader uses, with the "env value" supplied as a parameter
|
||||
/// instead of read from the process.
|
||||
#[test]
|
||||
fn overlay_embedding_prefixes_preserves_trailing_whitespace() {
|
||||
// The regression this guards: figment's `Env` provider parses an
|
||||
// unquoted value with its loose-value parser, whose bare-value
|
||||
// branch calls `.trim()` — so without this overlay a real
|
||||
// `AI_MEMORY_EMBEDDING_QUERY_PREFIX="query: "` would arrive as
|
||||
// `"query:"`, silently dropping the space the model publisher
|
||||
// requires. `Serialized` bypasses that parser entirely.
|
||||
let base = Figment::from(Serialized::defaults(Config::default()));
|
||||
let overlaid = overlay_embedding_prefixes(base, Some("query: "), Some("passage: "));
|
||||
let cfg: Config = overlaid.extract().unwrap();
|
||||
assert_eq!(cfg.embedding_query_prefix.as_deref(), Some("query: "));
|
||||
assert_eq!(cfg.embedding_document_prefix.as_deref(), Some("passage: "));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn overlay_embedding_prefixes_none_leaves_a_lower_layer_untouched() {
|
||||
// Simulates a `config.toml` value already merged in at lower
|
||||
// priority; passing `None` (the env var genuinely absent) must not
|
||||
// disturb it.
|
||||
let base = Figment::from(Serialized::defaults(Config::default())).merge(
|
||||
Serialized::default("embedding_query_prefix", "toml-query: "),
|
||||
);
|
||||
let overlaid = overlay_embedding_prefixes(base, None, None);
|
||||
let cfg: Config = overlaid.extract().unwrap();
|
||||
assert_eq!(cfg.embedding_query_prefix.as_deref(), Some("toml-query: "));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn overlay_embedding_prefixes_some_wins_over_a_lower_layer() {
|
||||
let base = Figment::from(Serialized::defaults(Config::default())).merge(
|
||||
Serialized::default("embedding_query_prefix", "toml-query: "),
|
||||
);
|
||||
let overlaid = overlay_embedding_prefixes(base, Some("env-query: "), None);
|
||||
let cfg: Config = overlaid.extract().unwrap();
|
||||
assert_eq!(cfg.embedding_query_prefix.as_deref(), Some("env-query: "));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn overlay_embedding_prefixes_empty_string_clears_a_lower_layer() {
|
||||
// Present but empty (`Some("")`) is a deliberate override — an
|
||||
// operator clearing a `config.toml` value via env without editing
|
||||
// the file — distinct from `None` (the previous test), which must
|
||||
// leave the lower layer untouched.
|
||||
let base = Figment::from(Serialized::defaults(Config::default())).merge(
|
||||
Serialized::default("embedding_query_prefix", "toml-query: "),
|
||||
);
|
||||
let overlaid = overlay_embedding_prefixes(base, Some(""), None);
|
||||
let cfg: Config = overlaid.extract().unwrap();
|
||||
assert_eq!(cfg.embedding_query_prefix.as_deref(), Some(""));
|
||||
}
|
||||
|
||||
/// Exercises the real `Config::load` end to end, through an absolute
|
||||
/// `config.toml` path and data dir from a `TempDir`. TOML strings were
|
||||
/// never subject to figment's `Env`-provider trimming in the first
|
||||
/// place, so this path already worked before the fix; this guards it
|
||||
/// staying correct.
|
||||
///
|
||||
/// This process's own environment is shared with every other test in
|
||||
/// this binary and could already carry one of the two prefix vars from
|
||||
/// the test runner's shell, which would make `Config::load` pick the
|
||||
/// env value over the TOML one below and this test would silently stop
|
||||
/// verifying the TOML-only path. Rather than assume the ambient
|
||||
/// environment is clean, the actual `Config::load` call runs in a
|
||||
/// separate child process with both vars explicitly removed via
|
||||
/// `Command::env_remove` — real isolation instead of an in-process
|
||||
/// assumption, and it does not touch this rule's target (mutating
|
||||
/// *this* process's env), since a spawned child's environment is its
|
||||
/// own.
|
||||
#[test]
|
||||
fn loader_toml_path_preserves_whitespace_with_no_env_var_set() {
|
||||
const CHILD_MARKER: &str = "AI_MEMORY_TEST_LOADER_TOML_PATH_CHILD";
|
||||
if std::env::var_os(CHILD_MARKER).is_some() {
|
||||
// Running as the child, with both prefix vars removed by the
|
||||
// parent below: do the real work and print the result for the
|
||||
// parent to assert on.
|
||||
let tmp = TempDir::new().unwrap();
|
||||
let config_path = tmp.path().join("config.toml");
|
||||
std::fs::write(
|
||||
&config_path,
|
||||
"embedding_query_prefix = \"query: \"\n\
|
||||
embedding_document_prefix = \"passage: \"\n",
|
||||
)
|
||||
.unwrap();
|
||||
let cfg = Config::load(Some(&config_path), Some(tmp.path().to_path_buf())).unwrap();
|
||||
println!(
|
||||
"query={:?} document={:?}",
|
||||
cfg.embedding_query_prefix, cfg.embedding_document_prefix
|
||||
);
|
||||
return;
|
||||
}
|
||||
// Running as the parent: re-exec this same test binary, filtered
|
||||
// to just this one test, as a genuinely separate child process
|
||||
// with both prefix env vars removed.
|
||||
let exe = std::env::current_exe().expect("current test binary path");
|
||||
let output = std::process::Command::new(&exe)
|
||||
.arg("--exact")
|
||||
.arg("config::tests::loader_toml_path_preserves_whitespace_with_no_env_var_set")
|
||||
.arg("--test-threads=1")
|
||||
.arg("--nocapture")
|
||||
.env(CHILD_MARKER, "1")
|
||||
.env_remove("AI_MEMORY_EMBEDDING_QUERY_PREFIX")
|
||||
.env_remove("AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX")
|
||||
.output()
|
||||
.expect("failed to spawn child test process");
|
||||
assert!(
|
||||
output.status.success(),
|
||||
"child test process failed:\nstdout: {}\nstderr: {}",
|
||||
String::from_utf8_lossy(&output.stdout),
|
||||
String::from_utf8_lossy(&output.stderr)
|
||||
);
|
||||
let stdout = String::from_utf8_lossy(&output.stdout);
|
||||
assert!(
|
||||
stdout.contains("query=Some(\"query: \")"),
|
||||
"child stdout: {stdout}"
|
||||
);
|
||||
assert!(
|
||||
stdout.contains("document=Some(\"passage: \")"),
|
||||
"child stdout: {stdout}"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn copilot_embedding_defaults_model_dim_and_reuses_copilot_auth() {
|
||||
let tmp = TempDir::new().unwrap();
|
||||
|
||||
@@ -1014,7 +1014,16 @@ fn run_wrapper_with_fake_docker_env(
|
||||
.env("AI_MEMORY_DATA_VOLUME", "test-ai-memory-data")
|
||||
.env("HOME", shell_path(tmp.path()))
|
||||
.env_remove("AI_MEMORY_SERVER_URL")
|
||||
.env_remove("CLAUDE_CONFIG_DIR");
|
||||
.env_remove("CLAUDE_CONFIG_DIR")
|
||||
// Guarantees the "unset" case in
|
||||
// `posix_wrapper_forwards_embedding_prefixes_by_presence_not_non_emptiness`
|
||||
// is actually unset rather than silently inheriting whatever the
|
||||
// test-runner's own ambient environment happens to hold; every
|
||||
// other case re-adds one of these via `forwarded_env` below, so
|
||||
// removing them unconditionally here is a no-op for every other
|
||||
// caller of this helper.
|
||||
.env_remove("AI_MEMORY_EMBEDDING_QUERY_PREFIX")
|
||||
.env_remove("AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX");
|
||||
if let Some(claude_config_dir) = claude_config_dir {
|
||||
command.env("CLAUDE_CONFIG_DIR", claude_config_dir);
|
||||
}
|
||||
@@ -1833,6 +1842,83 @@ mod slow {
|
||||
}
|
||||
}
|
||||
|
||||
/// The two embedding-prefix env vars are forwarded on PRESENCE, not
|
||||
/// non-emptiness, unlike every other var in the loop: an operator sets
|
||||
/// one to the empty string to clear a `config.toml`-configured prefix
|
||||
/// without editing the file (see `Config::load`'s figment overlay in
|
||||
/// `ai-memory-cli/src/config.rs`), and that override only reaches the
|
||||
/// server if the wrapper forwards the (empty) variable rather than
|
||||
/// dropping it the way a plain `[ -n ]` check would.
|
||||
#[cfg(unix)]
|
||||
#[test]
|
||||
fn posix_wrapper_forwards_embedding_prefixes_by_presence_not_non_emptiness() {
|
||||
const QUERY: &str = "AI_MEMORY_EMBEDDING_QUERY_PREFIX";
|
||||
const DOC: &str = "AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX";
|
||||
let has_e = |args: &[&str], name: &str| args.windows(2).any(|pair| pair == ["-e", name]);
|
||||
|
||||
// Unset: neither var forwarded (the loop must not invent a value).
|
||||
let args =
|
||||
run_wrapper_with_fake_docker_and_forwarded_env(&["llm-test"], "[name=seccomp]", &[]);
|
||||
let lines: Vec<&str> = args.lines().collect();
|
||||
assert!(
|
||||
!has_e(&lines, QUERY),
|
||||
"unset must not be forwarded; got {lines:?}"
|
||||
);
|
||||
assert!(
|
||||
!has_e(&lines, DOC),
|
||||
"unset must not be forwarded; got {lines:?}"
|
||||
);
|
||||
|
||||
// Empty: forwarded anyway — this is the override case.
|
||||
let args = run_wrapper_with_fake_docker_and_forwarded_env(
|
||||
&["llm-test"],
|
||||
"[name=seccomp]",
|
||||
&[(QUERY, ""), (DOC, "")],
|
||||
);
|
||||
let lines: Vec<&str> = args.lines().collect();
|
||||
assert!(
|
||||
has_e(&lines, QUERY),
|
||||
"an empty (but present) value must still be forwarded; got {lines:?}"
|
||||
);
|
||||
assert!(
|
||||
has_e(&lines, DOC),
|
||||
"an empty (but present) value must still be forwarded; got {lines:?}"
|
||||
);
|
||||
|
||||
// Whitespace-only: also present, also forwarded — this loop must
|
||||
// not apply any trimming/emptiness judgement of its own.
|
||||
let args = run_wrapper_with_fake_docker_and_forwarded_env(
|
||||
&["llm-test"],
|
||||
"[name=seccomp]",
|
||||
&[(QUERY, " "), (DOC, " ")],
|
||||
);
|
||||
let lines: Vec<&str> = args.lines().collect();
|
||||
assert!(
|
||||
has_e(&lines, QUERY),
|
||||
"whitespace-only must still be forwarded; got {lines:?}"
|
||||
);
|
||||
assert!(
|
||||
has_e(&lines, DOC),
|
||||
"whitespace-only must still be forwarded; got {lines:?}"
|
||||
);
|
||||
|
||||
// Non-empty: forwarded, same as every other var.
|
||||
let args = run_wrapper_with_fake_docker_and_forwarded_env(
|
||||
&["llm-test"],
|
||||
"[name=seccomp]",
|
||||
&[(QUERY, "query: "), (DOC, "passage: ")],
|
||||
);
|
||||
let lines: Vec<&str> = args.lines().collect();
|
||||
assert!(
|
||||
has_e(&lines, QUERY),
|
||||
"a non-empty value must be forwarded; got {lines:?}"
|
||||
);
|
||||
assert!(
|
||||
has_e(&lines, DOC),
|
||||
"a non-empty value must be forwarded; got {lines:?}"
|
||||
);
|
||||
}
|
||||
|
||||
// The Windows mirror of macos_wrapper_routes_urls_by_real_subcommand: Docker
|
||||
// Desktop gives Linux containers no host networking on Windows either, so the
|
||||
// helper container cannot reach the host-published server over loopback.
|
||||
|
||||
@@ -337,13 +337,18 @@ pub struct OpenAiCompatEmbedder {
|
||||
dim: u32,
|
||||
/// Prepended to every query text before embedding (before truncation).
|
||||
/// Empty by default: symmetric models (most OpenAI-compatible servers)
|
||||
/// see no behaviour change. Asymmetric models such as Nemotron-3-Embed,
|
||||
/// the E5 family, and Qwen3-Embedding require a `"query: "` /
|
||||
/// `"passage: "` (or model-specific) instruction string that the
|
||||
/// OpenAI-compatible `/v1/embeddings` wire format has no field for —
|
||||
/// the client has to prepend it instead. See `with_prefixes`.
|
||||
/// see no behaviour change. Asymmetric models need a query-side
|
||||
/// instruction the OpenAI-compatible `/v1/embeddings` wire format has
|
||||
/// no field for — the client has to prepend it instead.
|
||||
/// `nvidia/Nemotron-3-Embed-1B-BF16` and base E5 models use a simple
|
||||
/// `"query: "` / `"passage: "` pair; `e5-mistral-7b-instruct` and
|
||||
/// Qwen3-Embedding instead need a full task-instruction string with
|
||||
/// different exact spacing each (documents plain for both — see
|
||||
/// `ai-memory-cli/src/config.rs`'s `embedding_query_prefix` field doc
|
||||
/// comment for the two exact templates). See `with_prefixes`.
|
||||
query_prefix: String,
|
||||
/// Document-side counterpart of `query_prefix` (e.g. `"passage: "`).
|
||||
/// Document-side counterpart of `query_prefix` (e.g. `"passage: "` for
|
||||
/// Nemotron-3-Embed / base E5 — not every model needs one).
|
||||
document_prefix: String,
|
||||
}
|
||||
|
||||
@@ -648,6 +653,41 @@ pub fn cosine(a: &[f32], b: &[f32]) -> f32 {
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
/// Transport-level proof that `OpenAiEmbedder` (not just
|
||||
/// `OpenAiCompatEmbedder`, covered in
|
||||
/// `tests/suite/openai_compat_embedder.rs`) actually sends the
|
||||
/// configured prefix on the wire — `with_prefixes` alone only proves
|
||||
/// the fields are stored, not that `embed_query`/`embed_document` use
|
||||
/// them in the real HTTP request body.
|
||||
#[tokio::test]
|
||||
async fn openai_embedder_sends_the_prefix_on_the_wire() {
|
||||
use wiremock::matchers::{method, path};
|
||||
use wiremock::{Mock, MockServer, Request, ResponseTemplate};
|
||||
|
||||
let server = MockServer::start().await;
|
||||
Mock::given(method("POST"))
|
||||
.and(path("/v1/embeddings"))
|
||||
.respond_with(move |req: &Request| {
|
||||
let body: serde_json::Value = serde_json::from_slice(&req.body).unwrap();
|
||||
assert_eq!(body["input"], "query: find the runbook");
|
||||
ResponseTemplate::new(200).set_body_json(serde_json::json!({
|
||||
"data": [{ "embedding": vec![0.5_f32; 4] }],
|
||||
}))
|
||||
})
|
||||
.expect(1)
|
||||
.mount(&server)
|
||||
.await;
|
||||
|
||||
let e = OpenAiEmbedder::new(SecretString::from("sk-test"), "text-embedding-3-small", 4)
|
||||
.unwrap()
|
||||
.with_base_url(server.uri())
|
||||
.with_prefixes("query: ", "passage: ");
|
||||
|
||||
e.embed_query("find the runbook")
|
||||
.await
|
||||
.expect("embed_query succeeds");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn synthetic_embedder_produces_unit_vectors() {
|
||||
let e = SyntheticEmbedder::new(64);
|
||||
@@ -761,7 +801,7 @@ mod tests {
|
||||
let e = OpenAiCompatEmbedder::new(
|
||||
"http://localhost:9/v1",
|
||||
None,
|
||||
"nvidia/nemotron-3-embed",
|
||||
"nvidia/Nemotron-3-Embed-1B-BF16",
|
||||
2048,
|
||||
)
|
||||
.expect("embedder builds")
|
||||
|
||||
@@ -211,11 +211,17 @@ pub struct EmbedderConfig {
|
||||
/// `openai` and `openai-compat` embedders apply it; other providers
|
||||
/// (`google` has its own built-in task-type asymmetry, `voyage`,
|
||||
/// `local`, `copilot`) ignore it. Needed for asymmetric self-hosted
|
||||
/// models — Nemotron-3-Embed, the E5 family, Qwen3-Embedding — whose
|
||||
/// publisher specifies a `"query: "` instruction the OpenAI-compatible
|
||||
/// `/v1/embeddings` wire format has no field for.
|
||||
/// models whose publisher specifies a query-side instruction the
|
||||
/// OpenAI-compatible `/v1/embeddings` wire format has no field for.
|
||||
/// `nvidia/Nemotron-3-Embed-1B-BF16` and base E5 models
|
||||
/// (`intfloat/e5-base-v2`, multilingual E5, …) use a simple
|
||||
/// `"query: "` string; instruction-tuned E5 variants and
|
||||
/// Qwen3-Embedding instead need a full task-instruction string (their
|
||||
/// documents stay plain — leave `document_prefix` unset for those).
|
||||
pub query_prefix: String,
|
||||
/// Document-side counterpart of `query_prefix` (e.g. `"passage: "`).
|
||||
/// Document-side counterpart of `query_prefix` (e.g. `"passage: "` for
|
||||
/// Nemotron-3-Embed / base E5 — not every model needs one; see
|
||||
/// `query_prefix`'s doc comment).
|
||||
pub document_prefix: String,
|
||||
}
|
||||
|
||||
|
||||
@@ -5734,6 +5734,71 @@ mod tests {
|
||||
(tmp, store, server, ws, proj)
|
||||
}
|
||||
|
||||
/// A test-only embedder whose three `Embedder` methods each return a
|
||||
/// different, identifiable vector — mirroring `ai-memory-llm`'s own
|
||||
/// `health::tests::TaskAwareEmbedder` fixture (same idea, duplicated
|
||||
/// here because that one is private to its crate). `embed` deliberately
|
||||
/// returns the SAME vector as `embed_document`: that is what the
|
||||
/// pre-fix bug actually called (`AiMemoryServer::embed_query` invoked
|
||||
/// the generic `Embedder::embed` instead of `Embedder::embed_query`),
|
||||
/// so a regression back to that bug is what this fixture would surface.
|
||||
struct TaskAwareEmbedder;
|
||||
|
||||
#[async_trait::async_trait]
|
||||
impl Embedder for TaskAwareEmbedder {
|
||||
fn provider(&self) -> &'static str {
|
||||
"task-aware-test"
|
||||
}
|
||||
|
||||
fn model(&self) -> &str {
|
||||
"task-aware-test-model"
|
||||
}
|
||||
|
||||
fn dim(&self) -> u32 {
|
||||
2
|
||||
}
|
||||
|
||||
async fn embed(&self, _text: &str) -> ai_memory_llm::LlmResult<Vec<f32>> {
|
||||
Ok(vec![1.0, 0.0])
|
||||
}
|
||||
|
||||
async fn embed_document(&self, _text: &str) -> ai_memory_llm::LlmResult<Vec<f32>> {
|
||||
Ok(vec![1.0, 0.0])
|
||||
}
|
||||
|
||||
async fn embed_query(&self, _text: &str) -> ai_memory_llm::LlmResult<Vec<f32>> {
|
||||
Ok(vec![0.0, 1.0])
|
||||
}
|
||||
}
|
||||
|
||||
/// Regression test for the bug fixed alongside the embedding
|
||||
/// query/document prefix feature: `AiMemoryServer::embed_query` (the
|
||||
/// helper `memory_query` calls to vectorize the search query) called
|
||||
/// `Embedder::embed` instead of `Embedder::embed_query`. For any
|
||||
/// query/document-asymmetric embedder — Google's task-typed
|
||||
/// `embedContent`, or the new `openai`/`openai-compat` query/document
|
||||
/// prefixes — that silently embedded the search query as if it were a
|
||||
/// document, corrupting vector-stream retrieval. `TaskAwareEmbedder`
|
||||
/// makes the two paths return distinguishable vectors so the right one
|
||||
/// being called is a direct, exact-equality assertion rather than an
|
||||
/// inference from downstream ranking.
|
||||
#[tokio::test]
|
||||
async fn memory_query_embeds_the_query_with_embed_query_not_embed() {
|
||||
let (_tmp, _store, server, _ws, _pj) = setup_server().await;
|
||||
let server = server.with_embedder(Arc::new(TaskAwareEmbedder));
|
||||
|
||||
let query_vec = server.embed_query("anything").await;
|
||||
|
||||
assert_eq!(
|
||||
query_vec.as_deref(),
|
||||
Some([0.0_f32, 1.0].as_slice()),
|
||||
"memory_query's embed_query helper must call Embedder::embed_query \
|
||||
(query-side vector [0.0, 1.0]), not the generic Embedder::embed \
|
||||
(which this fixture deliberately aliases to the document-side \
|
||||
vector [1.0, 0.0] to make a regression to the old bug fail loudly)"
|
||||
);
|
||||
}
|
||||
|
||||
fn installed_ai_memory_prompt_surface() -> String {
|
||||
let mut prompt = String::from(ai_memory_core::SNIPPET_BODY);
|
||||
for skill in ai_memory_core::routing_skills::MANAGED_SKILLS {
|
||||
|
||||
+1
-1
@@ -1641,7 +1641,7 @@ If you set only the provider, ai-memory picks a sensible default:
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=openai` + `AI_MEMORY_EMBEDDING_BASE_URL=https://api.orcarouter.ai/v1` | `openai/text-embedding-3-small` via [OrcaRouter](https://www.orcarouter.ai) | Uses `EMBEDDING_API_KEY`, else reuses `LLM_API_KEY`, with the OpenAI-compatible embedding client. |
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=voyage` | `voyage-3` (1024-dim) | Voyage's current general-purpose recommendation. |
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=google` / `gemini` | `gemini-embedding-001` (768-dim) | Google-hosted embeddings via `embedContent`. Set `GEMINI_API_KEY` (or `GOOGLE_API_KEY`). |
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=openai-compat` | no default — set model, dim, and base URL explicitly | Self-hosted engines (Ollama, LM Studio, vLLM). Keyless by default; `EMBEDDING_API_KEY`, else `LLM_API_KEY`, is sent as a bearer token when present (gateways). Example: `AI_MEMORY_EMBEDDING_BASE_URL=http://localhost:11434/v1`, `AI_MEMORY_EMBEDDING_MODEL=nomic-embed-text`, `AI_MEMORY_EMBEDDING_DIM=768`. Switching an existing `openai`+base-URL setup to `openai-compat` changes the stored `{provider, model, dim}` triple — run `ai-memory embed --force` to re-embed. Asymmetric models need `AI_MEMORY_EMBEDDING_QUERY_PREFIX` / `AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX`, e.g. `nvidia/nemotron-3-embed` wants `query: ` / `passage: ` — see [`docs/llm-providers.md`](llm-providers.md). |
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=openai-compat` | no default — set model, dim, and base URL explicitly | Self-hosted engines (Ollama, LM Studio, vLLM). Keyless by default; `EMBEDDING_API_KEY`, else `LLM_API_KEY`, is sent as a bearer token when present (gateways). Example: `AI_MEMORY_EMBEDDING_BASE_URL=http://localhost:11434/v1`, `AI_MEMORY_EMBEDDING_MODEL=nomic-embed-text`, `AI_MEMORY_EMBEDDING_DIM=768`. Switching an existing `openai`+base-URL setup to `openai-compat` changes the stored `{provider, model, dim}` triple — run `ai-memory embed --force` to re-embed. Asymmetric models need `AI_MEMORY_EMBEDDING_QUERY_PREFIX` / `AI_MEMORY_EMBEDDING_DOCUMENT_PREFIX`, e.g. `nvidia/Nemotron-3-Embed-1B-BF16` wants `query: ` / `passage: ` — see [`docs/llm-providers.md`](llm-providers.md) for that and for Qwen3-Embedding/instruction-tuned E5, which need a different (query-only) format. |
|
||||
| `AI_MEMORY_EMBEDDING_PROVIDER=copilot` | `text-embedding-3-small` (1536-dim) | Reuses the `copilot` LLM provider's OAuth login (`ai-memory auth login copilot`, `COPILOT_GITHUB_TOKEN`, or `GITHUB_COPILOT_API_TOKEN`) — no separate API key. Calls Copilot's `/embeddings` endpoint following the OpenAI-compatible contract Copilot documents for chat; that endpoint's exact shape is not covered by a live test against Copilot here, so treat it as needing a real-Copilot smoke test. |
|
||||
|
||||
> **What we don't recommend:** reasoning-mode models (Claude with extended
|
||||
|
||||
+24
-5
@@ -220,20 +220,39 @@ respectively, before the existing truncation, so a long body is still
|
||||
truncated to the same overall input cap with the prefix included. Both are
|
||||
empty by default — no behaviour change when unset — and are not trimmed, so
|
||||
a publisher's trailing space is preserved exactly. For example, NVIDIA's
|
||||
`nemotron-3-embed` (served locally through vLLM/`openai-compat`) specifies
|
||||
`"query: "` and `"passage: "`:
|
||||
[`nvidia/Nemotron-3-Embed-1B-BF16`](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16)
|
||||
(2048-dim, served locally through vLLM/`openai-compat`) specifies
|
||||
`"query: "` for queries and `"passage: "` for documents per its model card:
|
||||
|
||||
```toml
|
||||
embedding_provider = "openai-compat"
|
||||
embedding_model = "nvidia/nemotron-3-embed"
|
||||
embedding_model = "nvidia/Nemotron-3-Embed-1B-BF16"
|
||||
embedding_base_url = "http://localhost:8000/v1"
|
||||
embedding_dim = 2048
|
||||
embedding_query_prefix = "query: "
|
||||
embedding_document_prefix = "passage: "
|
||||
```
|
||||
|
||||
The E5 family and Qwen3-Embedding use the same `"query: "` / `"passage: "`
|
||||
convention. Changing either prefix does not change the stored
|
||||
Base E5 models (`intfloat/e5-base-v2`, `e5-large-v2`, multilingual E5, …)
|
||||
use the same `"query: "` / `"passage: "` convention. Instruction-tuned E5
|
||||
variants and Qwen3-Embedding instead need a task-instruction string on the
|
||||
**query side only** — their documents are embedded plain, with no document
|
||||
prefix — but the two use **different exact spacing**, confirmed against
|
||||
each model card:
|
||||
|
||||
- `intfloat/e5-mistral-7b-instruct`:
|
||||
`embedding_query_prefix = "Instruct: {task description}\nQuery: "` — a
|
||||
trailing space after `Query:`.
|
||||
- `Qwen/Qwen3-Embedding-0.6B` (and the other Qwen3-Embedding sizes):
|
||||
`embedding_query_prefix = "Instruct: {task description}\nQuery:"` — **no**
|
||||
trailing space; the query text follows the colon directly.
|
||||
|
||||
Fill in your own task description for `{task description}`, leave
|
||||
`embedding_document_prefix` unset for both, and don't copy one model's
|
||||
exact string for the other — the trailing-space difference is
|
||||
publisher-specified, not a typo.
|
||||
|
||||
Changing either prefix does not change the stored
|
||||
`{provider, model, dim}` triple pages are keyed by, so existing vectors keep
|
||||
matching on the mismatch check but were embedded under the old (or no)
|
||||
prefix; run `ai-memory embed --force` to re-embed after changing one.
|
||||
|
||||
Reference in New Issue
Block a user