The parallel-scheduler contract failed twice on test-windows CLANG64 while
running fake suites only:
- PR #2394 (run 36364479687): "Windows timeout race refused without a
surviving descendant to refuse over". The refusal came 38s after the
forced leader exit (a cold powershell/CIM descendant probe, once in the
wave loop and again in the cleanup pass), while the fixture's descendant
was a `time.sleep(30)` that had already exited on its own.
- PR #2345 (run 36199749518): "suite 'hang_after_summary' taskkill could
not prove process-tree cleanup". The scheduler bounded taskkill.exe with
--kill-grace, which the contract passes as 1s, so a slow cold start of
the tool was reported as a failed cleanup.
Both verdicts were a race between a clock and runner latency. An audit of
the fixture found more of the same shape:
- Every hanging fixture ended on its own after 30s. A starved scheduler
could therefore see a "hung" leader exit 0, and a descendant the
scheduler failed to kill could vanish before the leak check looked.
- The 1s --timeout could kill a slow leader before it had printed its
summary (asserted pass=1), spawned its descendant or written the
descendant pid. That pid was then read back with int() from a file
written create-then-write, so a reader could see no file or an empty one.
The empty-file read recorded earlier was the scheduler's ready file,
which 3b9052c9 already publishes atomically; the fixture's own writer
had the same shape.
- The descendant installed its SIGTERM-ignore during its own start-up. A
SIGTERM that won that race skipped the SIGKILL escalation the POSIX leg
exists to cover, silently and with the verdict unchanged.
- On Windows, the stubborn-tree leak check was a single look right after
taskkill, but TerminateProcess returns before its target has exited.
Fix it by construction, not by budget:
- A gate. Every process that must outlive the scheduler's decision blocks
reading stdin. That is a pipe whose only write end the contract holds,
handed down through the scheduler, and the read returns only at EOF. The
contract closes the pipe once the verdict is recorded. If the contract
dies, EOF releases everything, so nothing is orphaned and no fixture
ends on a clock.
- The three hanging scenarios run under the scheduler's existing
pre-terminate barrier. They are released only once the fixture has
published `<suite>.established`, which happens after the summary is
flushed, the descendant has confirmed it is armed, and its pid is known.
The marker is written to a temp file and moved into place with
os.replace, so a reader sees either no file or the whole file.
- The Windows leak check waits on the descendant's process handle. That
wait is exact, not a race: the descendant is held by the gate, so the
scheduler's kill is the only way it can end.
- run-test-wave.py runs taskkill through windows_taskkill_tree() with its
own stable-state budget (WINDOWS_TASKKILL_SECONDS = 15), as 132e8fc3 did
for the descendant probe. --kill-grace keeps bounding what its name
says. Real runs already pass --kill-grace 15, so they are unchanged.
The contract pins this structurally, the same way it pins the probe:
exactly one taskkill site, no kill-grace there, and a helper bounded by
the constant that fails closed on a timeout, a start failure or a
non-zero exit.
Every assertion keeps its contract:
- A leader that hangs after a green summary is recorded as rc=124 pass=1.
- A zero-test child is recorded as rc=97.
- The wave continues after a bounded failure.
- A SIGTERM-resistant tree is killed in full.
- A leader that exits between the timeout decision and termination is
still cleaned up on POSIX. On Windows the scheduler refuses (rc=2,
naming the cleanup failure) over a descendant that is really alive.
A new self-check fails loudly if a held leader or descendant is not live
while the barrier holds, that is, if the fixture no longer exercises what
those assertions claim.
Proof, on macOS arm64 and a Linux arm64 container. The Windows branch was
reasoned through; the VM is down.
- Torture hook (proof only) that delays each hanging leader 2s, past the
1s timeout. The old fixture is RED 3/3 on every scenario (pass=0,
FileNotFoundError on descendant.pid, "lost its bounded result"). The new
one is green, including at 8s, and 10/10 under the hook plus CPU load.
- The #2394 mechanism: after an injected 31s refusal latency the old
descendant is gone and the new one is live. It exits 0.02s after the
gate closes.
- The taskkill pin is RED on the old scheduler and green on the new one.
- Stress, 20 runs each under CPU load. The old full contract passed 20/20
on macOS and 20/20 on Linux, so these hosts do not reproduce the CI
timing. The rebuilt held scenarios passed 20/20 on macOS and 20/20 on
Linux. The full contract after the change passes a single run on macOS
and on Linux, and 20/20 on each leg under CPU load (4 load workers;
the Linux container is capped at 4 CPUs).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
security-static (scripts/security-vendored.sh) failed on this branch:
the #2176 fix patches internal/cbm/vendored/grammars/rescript/scanner.c
(skip the template-string branch in the error-recovery state) and records
that patch in vendored/grammars/MANIFEST.md, but the checked-in digests in
scripts/vendored-checksums.txt still described the unpatched files.
Both changes are intentional and reviewed in the parent commit; this
records their SHA-256 via `scripts/security-vendored.sh --update`. Only
the two affected lines change, and the integrity check passes again.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Upgrading from v0.10.8 to v0.11.0 on a machine with OpenCode reported all
three OpenCode subagent profiles as "preserved modified profile" plus an
op=agent_install error each, and the activation stopped ("one or more agent
configurations failed; the published/current executable was kept"), although
nobody had edited those files.
Root cause: #1933 (v0.11.0) added the tool_search / tool_search_regex
permissions to the rendered OpenCode profile. Install and uninstall recognise
a profile as ours only when it matches the current rendering or one of the
known released renderings (the other access mode, the v0.9.1-rc.1 Codex
shape, the pre-tier Verify file). The v0.10.8 OpenCode shape was not in that
set, so cbm_text_migrate_owned_document classified the untouched v0.10.8
files as user-modified. Reproduced end to end: install with the v0.10.8
release binary, then install with v0.11.0/main, into an isolated HOME; the
three profiles are 1729/1880/1910 bytes, the same sizes as in the report.
cbm_render_graph_profile_opencode_v0108() renders the pre-#1933 shape
(both access modes), and install/uninstall now list it as a released
rendering, so those files are upgraded in place and uninstall removes them
as owned. A profile the user actually edited is still preserved. The
released-list assembly that install and uninstall duplicated moves into one
helper, and the two legacy renderers share one implementation (cli.c drops
three raw free sites; the memory-core baseline is tightened 173 -> 170).
The Codex and OpenCode op=mcp_install errors in the same report do not
reproduce from a pristine v0.10.8 config and depend on the reporter's own
config.toml / opencode.json; they are not addressed here.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
walk_defs keeps its pending frames in a growable stack that was bounded by
an 8M-frame ceiling (env CBM_WALK_DEFS_MAX). At the ceiling wd_push logged
"extract.walk_defs_capped" once and then stopped pushing. Children are
pushed last-to-first, so a file wider than the ceiling lost its FIRST
definitions: a work cap decided graph content.
Fix: no ceiling. The stack doubles on demand, heap frames now come from the
memory core (cbm_realloc/cbm_free, class EXTRACT) instead of safe_realloc and
raw free, and the doubling is overflow-guarded (int capacity and size_t
byte count). The only thing that can stop the walk is an allocation failure,
and that is loud: logged at error level as "extract.walk_stack_alloc_failed"
(walker=walk_defs, pending, path) -- the same event the complexity walkers
use -- and the file's result is marked has_error with
"definitions walk: stack allocation failed", which the definitions passes
already report as a per-file extract error. The arena-backed path
(ctx->scratch) gets the same failure handling. Traversal order is unchanged.
CBM_WALK_DEFS_MAX is removed. It was read only in wd_stack_max(); it was not
documented anywhere (docs, --help, config, tests), so nothing else changes.
extract_defs.c loses its raw free() call, so the memory-core baseline for
the file drops from 5 to 4.
Test (extraction suite): walk_defs_wide_file_extracts_every_definition
extracts a C file with 1000 top-level functions with CBM_WALK_DEFS_MAX=256
set, and asserts every wd0..wd999 definition is present. RED on main
(extract.walk_defs_capped limit=256; wd0 missing), GREEN with the fix, RED
again with the production change reverted. The extraction (379), complexity
(5), ast_profile (10) and semantic (35) suites pass.
Real input: Linux kernel, index-only, fresh cache and runtime dirs, same
machine. main and fix both produce 8,530,336 nodes and 15,764,128 edges,
SEMANTICALLY_RELATED 607 each, and the full node set (label, QN, lines) and
edge set (type, source QN, target QN) hash identically -- no kernel file
approaches the old ceiling. Index wall time 210.0 s (main) vs 194.4 s (fix),
within run-to-run noise.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
pc_module_var added a third compute-QN / find_by_qn / free sequence to
pass_calls.c, growing its memory-core ratchet (23 -> 24). Route all three
through pc_find_by_computed_qn, which leaves 22 raw sites, and ratchet the
pass_calls.c baseline down to match.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The lint chain stops at its first red target, so a clang-format failure hid the memory-core linter
until PR CI: the coverage line-offset table, the Cypher direction swap and the search label each
added a raw allocator site. The table now comes from cbm_alloc/cbm_free (extract class), the swap
frees through safe_str_free, and the BM25 row array moves to cbm_calloc/cbm_free so mcp.c's raw-site
count goes down (baseline 814 -> 813) rather than up.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Review on #2199: cbm_test_runtime_init ran more than a hundred lines
before the only EXIT trap, so under `set -e` a failure in the fixture
mktemp or its cygpath conversion left the private root behind, and the
fixture trap ran smoke_rmtree ahead of the runtime cleanup.
Arm `trap 'cbm_test_runtime_cleanup "$BINARY"' EXIT` immediately after
a successful init, as soak-test.sh and memlab.sh do. The fixture trap
that replaces it now retires the private daemon first and removes the
fixtures second, so no earlier cleanup step stands between the exit and
the runtime cleanup; on Windows the retirement is also what unblocks
the fixture rm.
tests/test_smoke_runtime_isolation_contract.sh pins the early trap to
the window between the init call and the first fixture; the pin fails
against the previous head.
Part of #1696.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Anton Standrik <astandrik@yandex-team.ru>
scripts/smoke-test.sh retires the account daemon from seven call sites,
but its wrappers sandbox only HOME/TMPDIR/CBM_CACHE_DIR and only
CBM_RUNTIME_DIR moves the daemon rendezvous, so every retirement reached
the operator's live daemon. Source scripts/test-runtime.sh so every
product process runs under a harness-owned runtime and cache, and clean
that root up from the EXIT trap.
Add tests/test_smoke_runtime_isolation_contract.sh, which fails before
this change, and wire it into scripts/test.sh.
Part of #1696.
Signed-off-by: Anton Standrik <astandrik@yandex-team.ru>
main fails the memory-core linter:
src/mcp/mcp.c: grew by 3 (814 -> 817)
PR #1723 (merged as 92abefa3) added a new early-return path to
handle_index_repository and, following the function's convention, freed
its three argument strings there with raw free(). Its CI lint was green
on 2026-09-09. On 2026-09-19 the memwaste merge cut this file's raw sites
and ratcheted the baseline down to 814, so by the time #1723 landed its
green was measured against a rule that no longer existed. The linter is
right and the merge-result check that let it through was incomplete: it
built and ran the suites and did not run the linter. It does now.
The fix is not cbm_free. The three strings come from
cbm_mcp_get_string_arg and resolved_repo_path_from_project_arg and are
raw heap memory, released raw by every other path in the function;
routing three of them through the core would free raw memory through the
accounting layer. The actual defect is that the function had 41 raw
free() calls of the same three arguments across a dozen early returns,
and #1723 added one more block in the same shape.
One helper, index_args_free(repo_path, mode_str, name_override), now
replaces twelve of those blocks. free(NULL) is a no-op, so the one path
where repo_path is still NULL passes NULL. On the cross-repo branch
mode_str used to be freed before handle_cross_repo_mode and the other
two after it; mode_str is neither passed to nor read by that call, so
releasing all three after it is not observable. Five sites on the
normal-completion path are left as they were: they free the strings at
different points as ownership is handed to the pipeline, and collapsing
them would change that order for no gain.
Raw sites in this file: 817 -> 785. The baseline follows the improvement,
814 -> 785, as the linter asks; it can only tighten.
Build: clean, -Werror, ASan/UBSan. mcp cli daemon_application:
692 passed, 4 skipped, 0 failed -- index_repository_honors_allowed_root,
which runs through a replaced early return, included.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
PR CI's lint job failed on src/foundation/mem_override_libc.c:143-144 while the
local lint lane passed on the same file. The two run different formatters: CI
installs LLVM 20, the local lane uses the Homebrew build, which is now 22. A
wrapped attribute is one of the few things they lay out differently, so the
wrapped form was stable locally and a violation in CI.
__attribute__((no_builtin("memcpy", "memmove",
"memset"))) static void *slow_move(...)
Written on one line the declaration is 117 columns, over the 100-column limit,
so it has to wrap and the two versions disagree about how. Spelling the
attribute as a macro brings the declaration to 73 columns, which leaves no
wrapping decision for either version to have an opinion about. The three
siblings below it already fit inline and are left alone.
Also repins scripts/vendored-checksums.txt for MANIFEST.md itself: the security
job's vendored-integrity layer checksums the manifest too, so editing it to
record the properties scanner patch invalidated its own entry. Produced by
scripts/security-vendored.sh --update; the one changed line is the manifest
hash, which is also the check that the earlier manual scanner.c repin was right.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Indexing a repo with .properties files in it produced a different graph from
one run to the next. Every few runs one small file - 253 bytes, valid, the same
file that had indexed fine a moment earlier - vanished from the graph, and the
coverage tree put raw-path File/Folder nodes where it had been. On the kotlin
corpus that showed up as node counts wandering between 33,486 and 33,494.
The vendored tree-sitter properties grammar keeps its external scanner's entire
state in a file-scope static:
static bool reached_eof = false;
bool ..._external_scanner_scan(void *payload, TSLexer *lexer, const bool *valid) {
lexer->result_symbol = FAKE_EOL;
return reached_eof = !reached_eof && valid[FAKE_EOL] && lexer->eof(lexer);
}
That flag is how the scanner refuses to emit a second end-of-input FAKE_EOL,
and the grammar needs exactly one to close a file. One static is one object for
the whole process, so with several index workers parsing .properties files at
once, whichever thread reached EOF first set the flag and the next thread's
`!reached_eof` was false: its parse never received the FAKE_EOL it was waiting
for and sat at end-of-input asking for a token that would never arrive. The
per-file parse budget stopped it five seconds later and the file was dropped.
Measured on a 253-byte file: 1,060,630 progress callbacks, ~106 M parse
operations, byte offset pinned at EOF, has_error false - the parser was not
recovering from an error, it was waiting.
The state now lives in the payload from ..._external_scanner_create(), which is
what tree-sitter's create/destroy contract is for. The serialize/deserialize
wire format is unchanged: the state is still the returned length.
Proof, interleaved under identical load so both binaries meet the same machine,
40 runs each: unpatched 6 stalls, patched 0. Under load the kotlin corpus now
returns 33494/203579 ten times out of ten and java 693129/5523595 three times
out of three; both used to wander.
All 103 vendored scanners were audited for the same shape. This was the only
one that WRITES a process-wide static. haskell's debug_parse_metadata and
purescript's res_cont/res_fail are read-only - constants missing a const - and
are left alone.
Upstream 6310671b24d4 still carries the static, so MANIFEST.md records this in
the local-patch table and scripts/vendored-checksums.txt is repinned.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
On 2026-09-18 a Windows leg printed "7925 passed, 0 failed" and
"=== All tests passed ===" and still exited 1, while the identical work run
through win.sh's other entry exited 0. The status travels
ssh -> cmd.exe -> msys2_shell.cmd and that chain loses it.
A channel that can turn 0 into 1 can turn 1 into 0, and that direction is a
false green: a red Windows leg reported as passing, which is the one thing the
local ladder exists to prevent. So the two channels are now split by what each
can actually know. The log decides the test outcome — it is written by the
runner on the VM and cannot be mangled in transit — and rc only decides what
the log cannot see. Green needs zero reported failures AND the completion
marker scripts/test.sh prints last, so a leg that stopped early has a summary
but no marker and is red. A lost exit status is reported, not obeyed.
The logic is a function, vm_verdict(), reachable as `--verdict <log> <rc>`, so
tests/test_vm_verdict_contract.sh drives the REAL code with synthetic logs
instead of a copy that could drift from it: lost exit status, failures under
rc=0, a marker that must not launder failures, failures in a later suite,
partial runs, and no summary at all (which keeps its own exit code 90, so
"never ran" stays distinguishable from "ran and failed"). Eight cases, wired
into scripts/test.sh as step 0y.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The memory-core linter is a ratchet: raw malloc/calloc/realloc/free/strdup
sites may not grow, because allocations belong on one accounted path
(src/foundation/mem_core.h) rather than spread across hundreds of call
sites. This branch's own work had added 35 raw sites across nine files --
the parallel k8s prep, the configlink key matching, the worker pool's
counted path, the Go shared stdlib registry, the spill namespace copy,
the route service array and the namespace-name arrays.
All of them now use cbm_alloc/cbm_calloc/cbm_realloc/cbm_free/
cbm_mem_strdup. One was a latent bug rather than a style point:
pass_k8s freed a buffer that k8s_read_file had allocated with raw malloc,
so converting only the free would have mismatched the allocator -- the
read function and every free of its result are converted together.
The Windows free-space path widens its argument through the memory core
instead of cbm_utf8_to_wide(), which allocates with raw malloc.
Two files improved enough to tighten their baselines, which the ratchet
requires so the gain cannot be quietly given back: sqlite_writer.c 92->74
(the index-cell arena), pass_k8s.c 13->6 and registry.c 19->17.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Indexing broke on x86-64 Linux: the index worker refused to start and
SIGKILLed its own process group, so every venue reported nothing but
"index worker ended with killed (exit=-1, signal=9)".
glibc places a thread's static TLS block inside the stack allocation it
makes for that thread. This branch's new thread-local caches -- a
512-slot field cache (~24 KB), per-depth tree cursors (~4 KB) and two
parked CBMArenas (~8 KB; an arena carries blocks[256] + block_sizes[256])
-- took static TLS from 47,905 to 93,513 bytes. The 64 KB parent-death
watchdog thread then could not be created at all, and a worker without
process-tree containment correctly refuses to index.
Four layers, because there were four things wrong:
- The big thread-locals move to the heap behind a thread-local POINTER,
allocated once per thread and reused, each with a correct fallback if
allocation fails (the field cache and cursor pool are caches; a miss
is slower, never wrong). Static TLS is back to 52,233 bytes.
- PARENT_WATCHDOG_STACK_SIZE 64 KB -> 256 KB. "The watchdog only polls,
a tiny stack suffices" was true when written and false once TLS
reached ~50 KB, because the TLS block comes out of that same
allocation.
- cbm_thread_create retries with the default stack on EINVAL. A stack
size is a hint about what a thread needs; the platform refusing the
hint is not a reason to fail to create the thread.
- tests/test_thread_stack_tls_contract.sh (Step 5a of scripts/test.sh)
keeps static TLS a quarter clear of the smallest requested stack,
reading that smallest size from the sources so a new small stack
tightens the gate automatically.
Two diagnostics that made this findable at all, and that stay:
smoke-test.sh prints the worker log its failure message names (CI deletes
that sandbox with the job, so the one artifact naming the cause was
unreadable), and worker_containment_unavailable() now reports WHICH of
its conditions failed, with pid/pgrp/ppid.
Verified in the amd64 container, the venue nothing on the local ladder
covers: baseline main indexes 9039 nodes/15067 edges, this branch before
the fix refuses to start the worker, and after it indexes 9039/15067.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
scripts/fuzz.sh drives the libFuzzer targets in two shapes from one entry:
a fixed --runs from a fixed --seed, so a CI smoke is a function of the code
and not a lottery (O9), and a time-boxed --max-total-time for exploration.
Checked-in seeds are copied into a working corpus under build/fuzz/, so a
run never rewrites the repository; a crash leaves its reproducer behind and
the script exits non-zero. Wired as a smoke into dry-run.yml and as
exploration into nightly-soak.yml.
The venue-parity contract rightly refused both workflow steps: a venue may
provision or plumb, but only a CANONICAL leg entry may exercise the
product, and neither script was registered. Registering them — rather than
working around the contract — is the fix, since both own their leg end to
end (build, corpus or binary, artifact paths, exit code) exactly as
soak-legs.sh owns the soak. _memwaste.yml joins the policed venue walk for
the same reason, and both scripts join the layer-5 interface probes.
Those probes immediately earned their keep, finding two real defects:
`fuzz.sh --help` exited 2 because the positional target consumed the flag,
and memwaste.sh read a typo'd flag as the repository path — which would
have indexed nothing and reported an empty lane as a clean one.
test_shell_line_endings.sh now degrades where the checkout has no
reachable repository metadata: the Linux container bind-mounts the working
tree, and a git worktree's .git is a pointer to the host repository, so
`git ls-files` cannot run there and the contract failed with "matched no
files — the glob set is broken", killing the leg before it compiled. It is
a repository property, so it stays gating in every venue that has the
metadata (macOS host, the Windows VM's real checkout at C:\cbm with 117
files, every hosted-CI checkout). test_version_metadata_contract.sh already
guards the same way for the same reason; the alternatives considered and
rejected are recorded at the case.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Answers a question the existing tooling could not: of everything we
allocate, how much is never used?
Lanes, all driven by scripts/memwaste.sh:
event every allocation observed through the memory core — size, age,
churn, and whether the bytes were ever written
access SanitizerCoverage load/store instrumentation, so "allocated and
never touched" is measured rather than inferred
scaling k vs 2k replicas through clang profile counters, to catch
super-linear work per file
padding allocator slack per size class
ratchet the numbers recorded per venue so a regression is visible
The event lane needed a foreign-free counter: on Windows the override sees
frees for blocks it never allocated (they come from the CRT), and counting
those as untracked tripped the soundness gate. mem_events now separates
them (14 daemon / 5 worker on the Go corpus).
Two portability fixes the lanes surfaced, both real:
- the enable decision must not be cached before `environ` exists. Under
-fsanitize-coverage the UBSan preinit runs dlsym -> malloc -> our
wrapper before glibc sets up environ, so the lane silently disabled
itself and produced no dump at all on Linux.
- the access map needs 48 bits: arm64 Linux hands out addresses above
2^47, which a 47-bit map dropped on the floor.
mem_profile is retired into this: it measured a strict subset and had its
own sampling story.
Local-only by decision — not a PR-CI gate. _memwaste.yml runs the lanes on
all three platforms report-only, and pr.yml/release.yml carry the wiring.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Second increment on the memory core. The first landed the core, the linter
and phase attribution with zero production call sites; this one migrates the
memory that matters, proves where it goes, removes the waste the proof
exposed, and makes the memory budget a promise: over budget, extraction
results go to disk and the run completes at the floor instead of aborting.
Every number below is from a real index on this machine (M5 Pro, 48 GB) with
CBM_MEM_PHASES=1 and the release build.
WHAT THE PROOF SAID (Go corpus: 21,875 files, 345 MB of source)
The graph buffer, blamed for the peak since the 2026-07 research, is 3% of
the memory: 0.55 GB peak. The per-file extraction results are the memory:
14.8 GB written into their arenas, of which 3.4 GB is reachable. The rest
is temporaries -- cbm_node_text copies at 496 call sites, per-node QN
strings, abandoned generations of every growable array -- that an arena
cannot free and that the result kept alive until resolve ended, because
the result owned the arena. The arenas were charged 28.8 GB of capacity
for those 14.8 GB. On the kernel the same shape was 8.9 GB used by 89,731
results, and resolve grew it to 13.3 GB used / 17.2 GB capacity because
the C walk allocated its working set in the result arena and stored arena
pointers into the shared registry. After that, 7 GB of the kernel's
post-resolve memory had no class at all.
WHAT CHANGED
Attribution. The graph buffer (87 raw sites -> 0), arena blocks, the
tree-sitter slab allocator, the bound tree-sitter/SQLite allocators, the
semantic path (pass_semantic_edges, pass_similarity, minhash: 80 sites), the
hash table (Verstable, a class per table via CTX_TY: cbm_ht_create_in) and
CBM_DYN_ARRAY allocate through the core. Classes arena, ts_tree, hash_table
and dyn_array join the table. Two cross-boundary frees of buffer-owned
strings (pass_importance, pass_complexity) go through a new setter,
cbm_gbuf_node_set_properties_json. The linter baseline drops 4086 -> 3916.
Result compaction (internal/cbm/result_compact.c). At the end of a file's
extraction, everything reachable from its result is measured, copied into
one exact-size arena (strings interned by content within the file) and the
working arena is handed back. Later appends restart growth at the default
block (cbm_arena_init_exact, arena grow_size). The stale duplicate
internal/cbm/arena.h, which shared the foundation header's include guard
and won or lost by include order, is a shim now.
Record slimming. CBMCall 320 -> 72 bytes: the eight inline argument slots
(256 bytes, mostly empty) are allocated on first capture. CBMUsage 48 -> 40.
The per-file overlay contract (src/pipeline/pass_lsp_cross.c). One
lifecycle for every language's cross-file resolve: a scratch arena, a
per-file overlay registry chained to the sealed shared base, the walk
against the overlay, resolved calls appended into the result, scratch
destroyed. type_registry gains chain-aware iterators (every yielded index
belongs to it->reg) and copy-on-write refinement
(cbm_registry_func/type_for_update), and the C, C#, Go and Python walks
use them. This is what an ASan heap-use-after-free in the
lsp_resolution_probe suite demanded: the C walk stored per-call arena
objects into the shared registry, and a kernel run with a naive scratch
arena lost 1,045 edges to the resulting garbage reads. It also ends
unsynchronized writes into the shared registry from parallel resolve
workers.
The 7 GB with no class, found and cleaned. gbuf_index tables and keys were
2.5 GB live and 7.1 GB at peak: the 18 worker buffers built the five
secondary indexes they never query (cbm_gbuf_new_worker builds none), and
18M per-key index arrays started at 8 slots (cbm_da_push_min starts them
at 2); edge property JSON is 81% "{}" (interned). The semantic pass
allocated a 512-token stride per function up front (packed tokens with
offsets: cbm_sem_corpus_add_docs_batch). Edge dedup keys are 128-bit
hashes instead of sprintf'd strings (edge_key_map_t). Each worker keeps
one reusable working arena between files (cbm_work_arena_take/give:
rewind, not free) and every phase mark returns freed pages to the OS.
Budget metric. The over-budget check compares the charged footprint --
max(phys_footprint, mimalloc commit) on macOS -- not resident_size: after
extraction the kernel worker sat at 17.4 GB RSS with 5.5 GB charged, because
pages mimalloc has purged stay resident until the OS wants them, and
phys_footprint under-reports after MADV_FREE_REUSABLE cycles.
Spill / admission control (internal/cbm/result_spill.c, the extract gate
in pass_parallel.c). At 15/16 of the budget (or over it, or with
CBM_MEM_SPILL=1) the gate latches spill mode -- early, because the gate
sees the crossing per file pull and the workers' in-flight files carry the
charge past the line before the first sweep lands: every compacted result
is parked on disk as header + block in one append-only file per writer
under <cache>/spill/, the results already cached are swept out by every
worker, and registry build, def collection, surface rows, resolve and the
infra-route passes load a result only for the moment they read it
(relocated by base delta). The abort path stays closed while a sweep runs
or a cached result remains; it fires only at the floor -- graph +
registries + in-flight files. A result that owns sub-results stays in
memory; a retained parse tree is dropped with the parked result and
re-parsed on load. Every consumer of the result cache reads a slot through
cbm_pipeline_result_acquire()/release() (pipeline_internal.h): a parked
slot is NULL in the array, and the passes that indexed it directly lost
every __route__infra__ node and its INFRA_MAPS edges before the contract.
The other things the transient loads found: a def label that was the
result's own pointer (SIGBUS in the Go cross registry builder), two JVM
helpers that passed their input through, a dedup table whose keys borrowed
sr->value (the insert spun on freed memory), and a latch that was visible
before the store was open (a peer found nothing to park and failed the run
while its neighbour wrote 9 GB). The arena rewind reuse branch had
defeated the nblocks = CBM_ARENA_MAX_BLOCKS OOM seam the suites use; it
reuses only blocks that exist. The incremental and probe routes still
index the cache array themselves and therefore never spill
(ctx->spill_allowed stays false there; follow-up).
The semantic pass under a budget (src/pipeline/pass_semantic_edges.c). The
pass held every function's tokens, the corpus's per-document token ids and
the vectors being stored at once: 4.9 GB on the kernel, on top of the
floor, and nothing sized it to the headroom. Now, when the charged
footprint plus that transient would cross the budget, phases 2-4 run per
batch of functions: tokenize -> count into the corpus -> free; finalize
once; tokenize again -> vectorize -> store -> free, with pages returned to
the OS after every batch. Tokenization is deterministic, so both passes
see the same tokens and the graph does not change; the run pays with a
second tokenize. CBM_SEM_BATCH=<n> forces the batch size (the
batched-equals-unbatched test). The corpus frees its per-document token
ids at the end of finalize (the co-occurrence pass is their only reader),
and mem.semantic.step lines (CBM_MEM_PHASES=1) put the class's live/peak
bytes and the charged footprint at every sub-phase.
The semantic corpus and tokenizer allocate through the core (54 raw sites
-> 0). This is not attribution only: cbm_sem_tokenize handed out libc
strdup blocks that pass_semantic_edges freed with cbm_free -- mi_free in
the production build -- since the pass moved onto the core. No libc-backed
test build can see a cross-allocator free, so the core now refuses a block
it does not own (mi_is_in_heap_region) when CBM_MEM_PHASES=1 is set,
fatally: every proof run in this PR ran with it armed, and the Go and
kernel indexes report zero foreign blocks.
Accounting without contention. The core charged every allocation and free
to shared per-class atomics; with 18 workers on 18 cores every block bounced
the same three cache lines, and the bench against the shipped v0.10.8 showed
it as CPU time rising faster than wall time for the same graph: Kotlin 38 ->
121 s CPU (+58% wall), TypeScript 210 -> 596 s (+67%), Go 177 -> 312 s. The
deltas are thread-local now and reach the shared counters once per 256 KB or
512 blocks of change per thread, when a parallel-for worker ends, and
whenever a reader looks (its own thread first). Phase marks read after the
join and are exact; a class peak is a diagnostic and can be low by threads x
256 KB. Result: Kotlin 10.9 s / 43 s CPU (the residual +22% wall against
v0.10.8 tracks its +34% nodes), TypeScript 33.9 s (v0.10.8: 33.4), Go 48.4 s
(v0.10.8: 52.6), Java 81.1 s (v0.10.8: 79.8); graphs unchanged.
SQLite on a heap of its own. The same bench put the kernel's write step at
132 s where v0.10.8 needed 9.7 s, all of it in one coverage publish
statement. A stack sample named the cost: mimalloc's free-page search and
its periodic heap collect, walking the graph's pages -- millions of them
once the graph lives on the core -- for every statement-journal chunk
SQLite allocates. SQLite now allocates from a mimalloc heap of its own per
thread (mi_heap_new; mi_free is heap-agnostic, so nothing else changes),
and the periodic collect is set to the maximum mimalloc allows since the
explicit release at phase marks and spill sweeps is where memory goes
back. Coverage publish 132 -> 9.4 s; profiled kernel 376 -> 248 s
(v0.10.8: 296 s).
Formatting without a process lock. On macOS every vsnprintf takes the
process locale under an unfair lock (localeconv_l inside __vfprintf);
eighteen workers formatting type names on the C# corpus collapsed into
it -- the 10 MB JIT test files took 68 s each instead of under a second
-- and cbm_arena_sprintf paid it twice per string (size pass, write
pass). Each thread now formats with a C locale object of its own
(vsnprintf_l), and the common short string is formatted once into a
stack buffer.
Per-file budgets that cover the whole file. The 5 s per-file budget
covered the parse only; the LSP walks after it had none. C# JIT stress
files (a 23 MB single expression among them) parsed inside the budget --
the old wall-clock parse timeout used to drop them, which hid the rest --
and then held a worker for 346 s each in the usage walk (tree-sitter's
ts_node_parent descends from the root: quadratic on a deep tree), and the
cross-file resolve ran on the same trees. Three rules now, one site for
every language (cbm_extract_file_ex, honoured by cbm_pxc_dispatch_file): a
parse that used more than half the budget disqualifies the file from the
LSP walks (lsp_skipped); the unified walk checks its thread CPU time every
1024 nodes against six budgets and stops there, keeping what it found
(walk_truncated, implies lsp_skipped); a file whose parse plus walk spent
the budget is skipped by the LSP walks as well. Each decision is logged
with the path (extract.lsp.skipped, extract.walk.truncated). C# extraction
355 -> 39 s; the whole index 463 -> 128 s wall and 27 -> 11 GB peak
against v0.10.8 with 299 more nodes. The walk cuts exactly the five JIT
stress files (hugeexpr1, hugeSimpleExpr1, HugeArray1, HugeField1/2); the
15 MB System.Runtime.Intrinsics reference file walks to the end and, with
four generic-nesting JIT regression tests, only skips the LSP walk under
the parse-plus-walk rule.
The crash behind the C# bench. dotnet/runtime is 42,555 files, and its
resolve phase died with SIGBUS on every run of this branch. A libc-backed
ASan build of the server named it: heap-use-after-free in c_adl_resolve,
reading a type name that another worker had allocated in its per-file
scratch arena and freed at the end of its file. The C++ class walk
refines a method's return type when the declaration is more specific
than the pre-registered one (NAMED -> pointer/reference/template), and
it did that by casting the chained lookup result to non-const and
writing a scratch-arena signature into it -- into the sealed shared base
when the entry lived there. The overlay contract had exactly one
bypass, and it was this cast; every lookup returns const, so a grep for
the cast is the audit. The refinement now goes through
cbm_registry_func_for_update: copy-on-write into the overlay, the base
untouched. Test: clsp_method_return_refinement_is_copy_on_write (a
sealed base with a NAMED method return, an in-class pointer declaration;
the base signature pointer is unchanged and the overlay copy carries the
refinement plus a marker only a copy has) -- RED on the cast, GREEN on
the accessor. The fixed ASan build then indexed the whole
corpus clean: 42,555 files, 1,224,981 nodes, 5,794,320 edges, worker
exit 0, no report.
Production build backs the core with mimalloc explicitly (never the global
override, which macOS's two-level namespace forbids); the test build keeps
libc so the sanitizers see every block -- the same split the tree-sitter and
SQLite bindings already use.
PROOF
Go peak RSS 16.7 -> 4.0-5.2 GB; result arenas 14,804/28,817 -> 856/856
MB; wall 55 -> 52-60 s; graph inside the pre-existing run-to-run
jitter (gRPC Route nodes: two identical runs differ by 21).
Go CBM_MEM_SPILL=1: peak RSS 1.7-2.3 GB, 21,875 results parked (950 MB
on disk), graph identical: 39 infra routes, 158 INFRA_MAPS, 21,875
surface rows, nodes/edges inside the Route jitter.
Go CBM_MEM_BUDGET_MB=1200: spill latched at 1,204 MB, the first sweep
took the charge to 768 MB, the run completed at 1.99 GB peak RSS.
Kernel DEFAULT budget (24,576 MB): aborted before, completes now without
spilling. Extraction 35.2 -> 16.5 GB RSS (5.6 GB footprint, 15.3 GB
committed); results 8.9 GB = capacity; resolve tracked 24.6 -> 19.1
GB; peak RSS 27.3 -> 20.1 GB. Nodes identical (8,529,729); edges
inside the +-400 run-to-run band the kernel shows between identical
runs.
Kernel CBM_MEM_BUDGET_MB=16000: aborted before, completes now through the
spill path. Latched at 17.2 GB charged; 89,731 results parked (9.3
GB on disk, 370,188 loads); extraction end 6.2 GB RSS / 5.1 GB
committed; resolve 8.6 GB RSS / 11.5 GB committed; peak RSS 16.6
GB (the semantic pass plateau, 13.3 GB RSS, is graph-derived and
outside what spill can move). Nodes identical; edges 15,773,598.
Kernel CBM_MEM_BUDGET_MB=15000, before the semantic batching: completes,
every phase mark under 15,000 MB by every metric, but the process
RSS high-water mark reached 16.9 GB inside the semantic pass (class
peak 4.9 GB against 0.8 GB live).
Kernel CBM_MEM_BUDGET_MB=15000, with it: completes; the charged footprint
stays under 15,000 MB through the whole semantic pass (14,018 ->
14,381 MB across its sub-phases, 6 batches of 125,644 functions);
every phase mark is under budget by every metric (max footprint
12.6 GB, max commit 11.5 GB); the process RSS high-water mark is
15.65 GB and is set in extraction, at the moment the gate latches
(15,002 MB charged) and before the first sweep lands -- a 4%
overshoot of in-flight work, no longer the semantic pass. Nodes
identical, edges 15,772,913, zero foreign blocks.
Kernel CBM_MEM_BUDGET_MB=15000, with the anticipatory latch: spill mode
enters at 14,062 MB (near_budget); the charged footprint never
exceeds 11.8 GB at any mark; the process RSS high-water mark is
15.1 GB in extraction and 15.3 GB during the final write-out --
1-2% over, the in-flight window and the dump transient. Nodes
identical, edges 15,772,543, zero foreign blocks.
Go CBM_MEM_BUDGET_MB=1200 with it: latched at 1,131 MB, extraction
peak 998 MB RSS, graph identical to the baseline.
BENCH AGAINST THE SHIPPED v0.10.8 (same machine, same driver, 2026-09-14)
corpus wall s peak RSS GB nodes / edges / CALLS
perl 5.7 -> 5.3 0.09 -> 0.05 -3 / -130 / -3
php 6.3 -> 5.8 0.66 -> 0.20 +21.9% / +1.4% / +23.8%
rust 6.8 -> 6.3 1.15 -> 0.67 +1 / +56 / -369
c 5.9 -> 5.7 0.16 -> 0.19 identical
kotlin 8.9 -> 9.4 1.55 -> 0.96 +33.5% / -22.8% / +12.1%
django 10.9 -> 10.2 3.13 -> 1.12 +6 / -14.5% / -28.4%
go 52.6 -> 42.9 17.2 -> 3.9 +14.2% / -10.3% / +2.2%
java 79.8 -> 73.8 27.6 -> 11.0 -121 / -1.9% / -0.6%
csharp 463.3 -> 128.1 27.0 -> 11.2 +299 / -0.9% / -0.6%
typescript 33.4 -> 33.6 12.0 -> 2.5 -186 / -3.5% / +130
kernel 294.8 -> 233.3 27.5 -> 25.1 -171 / -1.3% / -0.1%
No corpus is slower beyond noise; CPU time follows wall (kernel 2,442 ->
1,403 s). Every graph delta against v0.10.8 is main between the release and
the branch base (339b3f40), not this PR: a build of the base commit run
through the same driver reproduces the final graphs exactly on django and
kotlin, within the pre-existing gRPC Route jitter on go (30 nodes, 27
edges, CALLS identical), and on C# with the same node count. This PR's
own C# difference against the base is 5 CALLS (the five truncated JIT
stress files) plus one ijwhost swap and 563 edges of 5.79 M; the base
takes 448 s / 29.7 GB for that corpus.
Diagnostics that made this measurable stay in: the extract.arenas and
extract.census lines (behind CBM_MEM_PHASES=1), the CBM_MEM_RELEASE=1
release-to-OS probe, and peak_charged_mb on every mem.phase line -- the
high-water mark of the budget metric itself (cbm_mem_peak_charged), next
to the RSS peak that counts pages already purged but not yet reclaimed.
Tests: extract_compact_* and extract_spill_round_trip_keeps_every_field
(extraction), parallel_spill_mode_builds_the_same_graph (parallel: every
file parked, node count and every edge type equal to the in-memory run),
pipeline_semantic_batched_matches_unbatched (pipeline: one repo indexed
unbatched and with CBM_SEM_BATCH=5 -- byte-identical node vectors, token
vectors and SEMANTICALLY_RELATED edges),
extract_lsp_skipped_when_parse_used_ its_budget_share and
extract_walk_truncated_at_its_cpu_budget (extraction: the seams mark a
file / stop the walk, the LSP walk and the dispatcher skip it, the defs
stay), registry_overlay_chain_iterates_ and_copies_on_write (c_lsp), arena
exact-init/growth, mem charged/footprint/ pressure.
Not in this increment: run-to-run edge jitter (+-400 on the kernel, Route
class on Go) predates this work and has its own item; the incremental route
spilling; a tighter result encoding (def ids instead of QN strings in
resolved calls) is the next lever on the 9 GB the kernel's results still
take.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The memory-core linter scanned src/ and internal/cbm/. Widen it to src/ and every subtree under internal/, so a new tree added there is covered without anyone remembering to list it, and the count is for the whole project -- cli, mcp, daemon, store, pipeline, the extraction engine. Vendored code is now excluded wherever a vendored/ segment appears in the path rather than by one hardcoded prefix, so that exemption cannot drift either.
internal/ holds only cbm today, so the regenerated baseline is byte-identical: 92 files, 4086 raw sites. Widening lost no coverage and vendored stays out (0 sites listed). The red/green proof still holds under the new scope, and the failure message now names the NEW site -- one appended malloc reports grew by 2 (83 -> 85) at the appended line, instead of listing six pre-existing ones.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Memory in this project is now allocated through ONE core. This lands the
core, the gate that keeps new code on it, and the instrumentation that lets
an index run say which class of memory grew in which pass -- the three things
the 2026-09-13 audit found missing.
WHAT THE AUDIT FOUND
There was no centralized allocation layer. foundation/mem.h was policy and
measurement only (budget, RSS, over_budget, memory map, phase marks) with no
alloc/free wrapper, so the budget could be observed after the fact and never
enforced. Raw allocator calls in the tree: 4086 across 92 files. The
extraction engine had already adopted arenas (1301 call sites); the graph
buffer -- which holds the 35 GB peak of a kernel index -- used none,
allocating every 64 B node, 48 B edge and every name/qualified_name/
properties_json individually.
Two further facts made a self-accounting core the only workable design:
1. cbm_mem_map_collect() walks THIS thread's mimalloc heap only. Walking
the process-wide heap is a data race TSan caught on macOS, so it is
deliberately not done, and an 18-worker index attributes almost
nothing -- the rest lands in residual.
2. The mimalloc global override is ON for Linux/MinGW and permanently OFF
for macOS (two-level namespace: this binary's free becomes mi_free while
system libraries keep allocating from the system zone, and a pointer
crossing that boundary aborts). On macOS ordinary malloc is served by
the system allocator; the startup audit reports owned_classes=0/6. Any
accounting that assumes mimalloc owns the pointer is blind on an entire
platform.
THE CORE (src/foundation/mem_core.{h,c})
cbm_alloc / cbm_calloc / cbm_realloc / cbm_mem_strdup / cbm_free, each tagged
with an allocation class (gbuf_node, gbuf_edge, gbuf_string, gbuf_index,
extract, semantic, dump, store, other). Per-class live bytes, live blocks and
peak, kept in atomics -- thread-safe by construction, identical on every
platform, independent of which allocator serves malloc. libc semantics are
preserved exactly (alloc(0) is freeable, free(NULL) is a no-op, realloc(NULL)
allocates, a failed realloc leaves the block intact) so adoption is a
mechanical rename and never a behaviour change.
No per-allocation header: a {class,size} prefix costs 16 bytes on ~100M blocks
at kernel scale, 1.6 GB of overhead to measure a memory problem. Sizes come
from the platform usable-size query (malloc_size / malloc_usable_size /
_msize), which is correct under either allocator and free. Where no query
exists (BSD) the request size is charged, which understates -- the safe
direction for a diagnostic. A mismatched class on free clamps at zero rather
than wrapping, so a small drift can never masquerade as a colossal leak.
Arena-backed allocators report in bulk through cbm_mem_class_add_external /
remove_external so their memory appears in the same table; the extraction
engine must not be rewritten to per-object allocation, that would undo the
batching that keeps its allocation count low.
THE LINTER (scripts/lint-memory-core.py, wired into make lint-ci)
Raw malloc/calloc/realloc/free/strdup/strndup outside the core is a defect.
4086 sites cannot migrate at once, so the gate is a ratchet: a checked-in
baseline (scripts/memory-core-baseline.txt) records each file's count and
the build goes red when any file grows, or a new file appears with any.
Files only ever go down; the baseline line is lowered in the change that
migrates the file, and --strict fails a stale baseline. Comments and string
literals are stripped first, so prose that mentions malloc( cannot trip it --
the security audit already bit on exactly that with fork(. Proven both ways:
one appended malloc turns it red naming the file (83 -> 85), restoring turns
it green. Scans src/ and internal/cbm/: cli, mcp, daemon, store, pipeline.
PHASE ATTRIBUTION (src/pipeline/pipeline.c)
cbm_mem_phase_mark and the new cbm_mem_class_log now fire at pipeline.begin
and at every pass.timing site: parallel_extract, the six sequential passes,
the eight predump passes, tests and dump_and_persist, with peaks reset per
index. Both instruments existed in foundation/ and were wired only into MCP
request handling, never into the index pipeline -- which is where the memory
is. Locating the kernel peak this week required an external RSS sampler
because nothing in-process could say which phase it was in.
PRESSURE PRIMITIVES (system_info.c, mem.{c,h}, platform.h)
cbm_system_available_ram() -- macOS host_statistics64 (free + inactive +
purgeable), Linux MemAvailable, Windows GlobalMemoryStatusEx; 0 when unknown,
never cached -- and cbm_mem_system_under_pressure(), true below 12.5% of RAM
available and false when unknown so nothing ever aborts on a guess. Landed
UNWIRED: an earlier draft used them to let the over-budget latch press on
whenever the system was not under pressure, and that permitted the very
overshoot this work exists to end (Decision A's tests at
test_pipeline.c:13481 and :13583 went red under it, correctly). They are the
backstop for the next step, admission control, which keeps a run under budget
before an allocation rather than measuring it after.
mem suite 10 new tests green; pipeline + mem 339 passed 0 failed; lint-ci
clean including the new gate.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
scripts/memlab.sh gave the profiled run a private CBM_CACHE_DIR only, so
the process joined the operator's account daemon and cleanup removed the
cache without stopping any daemon. Source scripts/test-runtime.sh for a
harness-owned runtime and cache, and stop the private daemon before the
work directory is removed.
Add tests/test_memlab_runtime_isolation_contract.sh, which fails before
this change, and wire it into scripts/test.sh.
Part of #1696.
Signed-off-by: Anton Standrik <astandrik@yandex-team.ru>
scripts/soak-test.sh claimed to isolate daemon coordination from
interactive sessions through a private CBM_CACHE_DIR, but only
CBM_RUNTIME_DIR moves the daemon rendezvous, so the soak shared the
operator's account daemon and asserted that its own shutdown stopped it.
Source scripts/test-runtime.sh for a harness-owned runtime and cache,
and let the helper stop the private daemon before the root is removed.
Add tests/test_soak_runtime_isolation_contract.sh, which fails before
this change, and wire it into scripts/test.sh.
Part of #1696.
Signed-off-by: Anton Standrik <astandrik@yandex-team.ru>
The codeql-gate has never checked anything. It declared no permissions block,
so it inherited the workflow's contents:read and the code-scanning alert API
answered 403. gh writes the API error BODY to stdout, so
ALERTS=$(gh api '.../code-scanning/alerts?state=open' --jq 'length' 2>/dev/null || echo "0")
left ALERTS holding
{"message":"Resource not accessible by integration",...,"status":"403"}0
every [ ... -gt ... ] then failed as a non-integer comparison, the if took its
false branch, and the step printed "CodeQL gate passed (0 alerts)" and exited 0.
Observed in the log of a GREEN run (34700514102, job 103571322660) while the
repository had an open high-severity alert. The `|| echo "0"` reads as defensive
and is the exact opposite: it converts "I cannot see the alerts" into "there are
no alerts", which is the one direction a security gate must never fail in.
Grant the job security-events:read (and actions:read, which the CodeQL-run poll
needs for the same reason), and replace the fallback with a helper that treats a
non-zero gh exit or a non-numeric body as a hard failure. The helper returns
rather than exits because it runs inside a command substitution, where an exit
would end only the subshell and let the caller continue with an empty count --
the same silent pass in a new costume; both call sites check the status.
tests/test_security_gate_fail_closed.sh pins both halves and fails with six
assertions against the previous workflow.
Also bumps graph-ui's vitest to ^4.1.11 (GHSA-82fw-gwwq-j7x9, path traversal via
@vitest/mocker, CVSS 5.9). Development scope, and graph-ui ships in no release
artifact, so nothing released was ever exposed -- but it was the sole open alert
and the gate above is worthless while it stands unresolved. Raising the floor
past the vulnerable range also stops a fresh install resolving back into it.
graph-ui: npm ci clean, 47 tests in 12 files pass on 4.1.11.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The macos-15-intel leg killed the cli suite at its 900s wall clock while
reporting 7537 passed, 0 failed. Two separate things were wrong.
The suite was misclassified. cli spends 497s of the 900s default on macos-14,
the FASTEST macOS runner, while macos-15-intel in the same matrix is 2.4-3.6x
slower on comparable suites (daemon_runtime 842s vs 349s, stack_overflow_b
277s vs 76s). 497s at that ratio cannot fit. daemon_runtime, at 842s on that
same runner, passes only because it was already in SLOW_SUITES. cli belongs
there too -- it is deterministic and simply large, which is the only thing that
list is allowed to mean.
The drain test was also waiting on a clock instead of on the system. Profiling
the suite per test (315 tests, 300s wall / 166s cpu) put
cli_install_into_host_namespace_still_drains_host_cohort at 42s wall for 19s of
cpu -- 23s idle, the largest single idle block, where the next worst was 11s.
The existing cli_install_force_quiesces_active_cohort_before_replacing_binary
drains a cohort too and idles 0.5s, which is what pointed at the difference.
The cost was the host_serving probe. Its generous budget exists for the
POSITIVE question -- is this daemon still up? -- where a slow reply on a loaded
runner must not be misread as drained; that fixed a real flake and is kept.
Asked in the negative it inverts: no reply is coming, so the full 15s is spent
establishing silence and ASSERT_FALSE is decided by the timeout expiring rather
than by the daemon's behaviour. The drain is already proven positively in that
test -- install returns 0 only after the activation completed, and the host
child exits CLI_SCOPE_HOST_DRAINED -- so the probe only has to confirm it.
Also shortens the drained child's teardown budget, which retried service_free /
lease_release / manager_free against 10s on a path where the activation had
already torn the service down. Behaviour at the deadline is unchanged; it is
reached sooner when nothing is wedged. Measured at only -3s on its own, kept
because it is correct and free.
Local, same machine, ASan+UBSan: the drain test 42.0s -> 25.8s (-38%), the
suite 300s -> 279s, 315 passed before and after -- no test removed, no
assertion weakened. PR #2188, which runs the same leg without these two tests,
passed macos-15-intel in 37m34s, so the suite was not already over budget on
main.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
A reporter brought a full ProcMon trace to #1856 and still could not say which
check refused their install. That is not their failing: the binary discards the
answer before printing it.
cli_activation_production_context_init() validates the cache, rendezvous and
log directories -- it is the emitter MOST likely to hold a useful detail -- and
it printed CLI_ACTIVATION_REFUSED_MESSAGE bare, through the raw sink rather
than cli_activation_diagnostic(). So the reader got "Check the errors above"
with nothing above. That is the exact dead end #1416 and #1537 already fixed;
the property was repaired on the paths that had tests and left broken on this
sibling, which had none. Route it through the attributing helper: the refusal
now names the transaction refusal note or cbm_daemon_ipc_validation_detail().
The remedy was also wrong for the reported shape. `icacls <dir> /remove:g <sid>`
cannot remove an INHERITED ACE, and the stock C:\ grant for Authenticated Users
reaches every new child directory exactly that way -- so a reporter following
our advice watches the command succeed and the refusal persist. The refusal now
says whether the ACE was inherited, and the advice names
`icacls <dir> /inheritance:r /grant:r "%USERNAME%":(OI)(CI)F` for that case.
No safety check is relaxed. An install directory writable by other accounts is
still refused; it is now refused in terms the reader can act on.
The new contract test guards the CLASS rather than this one call site: any
emitter of a refusal constant that bypasses the attributing helper fails it,
which is what would have caught this. It fails on the previous tree naming
src/cli/cli.c:750, plus the missing inherited-ACE report and the missing
/inheritance:r remedy.
Local: cli 314 passed; activation_transaction daemon_ipc daemon_bootstrap
daemon_version 110 passed, 1 skipped (user namespaces are Linux-only, runs on
the Linux leg); make -f Makefile.cbm lint-ci clean.
Reported by zaferavci1.
Refs #1856
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Indexing accepts whatever a repository contains. A tree carrying a vendored
monorepo, a generated dump, or a runaway build directory is discovered in
full, and the first sign of trouble is a host under memory pressure with
nothing that attributes it to indexing.
Add two opt-in limits evaluated during discovery against accepted source
files only: index_max_files and index_max_source_mb. Both default to off, so
nothing changes until an operator sets one. Crossing a limit fails the whole
attempt with a structured resource_limit_exceeded result naming the resource,
the observed value and the limit; no partial graph is published, and an
existing serving index keeps answering.
Limits are read from the CLI-managed _config.db and are not MCP request
arguments. A supervised parent replaces any caller-supplied policy before
spawning its worker, and the worker rejects a missing or incomplete contract,
so the CLI, the daemon and the supervised worker all enforce the same
decision.
Both keys reach an operator through the existing config get/set/list/reset
with no new subcommand. `set` suppresses its own generic message for them
because the policy writer names the precise reason -- so that writer speaks
on every failure it can return, including a validated value whose write then
fails on a database that cannot be written. Exiting non-zero in silence is
not an acceptable answer from a CLI.
The two shell regressions that hand-roll the supervisor's worker argv carry
that contract as well. Without it the worker exits before either guard can
observe anything, and the guard would go quietly vacuous.
Signed-off-by: 刘冲 <mail@liuchong.dev>
The daemon-spawned index worker inherited the daemon process environment and
re-ran the allowed-root check without a session policy. The daemon's session
policy is installed only in application.c (the sole production caller of
cbm_mcp_server_set_session_context); after the daemon admits an
index_repository request, index_supervisor.c builds the worker argv
(`cli --index-worker --index-worker-build <fp> index_repository <args_json>
--response-out ...`, no root anywhere) and subprocess.c spawns it with the
daemon's environ. The worker then creates its own server with
cbm_mcp_server_new(NULL), so handle_index_repository runs with no session
policy and falls back to getenv("CBM_ALLOWED_ROOT") - the environment of
whoever STARTED the daemon (bootstrap.c spawns it with the starter's environ).
Observed: a daemon started with CBM_ALLOWED_ROOT=<rootA>; a client session for
<rootB>/tiny is admitted by the daemon (index.supervisor.reap outcome=clean
exit_code=0) and then refused by the worker with "<rootB>/tiny is outside the
allowed root. To allow it, run: codebase-memory-mcp allow-root ...". Any daemon
whose starter had CBM_ALLOWED_ROOT set broke every other session outside it.
Fix: in the worker arm of run_cli, right after the worker's server is created
and before any tool runs, read repo_path from the request args (both daemon
spawn paths rewrite the worker args with the canonical repo_path), canonicalize
it and install it explicitly as both session root and allowed root. The daemon
already authorized the request; the worker only executes it. A repository that
cannot be canonicalized does not exist and can only have been admitted by a
session with no declared boundary (containment needs a real path), so the
worker mirrors that with an explicit unrestricted policy - still never the
environment fallback - and the pipeline keeps reporting the missing repository
as the tool error it always was. A request without repo_path fails closed with
"request workspace scope invalid". No argv or IPC change, no process
environment mutation.
Regression test: tests/test_worker_session_scope.sh drives the real binary as a
supervised worker with CBM_ALLOWED_ROOT=<rootA> and repo_path=<rootB>/tiny and
requires an indexed response with no boundary refusal; a worker without
repo_path must exit nonzero on the scope check. Wired as Step 5f of
scripts/test.sh next to the other real-binary worker contracts. RED on main
(the exact "outside the allowed root" refusal above), GREEN with this change.
Distilled from #1925 with co-author credit.
Verification: test-runner suites index_supervisor, daemon_application, daemon,
mcp, cli green; tests/test_worker_error_response.sh, test_worker_watchdog.sh
and test_parent_watchdog.sh green against the fixed binary; make lint-ci and
scripts/check-no-test-skips.sh clean.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Co-authored-by: Zhiyu <zhiyuzhang001@gmail.com>
Path.write_text is open-then-write: it creates and truncates the file
first and the content lands afterwards. The parallel harness contract
polls for `<suite>.ready` to appear and then parses its content as the
leader pid, so a reader that wins that window sees an empty file and
fails with `ValueError: invalid literal for int() with base 10: ''` --
a spurious FAIL of the harness contract with a scheduler-side cause.
The same shape applied to `<suite>.leader-exited`.
Publish the barrier files through a helper that writes a same-directory
temp file and os.replace()s it onto the destination, which is atomic on
POSIX and Windows: a poller now sees either no file or complete content.
All three barrier publications go through it.
Pin the contract structurally (no timing): the scheduler may not write a
barrier file in place, and it must rename into place instead. The pin is
RED on the previous scheduler.
Distilled from #1188 with co-author credit.
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Co-authored-by: Kody <kaidi.shi.1121@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Fix-forward after rebasing onto main (545 commits, 26 conflicted-file
resolutions). Where main had shipped a behaviour in the meantime, main
wins; the PR keeps its output contract on top of it:
- search_code: the scan runs under main's supervised, deadline- and
output-bounded runner and feeds the branch's streaming classifier from
the runner's output file; the POSIX wrappers map grep's no-match to 0,
so every non-zero exit fails closed (no exit-1 special case).
- trace_path: strategy/confidence precede args in header and rows (#1542),
in both the tree table and the json legs.
- cypher: relationship expansion is never capped (a cap falsified
aggregates) and unlabeled MATCH scans every candidate; `truncated` is
raised only by a real ceiling — a clamped hop range is probed one hop
beyond the cap and reports truncated only when a candidate exists there.
The plain BFS and the trail BFS both probe one row past max_results so
saturation is observable instead of silent.
- search_code full mode and the final json use main's U+FFFD sanitizer
instead of '?' replacement.
- search_graph json keeps an empty `groups` array for semantic-only calls;
list_projects json echoes offset/limit next to next_offset and accepts
main's include_details spelling for the stats projection.
- discovery budget: main added get_file_outline, compare_graphs, manage_adr
set_sections and the debug/include_details parameters; their descriptions
are trimmed to the lean style and the ceiling moves from 15 to 18 KiB.
Tests that asserted a JSON default now request format:"json"; the
fail-closed search test uses an unreadable regular file (a missing or
non-regular indexed path is deliberately not a scan operand on main); the
non-regular fixture gives its real sources File nodes, as every indexed
project has; list_projects total counts listable projects (the 0-byte
ghost is never a row, so it is not a page either).
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>