Files
codegraph/__tests__/daemon-candidates.ts
Colby MchenryandClaude Opus 5.5 c405ac5e6c fix(windows): make load-sensitive tests wait on real signals; longer pid-file retry (#1773) (#2078)
* test: make Windows load-sensitive tests wait on the tree, not the clock (#1773)

Each failure from the loaded Windows run was traced to its mechanism and
reproduced deterministically before changing it. No assertion is weakened.

- daemon-lock-refresh: the PowerShell holder slept 200ms against a 375ms
  retry budget. Windows refuses a replace-rename while any handle to the
  target is open, even a Node read handle, so hold one in-process and close
  it inside the rename mock after the first real failure. Also asserts the
  first rename really failed with a retryable code.
- mcp-projectpath-lifecycle: the owner daemon armed a 500ms idle timer at
  start and exited (code 0) before a starved engine reached it. It now starts
  with idle exit off and is armed over IPC once the engines hold a session.
  The file waits for catch-up instead of the gate's 3s serve-anyway deadline,
  and its index-heavy hooks and eviction case get explicit timeouts.
- mcp-daemon / mcp-writer-lock teardown: racing launchers spawn detached
  daemon candidates whose cwd is the fixture; a loser still starting at
  teardown pinned it (EBUSY) or took over after the winner was stopped. A
  --require preload records candidate pids, and teardown waits for the
  losers before stopping the winner.
- mcp-writer-lock: writer.pid appears before the daemon binds and rewrites
  daemon.pid, so cleanup could probe a daemon mid-start ('unverified'); wait
  for both proxies to attach. Removal uses fs.promises.rm, because fs.rmSync
  gives up on the first Windows access-denied error despite maxRetries.
- mcp-daemon #1963: connect only after the proxy attached (lock before bind).
- sync-rebuild-convergence / multi-repo-workspace: explicit timeouts scoped
  to the cases that build real indexes; the global timeout is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(daemon): give the Windows pid-file replace a longer retry budget (#1773)

Windows refuses to replace a file while any handle to it is open, including a
plain Node read handle, so an antivirus scan, an indexer, or another session
reading daemon.pid can each block the daemon's startup refresh for a moment,
and for longer on a loaded machine. The retry waited 375ms in all; it now backs
off 25, 50, 100, 200, then 400ms three times, about 1.6s over eight attempts,
still well inside a launcher's ~6s connect window. Permanent failures and
non-Windows errors still abort at once, and an ownership change still stops
the retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 18:13:35 +00:00

64 lines
2.6 KiB
TypeScript

/**
* Teardown for tests whose `serve --mcp` launchers race to spawn the shared
* daemon (#1773). Each launcher that finds no daemon spawns a detached
* candidate; one wins the lock and the rest exit once they see it. Stopping the
* winner while a loser is still starting lets that loser take over the fixture
* being removed, and on Windows a live process whose working directory is the
* fixture makes the removal fail with EBUSY or EPERM. So the teardown waits for
* the losers, by pid, before it stops the winner.
*/
import * as fs from 'fs';
import * as path from 'path';
const SPAWN_RECORDER = path.resolve(__dirname, 'fixtures/record-detached-spawns.cjs');
/** Where launchers started for `fixture` record their candidates: beside it, never inside. */
export function spawnLogFor(fixture: string): string {
return path.join(path.dirname(fixture), `${path.basename(fixture)}.daemon-spawns`);
}
/** Node arguments and environment that make a launcher record its candidates. */
export function recordSpawns(fixture: string): { args: string[]; env: NodeJS.ProcessEnv } {
return { args: ['--require', SPAWN_RECORDER], env: { CG_TEST_SPAWN_LOG: spawnLogFor(fixture) } };
}
function recordedPids(fixture: string): number[] {
let raw = '';
try { raw = fs.readFileSync(spawnLogFor(fixture), 'utf8'); } catch { return []; }
return raw.split('\n').map(Number).filter((pid) => Number.isSafeInteger(pid) && pid > 0);
}
function isAlive(pid: number): boolean {
try { process.kill(pid, 0); return true; } catch (err) {
return (err as NodeJS.ErrnoException).code === 'EPERM';
}
}
/**
* Resolve once every recorded candidate except the current lock holder has
* exited. Call it after the launchers have exited, so no new candidate can
* appear. The losers leave on their own; this never signals one, so a loser
* that stays alive fails the teardown instead of being hidden.
*/
export async function settleLosingCandidates(
fixture: string,
holder: () => number | null | undefined,
timeoutMs = 30_000,
): Promise<void> {
const deadline = Date.now() + timeoutMs;
for (;;) {
const current = holder();
const pending = recordedPids(fixture).filter((pid) => pid !== current && isAlive(pid));
if (pending.length === 0) return;
if (Date.now() > deadline) {
throw new Error(`Daemon candidates ${pending.join(', ')} were still running with ${current ?? 'no process'} holding the lock`);
}
await new Promise((resolve) => setTimeout(resolve, 25));
}
}
/** Remove the spawn record once the fixture's processes are gone. */
export function removeSpawnLog(fixture: string): void {
fs.rmSync(spawnLogFor(fixture), { force: true });
}