mirror of
https://github.com/colbymchenry/codegraph.git
synced 2026-10-07 23:02:51 +08:00
Three bugs in the self-hosted telemetry services (telemetry-worker/ and
telemetry-dashboard/), reported against 6560052a and all still on main.
1. An index run that uploads late never counted toward activation. The
dashboard's funnel reads machine_first_seen.first_index_day, which only the
nightly rollup wrote, and the rollup re-rolls just the last three days plus
days with no daily_machines row. Ingest accepts timestamps up to 30 days old
and the client keeps an event's original timestamp when it re-queues it, so
an index event 4+ days late landed on a day nothing revisited and
first_index_day stayed NULL. The ingest upsert of machine_first_seen now
lowers first_index_day the way it lowers first_day:
coalesce(min(old, new), old, new), because SQLite's min() is NULL if either
side is. It is the same statement and row, with no new index and no
migration (the column exists since 0002). The rollup still re-derives it
from raw events, which covers events stored before this change.
2. ingest_stalled read only max(day) FROM events, but usage counters have
lived in usage_daily since 0003, so a day with usage and no lifecycle events
read as "The ingest worker ... is not storing anything". /api/meta now also
reads max(day) FROM usage_daily (a primary-key lookup) and returns
latest_usage_day and latest_ingest_day. Ingest counts as stalled only when
there is no lifecycle event from yesterday or later and no usage counter
from the day before yesterday or later, because clients upload a day's
counters only after that day ends. The banner names latest_ingest_day.
3. runNightly rolled up the 3 regular days, then up to 31 missed days, and ran
purgeOldEvents last. Each day first folds its legacy usage_rollup rows,
which takes minutes on a heavy day, so a backlog could run past Cloudflare's
15-minute Cron Trigger limit and skip the purge and the summary line night
after night. The purge now runs first: it is bounded and touches only days
past the window, which no rollup reads. After that, rollups stop starting
new days or fold chunks 10 minutes in (NIGHTLY_BUDGET_MS). A day stopped
mid-fold gets no rollup at all, so it stays a missed day and the next night
continues from its last committed chunk. The summary line counts those days
as `deferred`, and `caught_up` now counts missed days actually rolled up.
Verification: __tests__/telemetry-services.test.ts (new; the root vitest
config picks it up) runs the worker's fetch/cron code and the dashboard API
against the checked-in migrations in in-memory node:sqlite through a small
D1-shaped adapter. With the source fix reversed, all 5 tests fail (activation
0, first_index_day null, ingest_stalled true, and "nightly rollup incomplete:
4 day(s) failed, purge failed" once statements pass the 15-minute mark). With
the fix, all 5 pass. `npm run check` is clean in both packages. Against
wrangler dev, run locally only: smoke:rollup 71/71, smoke:cutover 62/62,
smoke:api 122/122 and render-check 87/87. smoke (ingest) is 46/47; the one
failure is the 70 KB oversized-body curl argument hitting the Windows
command-line limit, and main fails it the same way. The new smoke assertions
fail on main.
Takes effect only once both workers are redeployed; no migration.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>