Files
Colby MchenryandClaude Opus 5.5 a192df40f9 fix(telemetry): count late index uploads, treat usage as live ingest, purge before catch-up (#2333) (#2356)
Three bugs in the self-hosted telemetry services (telemetry-worker/ and
telemetry-dashboard/), reported against 6560052a and all still on main.

1. An index run that uploads late never counted toward activation. The
   dashboard's funnel reads machine_first_seen.first_index_day, which only the
   nightly rollup wrote, and the rollup re-rolls just the last three days plus
   days with no daily_machines row. Ingest accepts timestamps up to 30 days old
   and the client keeps an event's original timestamp when it re-queues it, so
   an index event 4+ days late landed on a day nothing revisited and
   first_index_day stayed NULL. The ingest upsert of machine_first_seen now
   lowers first_index_day the way it lowers first_day:
   coalesce(min(old, new), old, new), because SQLite's min() is NULL if either
   side is. It is the same statement and row, with no new index and no
   migration (the column exists since 0002). The rollup still re-derives it
   from raw events, which covers events stored before this change.

2. ingest_stalled read only max(day) FROM events, but usage counters have
   lived in usage_daily since 0003, so a day with usage and no lifecycle events
   read as "The ingest worker ... is not storing anything". /api/meta now also
   reads max(day) FROM usage_daily (a primary-key lookup) and returns
   latest_usage_day and latest_ingest_day. Ingest counts as stalled only when
   there is no lifecycle event from yesterday or later and no usage counter
   from the day before yesterday or later, because clients upload a day's
   counters only after that day ends. The banner names latest_ingest_day.

3. runNightly rolled up the 3 regular days, then up to 31 missed days, and ran
   purgeOldEvents last. Each day first folds its legacy usage_rollup rows,
   which takes minutes on a heavy day, so a backlog could run past Cloudflare's
   15-minute Cron Trigger limit and skip the purge and the summary line night
   after night. The purge now runs first: it is bounded and touches only days
   past the window, which no rollup reads. After that, rollups stop starting
   new days or fold chunks 10 minutes in (NIGHTLY_BUDGET_MS). A day stopped
   mid-fold gets no rollup at all, so it stays a missed day and the next night
   continues from its last committed chunk. The summary line counts those days
   as `deferred`, and `caught_up` now counts missed days actually rolled up.

Verification: __tests__/telemetry-services.test.ts (new; the root vitest
config picks it up) runs the worker's fetch/cron code and the dashboard API
against the checked-in migrations in in-memory node:sqlite through a small
D1-shaped adapter. With the source fix reversed, all 5 tests fail (activation
0, first_index_day null, ingest_stalled true, and "nightly rollup incomplete:
4 day(s) failed, purge failed" once statements pass the 15-minute mark). With
the fix, all 5 pass. `npm run check` is clean in both packages. Against
wrangler dev, run locally only: smoke:rollup 71/71, smoke:cutover 62/62,
smoke:api 122/122 and render-check 87/87. smoke (ingest) is 46/47; the one
failure is the 70 KB oversized-body curl argument hitting the Windows
command-line limit, and main fails it the same way. The new smoke assertions
fail on main.

Takes effect only once both workers are redeployed; no migration.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-05 20:06:42 +00:00
..