Files
Colby MchenryandClaude Opus 5.5 217190680d fix(telemetry): recover from the D1 10 GB outage; dashboard starts on today (#2317)
The telemetry database hit D1's 10 GB cap on 2026-08-11 and refused nearly
every write for seven weeks: the dashboard kept showing Aug 9 as the latest
day, and its activation query (a cohort join over raw events, ~55 s per
week) stalled D1's single query lane until every panel failed.

Root cause was the client: usage counters were aggregated per process and
appended on exit, so every `serve` launch, CLI command and prompt hook
uploaded its own `count: 1` line - 30.3M of the 30.8M event rows.

- Client: merge count lines per (day, kind, name, client) on append,
  stale-claim recovery (which also bypassed the size cap) and send.
- Ingest: usage counters ADD into usage_daily, one row per machine x day x
  tool (migration 0003), so storage no longer depends on upload frequency.
- Rollup: reads usage from usage_daily, folds legacy usage rows out of
  events in 50k-row transactions, maintains machine_first_seen.first_index_day
  (0002), catches up on days a failed run missed, purges usage_daily with
  retention. 0004 drops the now-unused events_machine_day index.
- Dashboard: ranges end today; rolled-up series return null past the
  rollup's last day instead of a phantom zero; no panel reads raw events
  (activation reads first_index_day); min/max lookups stay index-only;
  one retry on transient 5xx; banners when ingest stalls or the rollup
  falls behind.
- scripts/backfill-rollup.sh re-runs the rollup over a day range.

Already applied in production: DB 10 GB -> 467 MB, all usage counts kept,
rollups current through 2026-10-02.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 05:58:44 +00:00
..