On 2026-09-03 one admin browser tab took the site down for two hours. `/admin/archive/panel` polls
every 5 s with no in-flight guard and fanned a `/admin/archive/keys` request out per endpoint row;
6,882 of them arrived from that one tab and held a third of the single 15-slot pool continuously.
Render's edge answered 15,480 requests `503 treg_saturated`, ~10,300 of them real `/call/` traffic.
PostHog showed 894 because capture_fault's token bucket drops past 10/min per key.
Sizing would not have prevented it and neither would a semaphore. `audit.py` had the right idea
(`_MAX_CONCURRENT_WRITES = 4`, "never starve the request pool") but a semaphore bounds only the
module that remembers to take one — `archive.py`, written later, took none. #319 has since added a
third independent copy of that primitive. A pool bounds every module routed to it, which is the
difference this buys: the fourth background writer is safe without its author reading the first three.
`POOL_SPECS` in infra/db.py, three engines against one database:
api session_maker 5+10 every request handler
admin admin_session_maker 3+0 /admin/* only
background background_session_maker 8+0 audit, archive writes, ads worker,
the observation reader, the evidence sweep
Overflow is 0 on both minor pools — it is the escape hatch a bulkhead must not have. Two sizing
rules, both learned by getting them wrong in review first: background must be at least the sum of
the semaphores aimed at it (audit 4 + archive 4), or the two bounds fight and the loser waits
pool_timeout then drops its row; and a handler that opens a second session needs a spare slot, which
at admin=2 is how /admin/errors nesting the retention sweep inside itself ate the whole pool. That
sweep now runs on background, where a sweep belongs.
Statement timeouts are deliberately NOT here. They are a behavior change on every query, can only be
validated against production data volumes, and would land on reads already known to be far over any
sane bound. Splitting them out keeps "the bulkhead broke something" and "a timeout broke something"
distinguishable. TREG_DB_POOL_OVERRIDES patches POOL_SPECS from the environment, because pool sizing
can only be validated in production and a wrong number must be a dashboard edit, not a deploy.
Routing details worth review: `get_admin_session` is a separate dependency callable rather than a
flag because FastAPI caches dependencies per request by identity, so `require_superadmin` moves onto
it too. `archive.lookup` stays on the API pool — it runs inside a caller's /call/ and must not queue
behind the archive's own writes. `_touch_write` now takes the same semaphore `_store` does: it was
the one archive writer with no bound at all, and what it drops is the demand signal `prune_once`
reads, so a burst of served hits would have made those keys look prunable.
tests/test_db_pool_isolation.py classifies every module that opens a session, so a new one has to
declare its pool rather than defaulting into the API's. Writing it caught two holes in itself: the
first version was parametrized over a single file, and it missed `application/auth.py` reaching a
maker as a module attribute. Suite: 2528 passed, 5 skipped, five consecutive runs.
Root-cause fix for the 2026-08-15 outage. The #122 deploy's startup migration ran
`ALTER TABLE callrecord ADD COLUMN ...` — an ACCESS EXCLUSIVE lock on the table every live call
writes. With no lock_timeout the ALTER queued behind live traffic, every NEW query then queued
behind the waiting ALTER, both instances starved, the health check failed, and the shared Postgres
was left wedged for the instance that was healthy. An app restart could not fix it (the held slots
were server-side); only a database restart cleared it.
Two changes, each pinned by a test:
* **`init_db` sets `lock_timeout = 5s` (and `statement_timeout = 120s`) before migrations, on
postgres only.** A contended deploy now fails in seconds, cleanly: prod keeps serving the old
code, nothing wedges, and the deploy is retried at a quieter moment. The test asserts the
timeout is set BEFORE `_migrate_to_orgs` runs.
* **The pool shrinks from 20+40 to 5+10 per instance.** The pool is per INSTANCE and a rolling
deploy runs two against one database — 60 each meant 120 potential connections against a
basic-plan ceiling of ~100, so a deploy could starve Postgres with no bug anywhere. The test
computes double the configured bounds and requires them under 90.
ops/deploy.md records the operational half: a lock-timeout deploy failure is the mechanism WORKING
(retry quieter, do not raise the timeout), and a wedged database needs the POSTGRES resource
restarted, not the web service.
Why no test caught the original: SQLite has no pool and no lock queue, so this class is invisible
locally — also now written in deploy.md as the rule that a db.py change needs a Postgres-shaped
deploy plan.
2 new tests, 1512 pass.
Found verifying the 0.9.0 release on the live path, and my own report of it was wrong first.
**/meta already had `app_version`**, so "no version field" was inaccurate as I said it. But
`app_version` is a HASH OF index.html — it answers "has the dashboard bundle changed?", which is what
a long-lived tab compares to offer a refresh. It is not the release version, so the real gap stood:
after publishing 0.9.0 there was no way to confirm from outside WHICH VERSION was serving, only a
commit id.
`/meta` now returns both, because they answer different questions, and a test asserts they are not
the same value so nobody later "simplifies" one away.
`_treg_version` reads installed metadata directly rather than calling `cli.cli_version`, which does
the identical thing: importing `treg.cli` from an endpoint cost ~200ms on first call and pulled the
whole CLI into the server process for one string. Cached — 2ms first, 0.04ms for a thousand after —
and a measurement confirms `treg.cli` stays out of the server.
The matching runbook fix could not be committed: `.claude/` is gitignored, so it stays local. It
corrected two things — `/demo/sandbox` is POST-only (the runbook had releasers GET it, which returns
405 forever and reads as a broken endpoint), and the live-path check now compares `treg_version`
against what was just published, which is what would have made this release conclusive rather than
inferred.
1328 pass.
main was RED before this branch touched anything, and every PR was failing CI behind it. The repo
rename to superdesigndev/treg rebranded the landing page (185 occurrences of 'treg', none of the old
name), but test_dashboard_served_at_root still asserted 'tools-registry' in the markup.
Fixed here rather than left for a separate PR because nothing else can merge until CI is green, and
the change is one assertion.
* feat(web): first-run agent onboarding + Getting started revamp
New signups land on #connections (the default view for any signed-in
arrival with no deep link). The welcome modal grows two steps after
team creation: an agent picker (OpenClaw / Hermes / Claude.ai / Claude
Code / Codex + more) and the per-agent setup line. The old docked
"Getting started" stepper (onb.*) is removed entirely.
Getting started is rebuilt as numbered cards — ① Set up your <agent>
(setup line + the already-created API key, masked) and ② Try it out
(four copyable live-data prompts, each verified through the catalog,
plus OAuth connect chips for social/ads/site-SEO) — moved to the top
of the nav, above Catalog. Main content column is now centered;
"Your tools" renamed to "Your vault"; page titles reworded.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* feat(skill): rename the installed skill tools-registry -> treg
The skill installed into agents was named/foldered "tools-registry",
which matched neither the `treg` command nor the brand agents see —
confusing both agents and humans. Now: skill.md frontmatter name is
treg, `treg skill bootstrap` writes <agent>/skills/treg/ (and retires
a leftover tools-registry/ folder when it holds only our SKILL.md),
install.sh's fallback path and the Claude plugin folder follow suit.
The PyPI package name (tools-registry) is unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* feat(skill): sell the skill in its description, not describe it
The old description led with mechanics (catalog vs own-tools split).
Now it leads with when to reach for it — data access (SEO/SERP,
keyword volume, backlinks, social & trends, enrichment, ads,
scraping) and acting on connected accounts (posting, ad campaigns,
site SEO via OAuth) — and pitches treg as OpenRouter for agent
tools. Kept YAML-safe (no ": " inside the plain scalar).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* feat(agents): one setup line everywhere — llms.txt carries the flow
The ~30-line "Agent instruction" duplicated what llms.txt and the
skill already teach (catalog search, prices, scan/upload); the only
things a prompt can carry that a file cannot are the token, the team,
and the authorization to run. So the instruction everywhere is now
exactly: set up treg — {BASE}/llms.txt with token <T>, team <slug>
(welcome modal, Getting started, sidebar copy row, preview modal —
all one buildAgentPrompt).
llms.txt absorbs the rest: a "set up treg" agent flow (permission
once, work straight through, prove it with one relevant call), the
star-the-repo ask after first real data, and its money rules drop
the confirm-price-with-the-human ritual — prices are visible, calls
are fractions of a cent, just call.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* feat(catalog): Try-it is the primary CTA, with agent/CLI/API/manual tabs
On an endpoint row, "▶ Try it" is now the ink-fill primary action and
the old "Connect {provider}" is demoted to a secondary "Bring your own
key" (its real meaning — your key wins over treg's and isn't metered).
The Try-it drawer gains four tabs (AI Agent default): AI Agent shows a
one-line setup (team + token embedded here only, a run-now context) and
a ready "Use treg to call <id>" prompt; CLI shows install/login →
catalog get → the filled treg call; API shows the curl /call/ form with
the token header (+ X-Treg-Org in session mode); Manual is the existing
live test form.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* fix(catalog): OAuth endpoints keep Connect as primary; fix dead Manual Run
Which run action is primary now depends on the provider's auth_kind
(mkOauth): OAuth providers (act as your account, can't use treg's key)
keep Connect as the ink-fill primary with Try-it secondary; key/token
providers keep Try-it primary with "Bring your own key" secondary.
The Manual tab no longer strands a disabled Run hugging the tab bar
when the endpoint isn't callable from this org — it shows a banner
pointing at Connect / the AI Agent / CLI / API tabs instead, and a
"no parameters — just run it" hint for the param-less callable case.
Retitles the markup test to the new run-actions intent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
* feat(landing): 0% additional fee band after pay-for-results
Centered pixel-type headline plus a dark mono equation card that types
out "provider rate + $0.000 markup = your price" on scroll, with a
one-line note that treg earns on vendor volume pricing, not markup.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018UnVdcbzN8H4mXQaTkHdvD
* feat(web): billing — one-click fund cards + auto top-up toggle
Add funds is now four preset cards ($5/$10/$25/$50); clicking one goes
straight to Stripe checkout — no dropdown-then-button. Auto top-up is a
toggle: switching it on reveals the settings panel (amount, threshold,
the PSD2 consent mandate, confirm); the toggle alone can only disarm or
open the panel — arming still requires the recorded consent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D8t2aMqbfstGqStPLsoxDq
* feat(web): default landing is Getting started, not the catalog
New signups (after the welcome modal) and every signed-in arrival at
/app with no deep link or hash now land on #start (Getting started)
instead of #connections. welcomeFinish and both boot paths updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The dashboard's app div gained v-cloak (<div id="app" v-cloak>), which
broke the walking-skeleton's exact '<div id="app">' substring asserts.
Match 'id="app"' instead so attributes on the root div don't fail CI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Brings the public tree up to the tip of the private repo: the monologue-UI
landing page (web/landing.html), auth/email + CLI fixes, and web empty-state
work. Re-applies the sandbox.py demo-value sanitization. 520 tests pass.
Note (see OSS-PREP-NOTES.md): the new landing page pulls Google Fonts + a
brand-icon CDN in addition to Vue — vendor these before public launch.
A remote registry that turns a team's skills into shareable, callable tools:
any member's agent or a human calls a tool without owning its credentials —
a proxy injects the secret server-side.
This is the curated public tree (internal handoffs, plans, meeting notes,
and dev journal are kept in the private archive). Still WIP before going
public: see OSS-PREP-NOTES.md for the remaining genericization + the LICENSE,
CONTRIBUTING, AGENTS, and SECURITY items to finish.