The Catalog box asked the finder on every typing pause, and its answer
replaced the filtered shelves - even when the answer was empty. Now a
pause asks only when the box's name filter shows nothing (platforms on
the Catalog page, tools on a shelf) and the text is more than one word
under four letters: a person navigating by name keeps navigating, and
typing fragments no longer reach the judge.
An auto answer is a section above the still-filtered shelves; while it
reads, and when it is none or empty, it is one line. Enter, or the
suggestion row, still asks for the full answer, which unfilters the
shelves and lights where it landed, as before.
One transient upstream error in one batch failed the whole card-vector
build, so the semantic channel reported failed although a rerun
succeeded. A batch that fails with a timeout, 429, 5xx or a dropped
connection is now tried twice more with a short doubling backoff; a
refused key or a vector of the wrong size is not retried.
v2's recall gains its second channel. infra/embed.py is an
OpenAI-compatible /embeddings client that never raises and caches a
query's vector in-process by model and folded text. application/
find_index.py builds one card matrix per catalog in the background on
the first v2 find: vectors are read from the archive's object store
under find-vectors/<model>/<card sha256>, only missing cards are
embedded, and those are written back. Without a store they live in the
process; without the API the channel stays off and the build is retried
later. The judged event and the searchlog row record the query's
embedding time and error.
The object store gains named objects for these vectors only, each body
carrying its own card hash and size. numpy joins the server extra for the
matrix product. Settings find_embed_api_key (falls back to treg's
OpenRouter key on the OpenRouter URL), find_embed_model, find_embed_url,
find_embed_timeout_s.
The bench builds the vectors before scoring when a key is set, and its
judge cache key now includes the job question's wording.
A find_engine setting (v1 | v2 | shadow) puts a second find behind
/catalog/find. v2 recalls by JOB: a capability is one unit carrying every
vendor, a member whose own words fit better than its job's card is judged
on its own, and uncatalogued endpoints stay units. A whole-word lexical
channel (light stemming, aliases, a doubled platform weight, half-weight
name prefixes) max-pools jobs over members and fuses by reciprocal rank
with an optional semantic channel that nothing supplies yet.
One judge request reads the units (a job question for jobs), asks whether
the text is only a name and, off a shelf, which platform it needs (a Jev
Choice). Eight rules decide the verdict; a name table (platform, provider,
model names) answers bare names; a strong job lists every vendor in the
evidence rerank's order, a closer one its first five. An empty answer
says whether it is a catalog gap or not a job. shadow serves v1 and logs
both.
Migration 0054 records the engine and v2's readings on searchlog and the
reason on searchmiss. The dashboard shows the reason, where a fit came
from, and how many vendors a folded job has. /catalog/search, MCP search
and the search experiment are unchanged. New fragment find.md; the find
section of search-experiment.md points to it.
scripts/find_bench.py scores /catalog/find against labeled queries: a
recall tier with no API calls, and a judge tier that runs the whole answer
with judge answers cached on disk by (model, query, unit ids, questions).
Reports verdict accuracy by stratum, false-strong, false-none, top-1, MRR,
tokens, latency, job coverage (micro, macro, full) and a paired diff
against a baseline run. A label that no longer matches the catalog stops
the run.
tests/fixtures/find_bench.yaml holds 30 synthetic queries; CI runs the
recall tier on them. The find section of search-experiment.md describes
the bench.
Activity fetched one page of 100 calls and runs with no way to go further, and /runs read a
team's whole call history to find its local runs.
- GET /activity merges calls, server runs and local runs newest first and pages them by one
(created_at, source, id) cursor, so Load older never lands rows above ones already shown.
Local runs and owned polling are excluded from its calls before the limit.
- /calls, /runs and /activity share one row format (routers/activity.py).
- 0054 adds (org_id, id) for /calls and a partial (org_id, created_at, id) over local runs,
built concurrently.
Fragments: interface/api.md, interface/dashboard.md, architecture/data-model.md,
architecture/multi-tenancy.md
A local dev server and production look identical in a browser tab. On a local
server every HTML page's <title> now starts with "[dev] ", whichever route
served it: the dashboard, a server-rendered catalog page, or a static page.
_DevTitleMiddleware is a pure-ASGI wrapper registered only when
Settings.local_dev holds: a local sqlite database behind a loopback
public_url. It buffers HTML responses only, fixes Content-Length, and passes
everything else through untouched.
local_dev is now the one test for "this machine, not a deploy":
single_user_ok is single_user and local_dev, so the no-login guard and the
cosmetic marker cannot drift apart on the host list.
A shelf was one grid of equal cards, comparisons and single tools mixed. Long
job descriptions stretched whole rows, the Auto-route pill repeated on most
cards, and nothing told the eye where to land.
- Everything on a shelf is ranked by 30-day observed calls, whether one
provider serves it or several. The six most used are cards of one height:
a two-line title, who serves it with the price on the right, and the
provider count with auto-route as quiet text under it.
- The rest is one list (tool, provider, price) on DataTable: a compared job
shows its logo stack and opens its comparison, a single tool opens the
drawer. Account and setup shows its first six until asked for the rest.
- With no counts yet (a fresh server) the cards read Featured, not Most used.
- A compared job's card shows its short title (capability_titles) and the
full description on hover.
- The old shelf-only rules (the Auto-route pill and the card footer) are gone.
Catalog and Connections moved to the platform shelf's look; Your own tools
(Skills & tools, Secrets, Team resources) and Activity still used the old
full-width headers, underline tabs, bordered tables and a floating search box.
- One hero for the three Your own tools pages, the lede constant so the tabs
never move; each page's own note sits under the tabs.
- A filter row on every page: FilterBox (the search box Connections had,
now one component for three pages) beside a segmented switch or pickers.
- Tools and skills are cards shaped like a Connections card: a status pill
(API, Server, Local only, Skill) and Try it / Copy snippet in place of the
icon row. Secrets, team resources and the call feed stay lists, on
DataTable, which also gives them the phone card layout.
- Activity's feed reads when / caller / call / status / cost, the key and
runtime under the caller and the action and tags under the call; the
"Show all (6)" toggle becomes a Succeeded / All switch. Usage gets figure
cards, a 7 / 30 / 90 switch and sections under a rule.
- A .pl-card lifts and shows a pointer only when it is a link or button; a
panel or figure no longer has to undo that. Connections' credential cards,
which open nothing, stop lifting.
- The empty-state "load from your machine" block, duplicated on Tools and
Secrets, is one component.
A compared job's capability description is written for agents as much as for
people ("Check an email address you ALREADY have is real and deliverable, not
for finding one"). On a platform shelf it made cards three lines tall and hard
to scan, but shortening it would weaken agent search, which ranks on it.
capabilities.yaml gains an optional `capability_titles:` block (id -> short
title). domain_rows carries it on a merged row as `title`; the description
stays whole. A title naming no capability fails the catalog load, so a typo
cannot ship silently. The titles themselves follow as a catalog data change.
Preserve public page sources and portable maintenance tools while moving hosted records and pricing evidence out of the public tree. Require explicit external pricing evidence and remove hosted database access helpers.
Update the catalog, ads conversion, archive, super-admin, data-model, MCP OAuth, API, SEO and skill context fragments. Merge the companion private import before this change.
`MCPServer(...)` calls `logging.basicConfig(level="INFO")` by default, so the root logger ran at
INFO and httpx/httpx2 logged "HTTP Request: GET <full url>" for every provider call. Providers that
take their key in the query string (SerpAPI, ZeroBounce, Datagma, Twelve Data, ...) had the key in
the production log, next to the caller's search values.
The four HTTP client loggers (httpx, httpx2, httpcore, httpcore2) now log at WARNING. treg's own
lines keep INFO.
Tests: tests/test_log_redaction.py (fails without the fix).
On Linux the nav's words render wider than on macOS, and at 1081px with the
hub's button the bar ran out of room (CI's layout test caught it); 1251px had
as little slack. Each breakpoint moves up to leave at least 60px: the compact
bar from 1300px, two rows from 1120px. The test checks both sides of each.
- A group heading leaves its rule out of the template instead of a page
selector hiding it; the kind switch and the filter box share one surface
rule; the catalog's person mark is sized to the name's line, not held in by
a negative margin, and uses the hover token.
- Gone: the --pl-edge variable no status sets any more, the Connections
footer's rules (the catalog's is .cat-foot), two margins another rule
already decides, a nav gap with no icon to space, platConnected.
- After a connect, a key save or a disconnect, loadAll alone re-reads the
connections; each did it twice. provName reads the provider index.
No short word said it: Connected read as "only these work", Your account
needed a verb, and a verb made it long. The mark is now a small person in a
neutral circle beside the platform's name; its hover and accessible name say
which provider's calls use your account or key, unmetered, and that the rest
run on treg's key.
A green Connected on a few platforms read as "only these work", when every
platform works on treg's key and the mark means calls there use the team's own
credential. It now says Your account (an OAuth provider) or Your key, in a
neutral pill, and its hover names the provider and that those calls are
unmetered.
A status pill's tint already carries its tone, so the dot inside it goes. The
Your account / API key switch was a shorter, outlined, monospace control beside
a tall surfaced box; it now shares the box's height, surface and corners.
A status read as a line of bold coloured text, and a card needing a person
wore a coloured ring besides. Every status, the catalog's Connected included,
is now one soft pill in its tone; the card stays plain and its primary button
is the call to act. The Connections count in the nav goes: the page lists the
cards needing a person first.
Taking every rule out left the tab bar without its baseline and the main
sections (Connected, Add a connection, Tools) without an edge. Those rules are
back; a group inside one list (a catalog category, a provider category, a
quiet heading) keeps none, its count beside its title.
A shelf, a comparison and a provider's page each opened with a line of joined
fields ("107 tools · 12 compared across providers · 36 providers", "29
providers · free – $1.2", "AI generation · API key · https://…") that read as
generated and repeated what the section counts, the Auto-route card, the table
and the buttons already say. A head is the breadcrumb, the title and one line.
Geist Pixel is drawn for the title size Getting started uses; at fifty pixels
its steps showed. Page titles are 26px now, a comparison's 22px, with the
spacing tightened to match. The rule after each section heading and under the
catalog's tabs striped the page where space and the cards already part it:
the count sits beside the heading instead.
The count line repeated the section headings and the nav badge, the
three-line pitch said one thing worth keeping, and the terminal tip at the foot
was for the few who save keys from a shell. The head is the title and "Your
own accounts and keys. Calls with them are never metered."
The head carried a count line, the title, a two-line pitch and three buttons
above the box a visitor came to type in. It is the title, with the size folded
back into it, and the box. Your own keys have Connections in the nav, so the
Bring your own key button goes; asking for a tool and listing one move to a
line at the foot, and a search that finds nothing still offers Request a tool.
The dashboard fragment had grown to 1,300 lines narrating computed names,
class names and past layouts that the code and its comments already carry, so
it drifted and every frontend change had to rewrite it. It now keeps what no
single file shows: how the app is built and served, its views and routes, the
rules every change keeps (the two view whitelists, dialogs at the root, the
load tickets, storage, inline confirms), and the decisions behind Catalog,
Connections, the shelf and the top bar.
Its sources are the entry points and the src/ files only it covers; the rest
are documented by other fragments. The backdrop-filter build quirk moves to a
comment beside the rule it constrains.
With Connections the bar holds seven destinations, eight with the hub, and it
had grown an icon per destination, two social icons, an offer pill and a
six-decimal balance. It now reads in words: no nav icons, GitHub and Discord in
the account menu beside Follow on X, and the balance to the cent, never rounded
up, the exact figure on hover. The pill says Referral, since it leads to the
friend credit and the affiliate programme alike; the credit offer is its hover
title and accessible name. Below 1250px it keeps only its gift.
The layout test now covers the widths either side of each breakpoint.
A key saved under a provider's name is told from a connection by comparing the
two lists, and they were fetched at different times: after a disconnect the
removed connection's secret came back as a saved key, a team switch showed the
previous team's saved keys, and a slow /connections showed every pasted-key
connection as saved for a moment. loadConnections now fetches /secrets in the
same Promise.all and replaces both together.
- A failed Remove of a saved key reported to secretErr, which only Secrets
showed; Connections and a provider's page show it now.
- A grant with no catalog provider (an own-app OAuth connect) got a card whose
Manage opened a blank page; it stays on Secrets.
- Verify and connect opened the paste dialog and never used the saved key; it
is gone. A saved key's card gets Manage like any other.
A key saved as a secret named for its provider is now a connAccounts row drawn
by ConnectionCard like any connection, so a provider's page lists it and the
catalog's Connected mark and a provider card's Replace key count it.
- A connection reads "Needs a second credential" from the server's
needs_extra_credential, which clears once it is supplied; the note it read
before is always sent, so the card and the nav badge never cleared.
- The provider list is fetched once a session, not on every view; Secrets
fetches connections only when it has none; a Bring your own key jump scrolls
from one place.
- Rows carry pasted and the connected count, so cards stop re-deriving them;
connectLabel and renewConnection replace four hand-kept copies.
- The switch reuses .seg, the catalog's Connected mark reuses .cn-st, the focus
ring reuses .pl-card.on; dead .mk-card rules and top-bar rules the 1600px
breakpoint made redundant are gone.
"Sign in" read as signing in to treg. A provider you log in to now reads
"Your account", the catalog's own word for it, with a Connect account button;
one you paste a key for keeps its key's name and Add key. The switch reads All ·
Your account · API key.
Signing in with an account the team already holds and pasting a vendor key are
different errands, and the sign-in providers were scattered alphabetically
among eighty key vendors. Each category now lists them first, and a switch
beside the filter narrows to one kind, each option counting what it would show.
The section reads Add a connection: Add another meant nothing to someone with
no connection yet.
A provider this deployment holds no client credentials for showed a card whose
only word was "Unavailable here": a button that could never work. Add another
now lists the configured providers; an account already connected to one still
shows under Connected.
A shelf of forty cards, each on a full shadow, read as a pile of floating
tiles. A card now rests on the small shadow and lifts to the medium one on
hover; the search box drops to match, and a card that needs a person draws its
status in its edge instead.
providers under a second taxonomy, and the page named for connections showed
none. The Catalog (#catalog, /catalog) now answers what an agent can call; a
new Connections page answers whose credential it calls with.
Connections lists every connected account and key, the ones that need a person
first, each with its status and the one step that fixes it in place (reconnect,
replace the key, add the second credential, choose the account). A provider key
saved as a secret named for the provider is listed there too, since the
credential ladder uses it the same way, so the Secrets page now holds only the
keys the team's own tools use. Below, every provider is a one-line card with an
outlined Connect or Add key. Bring your own key, from the catalog or a tool,
lands there with the provider in view. The nav's Connections entry counts the
accounts that need a person.
The Catalog, Connections and a provider's page share the platform shelf's hero,
section rules and surface cards. A pasted-key provider's credential is named by
its own token_label, so Fish Audio and the other API-key providers no longer
read as "your own bot".
The bar centred the nav by giving both sides an equal share, however wide
their content: from about 1060px to 1500px the account strip was squeezed and
the Discord icon slid under the balance, and a wider nav ran under the
referral pill. The sides now never shrink below their content and split only
the leftover width, so the nav stays centred whenever it fits; below 1600px the
nav drops its icons and the referral keeps only its gift, which is what fits.
A layout test walks the one-row widths with the hub's button added.
* fix(call): close abandoned idempotency claims with a stored 410 instead of 409 for 24h
A served call whose done-marking failed (a DB pool timeout) left its claim pending, and every
retry of that key answered 409 idempotency_in_progress until the 24h window expired.
- A claim is a lease owned by its call: it carries call_ref from the claim, the owner renews it
every 5 minutes while running, and store, release and renewal are fenced on call_ref.
- A lease older than 15 minutes with no open hold and no pending async task under call_ref (and
its call_ref: children) is closed by compare-and-swap on owner and lease timestamp with a
stored terminal 410: idempotency_response_lost with the charge when the owner's ledger shows
one (GET /calls/{id}/result may still have the answer), else idempotency_outcome_unknown.
The key is never run again: a lapsed lease does not prove the owner stopped, and a second run
under the same key is the double charge idempotency exists to prevent. Legacy rows without a
call_ref keep answering 409 until they expire.
- Storing, releasing and renewing a claim retry once on a pool timeout, on a fresh session.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(db): index the idempotency sweep and slow the arena refresh
Both run on the 1 vCPU primary whose CPU pinned during the pool-saturation windows that
stranded idempotency claims and settlements:
- _claim_idempotent sweeps expired labels on every keyed call, but idempotentcall had only
membership_id, so the DELETE read every label the caller holds (tens of thousands for a
batch caller). (membership_id, expires_at) makes it a range over the expired rows (0052).
- The arena collector re-aggregated the whole 30-day window every 120 s: three multi-second
aggregates plus a retention DELETE over a multi-GB table, about a quarter of DB time.
A 30-day leaderboard refreshes every 30 minutes now.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>