\set after -v overwrote the mode, window and followup psql was given, so
-v mode=v2 still read the interleave arms; the defaults now apply only where
no variable was given. Block 2b probes callrecord once per page on its
(org, email, created_at) index, as 2c does, instead of hashing the window's
calls against the jobs table, which timed out on the replica over one day.
decide's last rule read "nothing kept" as "not a task". Most such queries are
tasks on a platform the catalog lists, for a job it does not: posting to
Threads where only reading it is listed, a Telegram bot, an options quote,
a sound effect from text. The judge names the platform with confidence and
keeps nothing, and the answer was "try describing the data you want" (a
person) or the lexical page (an agent): for "publish post to Threads" the
Threads search endpoints, which the agent then re-queried past, and the miss
list never saw the gap.
A confident platform Choice with nothing kept is now a gap (rule 8; "not a
task" is rule 9), so the empty page says so, the agent's hint names the
platform ("the catalog has Threads but no tool for ..."), and the search is
recorded as a gap to add. On the agent eval set: 29 of 33 "not a task"
answers become gaps, job-hit unchanged, false-none 6% -> 7% (one of them a
lexical page that did carry the endpoint called).
Fragments: find.md (the rules), search-experiment.md.
CodeQL js/xss-through-dom: the result line was an HTML string assembled from
chapter titles and data-ch values. It is now built with createElement and
textContent, and chapter ids are checked against a fixed shape. The mobile
table of contents is cloned from the desktop one instead of copied as HTML.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bundled pages and llms.txt can now mark hosted-only links with
<!--hosted-->...<!--/hosted-->; _strip_hosted unwraps them on the hosted
deployment and removes them elsewhere, as the homepage already did. The
playbook's inbound links use it, and a test checks a self-hosted registry
serves none of them. Summaries (meta, structured data, blog card, llms.txt,
Related cards) now say the data steps were tested on recorded runs rather
than every chapter. The diagnostic also marks chapters in the mobile menu.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each chapter ends with a short 'Go deeper' line to the job page, provider
page, bench or workflow it draws on. The playbook is linked from the
homepage footer (hosted-only block), the /blog launches, the /use-cases and
/workflows hubs, and the Related blocks on /people-search and
/leads-signals; the page test now checks those inbound links.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rebuilds /gtm-engineering as a full playbook: a diagnostic that turns the
problems practitioners raise into symptoms that point at chapters, then 16
chapters in six parts (define, build the list, timing, reach out, inbound,
operate) and four appendices, with a sticky table of contents. Each chapter
has a symptom, the play, a prompt to copy, a visual from a recorded run with
its bill, the rule, and what treg does and does not cover; chapters with no
run behind them are marked as method.
Headings answer the questions answer engines are asked, and the FAQPage
JSON-LD answers the same ones. /gtm-engineering.md serves the playbook as
Markdown (noindex, rel=alternate, linked from llms.txt); a test fails if a
chapter heading on the page is missing from it.
Docs: docs/context/interface/seo.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first version led with our own proof. This one is built from what GTM
practitioners raise on Reddit and LinkedIn: each of seven chapters is one of
those problems in their words, the technique that fixes it, the rule, a
visual from a recorded run, and a plain line on what treg does and does not
cover, plus a chapter on the parts treg does not do.
It follows the /people-search design: the rotating agent in the H1 (the
served HTML names Claude Code), a self-playing demo that replays the
lead-list run and is captioned as abbreviated and illustrative, and
reveal-on-scroll visuals. Adds a gtm_chapter_view event per chapter.
Shorter title that fits a search result.
Docs: docs/context/interface/seo.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A hand-written hub framed as GTM engineering with Claude Code. It leads with a
recorded workflow run and the work-email bench, links the seven existing job
pages (no new routes behind it), lists the six workflows with their receipts,
and a dated table of GitHub skills next to the catalog job that covers each
one's data step, stating that those skills have not yet been run with treg.
Hand-written rather than _page()-rendered so it loads sitetrack.js and its
reading can be measured. Value events fire only on real outcomes (a clipboard
write that succeeded) and are flushed once PostHog initialises. Hosted-only
like /workflows: 404 and absent from the sitemap on a self-hosted registry.
Linked from the _page() footer, the launch-page footers, the use-case
landers' next steps and /resources.
Docs: docs/context/interface/seo.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MCP catalog_search answered from the lexical ranker with the v1 judge's say
over it. The judge recovered a quarter of the searches the lexical gate
admitted nothing for and barely moved the rest: its page was the lexical
candidates filtered and reordered, with no vendor the words did not reach, and
an empty page told the agent to try other words, so half of all searches were
followed by another search. The engine behind /catalog/find (recall by job and
by meaning, one judge request, every vendor of a fitting job, a verdict that
can say the catalog lacks it) now answers agents in a new experiment mode.
- search_experiment gains mode v2: arm_for deals v2 as the majority arm, the
same two holdouts keep the pure lexical and the pure v1 judged page, so v2
is read against what it replaces. Arms are dealt by team and email where the
search resolved them (an OAuth token rotates hourly), else by token.
- catalog_find: decide takes the verdict for its rule 8 (a person gets none,
an agent keyword: its input always means something), an abstain keeps the
judge's reason, expand returns its groups (expand_groups) and answer_v2
splits into judge_and_decide and lay_out, so a holdout that only records
lays out nothing. store gains routed_discovery_on and routed_parent, the
one reader of each where four were.
- application/catalog_search lays the answer out for an agent (agent_page):
the best limit // 2 jobs, rows dealt round-robin and laid out job by job,
members the judge rated on their own leading their job, a routed parent
leading a strong or closest job's group, a listed hub tool joining its job
with no lexical gate, and per job the vendors the page left out.
- The verdicts an agent sees: strong, closest, name, and none only for a
catalog gap (an empty page that says so, with catalog_request and no near
misses). Not-a-task, an abstaining judge, a failure or a caller past the new
per-caller cap (search_judge_max_per_caller_hour, the only bound on what an
unmetered search can spend; it guards every dealt caller before any judge)
serve the lexical page under keyword.
- The response adds verdict, reason, jobs, and per row job and more_providers;
score is null on a judged answer (no probability reaches an agent); the hint
says only what to do next. Both MCP surfaces answer alike.
- SearchLog: every served row records its job in every mode, so the report
credits a call to any vendor of a job the page showed (a v2 hint sends the
agent there); v2 rows carry find's readings and the verdict; the lexical
holdout still has v2 judged and recorded, the counterfactual on a false
none. The report gains conversion by job per arm and verdict, calls after a
none, and the re-query rate per verdict.
- scripts/search_agent_bench.py scores the v2 answer offline against searches
agents made and what they called next (job-hit, hit@limit, false-none by
arm), beside the pages the log served, on find_bench's harness (the judge
and the query vectors cached on disk, the card vectors warmed once).
- find_index: stored vectors decode as arrays, the index builds off the loop.
The HTTP route and the CLI still answer from the lexical ranker; they follow
once the route's hub read holds no session through a judge call and the CLI
sends its token.
Fragments: search-experiment.md, find.md, catalog.md, mcp-oauth.md; llms.txt
and skill.md (and the generated SKILL.md copies) say what the verdict means.
A v2 answer's verdict was only implied by its SearchLog row (an empty shown for
none, owner judged for both strong and closest), and the experiment report
stratifies by it. Migration 0056 adds searchlog.verdict, the reason after a
colon (none:gap), written by v2 finds.
scripts/search_experiment_report.sql: the interleaving-credit block never ran
on Postgres (round(double precision, int) does not exist; cast to numeric), the
mode its arms are read from is a psql variable, and latency percentiles leave
out answers served from the judge's in-process cache (0 ms).
Fragments: find.md, data-model.md, search-experiment.md.
The semantic channel's vectors were built by the first v2 find in each process,
so after every deploy each worker answered lexically (embed_error not_ready)
until a find happened to arrive and the build finished. The lifespan now starts
the build as a background task on every role (parsing the catalog off the event
loop, building on the one cached catalog object), cancelled with the other
workers on shutdown; the first find still starts it where that did not finish.
Fragments: find.md (semantic channel), composition.md.
The orchestration behind MCP catalog_search (band, evidence rerank, hub merge,
routed groups, the discovery experiment and the records) lived in mcp.py; the
layering table says a router holds no query orchestration. It now lives in
application/catalog_search.py and the MCP layer only resolves who is asking and
shapes the rows. No behaviour change: both MCP surfaces answer byte-identically
over a fixed query set with routed discovery on and off (verified locally with a
snapshot, not committed because catalog ids change weekly).
Fragment: search-experiment.md names the use case.
A sweep reading below EMPTY_BELOW refuses the provider's shared-key calls, so
three collectors could refuse calls the account still serves:
- fiber_ai read only the first credit pool, a spent trial pool, and reported 0
while a paid pool held credits; it now sums every pool.
- spyfu reported a negative balance once the monthly allowance was spent, but
further units bill as overage; a spent allowance is now informational.
- getleadsio raised on Unlimited plans, which answer fair-use windows instead
of credits_remaining; it now reports the tightest window.
A new test fails any platform-key slot that is on neither BALANCE_ROUTES nor
NO_BALANCE_API. It found seven undecided slots and a NO_BALANCE_API entry
spelled google-ai that never matched its google_ai slot. AnyAPI, cloro, reAPI
and PiAPI get collectors from their documented free balance routes; MiniMax,
OpenRouter and Replicate are recorded as having none usable with our key.
- Back into the first entry of a page opened at /catalog (or
/catalog/<slug>) showed "Your own tools": that entry has no history
state and no hash, and popstate fell back to the tools view. It now
resolves the view from the path, as the first load does. A bug on main
that the find flow walks into.
- A result opened from the Catalog page landed on its platform's shelf
with the shelf's box filled with the row's name, which the always-ask
box then filtered and searched by. It now opens the job: its
comparison once the shelf has loaded (replacing the shelf's history
entry, so Back returns to the answer), else the tool in the drawer; the
shelf's box is never prefilled. openPlatform returns its load.
With the uniqueness rule, "instagra" matched Instagram and Meta Ads
(whose label mentions Instagram) and so named no platform, falling to
the Instagram provider page. A query that starts exactly one platform's
own name without being a whole word of any matched label is that
platform; a whole word several platforms share ("video", "search")
still names none.
_stream_v2 awaited the query embedding (up to its timeout) before its
first event, so the page's reading animation started late. The lexical
recall now goes out first; the embedding, the fused recall and, when the
meaning changed the units, a second candidates event follow, then the
judged event. The pages already replace the candidate list on each
candidates event.
children_hidden counted the endpoints cut from a job under the strong
cut, but the pages say "N of M providers": a provider with two rows
made the total disagree with the vendor count the judge was shown. A
folded job now shows the first row of each of its first five providers,
and children_hidden counts the job's providers not on the page.
In shadow mode v1 and v2 both wrote a SearchMiss for the same empty
query, and searchmiss had no column to tell them apart, so every
double miss counted twice in the demand report. Only the served engine
now files a miss (v2's empty answers stay visible in its SearchLog
row), and searchmiss gains `engine` in migration 0054, which no
deployment has applied yet.
With no units the judge sent nothing and returned no extra answers, so
decide fell through to not_task: an out-of-catalog word like "weather"
hid the request link and filed its miss as not_task. The judge now asks
the extra questions alone when there are no candidates, and v2 passes
its name Noul and platform Choice, so an empty recall can be a gap. v1
still sends no request for an empty recall.
The bench records each case's embedding error, so a query that ran
without the semantic channel is visible in the run file.
- A platform is named by its whole name as a token sequence (slug words
and the label's short form, longest first), not by a stem of its
hyphenated slug, which no query token could equal: "tiktok ads library"
named TikTok, and "meta ads library" and "search console clicks" named
nothing, so multi-word platforms never got the double weight or the
reserved seats.
- The name table no longer lets a short or shared word stand for a name:
a platform or provider prefix needs four letters and must match exactly
one of them. "video", "search", "ads", "ai" matched several platforms
and rule 3 replaced a judged answer with a platform list.
- Fusion admits the semantic channel's most similar units whatever the
sign of the similarity, so a query with no word on any card fills its
seats by meaning; only the lexical channel requires a hit.
A shelf's find reads that shelf's units only, so it cannot know whether a
tool exists elsewhere, yet an empty one showed the catalog-gap copy and
offered "Request it". On a shelf an empty answer (none, or closest with no
rows) is now one line, "Nothing in <platform> for ...", with "Search all
tools", which moves to the Catalog page with the box prefilled and asks
the same words unscoped. The closest page on a shelf offers the same
instead of a request. "Request it" stays on unscoped answers.
On the server a shelf's none is reason `scope`, never `gap`, so its
SearchMiss row does not claim a catalog gap.
- find_recall: the tokenizer cut is store's; platform names, job counts
and each platform's providers are built with the index instead of per
query (a shelf's name lookup scanned every unit per provider); the
index no longer caches whichever provider-display function its first
caller passed - name_of takes it; Candidate drops an unread score.
- catalog_find: v1 and v2 name pages share one jobs-first order, and v2
keeps name_of's platform order instead of re-sorting it; the label cut
is find_recall's; the two name rules are one branch; the shadow drains
the v2 stream instead of repeating it, and the log reuses the reached
endpoints the stream built.
- judge: one question builder for both unit kinds, one cache key.
- object_store: named objects share the put and get bodies, and the name
shape is generic, the find-vectors prefix owned by find_index.
- find_index: a typed store, one build path for the background and the
bench, cards hashed once.
- find.js: engine and platform are read from the event for analytics,
not stored.
The product-name layer admitted "tts", "t2v" and "i2v": model names
share them, and names elsewhere rarely do. They are keys of aliases.yaml -
ways of saying a job - so a bare "tts" got a two-row product page instead
of the voice-generation jobs. Alias keys are now left out of the layer.
Now that an auto answer sits above the still-filtered shelves instead of
replacing them, the gates on it only withheld good answers: a bare name
being typed, a model or platform word, a short job word each get a v2
answer worth showing. One rule remains: a 700 ms pause on two characters
or more asks. The name filter keeps filtering while its answer shows
above; Enter still asks for the full answer.
The case for an auto answer above the shelves failed on a reviewer's
machine and passed locally under every load tried. The spec typed as soon
as the first shelf painted, while the session check and the connections
request could still be in flight, and typed key by key, so every partial
state passed through the gate. It now waits for the network to settle,
enters the sentence in one input event, checks the box holds it, and
waits for the /catalog/find request with its own message before
checking the drawn answer. The mocked row carries a real cost shape.