2617 Commits
Author SHA1 Message Date
SToneX 7a5a3111ea Merge pull request #775 from superdesigndev/fix/report-mode-variable
fix(scripts): the experiment report takes its variables from the command line
2026-10-01 22:23:58 +08:00
SToneX 2f89c7cf41 fix(scripts): the experiment report takes its variables from the command line
\set after -v overwrote the mode, window and followup psql was given, so
-v mode=v2 still read the interleave arms; the defaults now apply only where
no variable was given. Block 2b probes callrecord once per page on its
(org, email, created_at) index, as 2c does, instead of hashing the window's
calls against the jobs table, which timed out on the replica over one day.
2026-10-01 22:17:31 +08:00
SToneX 4ed47d4cad Merge pull request #774 from superdesigndev/fix/find-gap-on-a-known-platform
fix(find): a job missing from a platform the catalog has is a gap
2026-10-01 21:55:35 +08:00
SToneX a8800fcad9 fix(find): a job missing from a platform the catalog has is a gap
decide's last rule read "nothing kept" as "not a task". Most such queries are
tasks on a platform the catalog lists, for a job it does not: posting to
Threads where only reading it is listed, a Telegram bot, an options quote,
a sound effect from text. The judge names the platform with confidence and
keeps nothing, and the answer was "try describing the data you want" (a
person) or the lexical page (an agent): for "publish post to Threads" the
Threads search endpoints, which the agent then re-queried past, and the miss
list never saw the gap.

A confident platform Choice with nothing kept is now a gap (rule 8; "not a
task" is rule 9), so the empty page says so, the agent's hint names the
platform ("the catalog has Threads but no tool for ..."), and the search is
recorded as a gap to add. On the agent eval set: 29 of 33 "not a task"
answers become gaps, job-hit unchanged, false-none 6% -> 7% (one of them a
lexical page that did carry the endpoint called).

Fragments: find.md (the rules), search-experiment.md.
2026-10-01 21:48:11 +08:00
SToneX ce891a1b88 Merge pull request #773 from superdesigndev/feat/search-v2
feat(search): serve the job-first answer to agents (search_experiment v2)
2026-10-01 21:20:44 +08:00
Jason Zhou 71daa9706c Merge pull request #749 from superdesigndev/feat/gtm-engineering-hub
feat(web): the GTM engineering playbook at /gtm-engineering
2026-10-01 05:33:44 -07:00
Jason ZhouandClaude Opus 5.5 326542ed80 fix(web): build the playbook diagnostic's output from DOM nodes
CodeQL js/xss-through-dom: the result line was an HTML string assembled from
chapter titles and data-ch values. It is now built with createElement and
textContent, and chapter ids are checked against a fixed shape. The mobile
table of contents is cloned from the desktop one instead of copied as HTML.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:04 +10:00
Jason ZhouandClaude Opus 5.5 ceb695b6a9 test(web): assert the playbook's inbound links both ways, and that pages still answer
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:04 +10:00
Jason ZhouandClaude Opus 5.5 5196ade450 fix(web): keep GTM playbook links off self-hosted registries; honest summaries
Bundled pages and llms.txt can now mark hosted-only links with
<!--hosted-->...<!--/hosted-->; _strip_hosted unwraps them on the hosted
deployment and removes them elsewhere, as the homepage already did. The
playbook's inbound links use it, and a test checks a self-hosted registry
serves none of them. Summaries (meta, structured data, blog card, llms.txt,
Related cards) now say the data steps were tested on recorded runs rather
than every chapter. The diagnostic also marks chapters in the mobile menu.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:04 +10:00
Jason ZhouandClaude Opus 5.5 99409da4a0 feat(web): link the GTM playbook into its chapters' pages and back
Each chapter ends with a short 'Go deeper' line to the job page, provider
page, bench or workflow it draws on. The playbook is linked from the
homepage footer (hosted-only block), the /blog launches, the /use-cases and
/workflows hubs, and the Related blocks on /people-search and
/leads-signals; the page test now checks those inbound links.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:03 +10:00
Jason ZhouandClaude Opus 5.5 cfbb085420 feat(web): the GTM engineering playbook, with a Markdown twin
Rebuilds /gtm-engineering as a full playbook: a diagnostic that turns the
problems practitioners raise into symptoms that point at chapters, then 16
chapters in six parts (define, build the list, timing, reach out, inbound,
operate) and four appendices, with a sticky table of contents. Each chapter
has a symptom, the play, a prompt to copy, a visual from a recorded run with
its bill, the rule, and what treg does and does not cover; chapters with no
run behind them are marked as method.

Headings answer the questions answer engines are asked, and the FAQPage
JSON-LD answers the same ones. /gtm-engineering.md serves the playbook as
Markdown (noindex, rel=alternate, linked from llms.txt); a test fails if a
chapter heading on the page is missing from it.

Docs: docs/context/interface/seo.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:03 +10:00
Jason ZhouandClaude Opus 5.5 d3387f5495 feat(web): rewrite /gtm-engineering as a problem-first playbook
The first version led with our own proof. This one is built from what GTM
practitioners raise on Reddit and LinkedIn: each of seven chapters is one of
those problems in their words, the technique that fixes it, the rule, a
visual from a recorded run, and a plain line on what treg does and does not
cover, plus a chapter on the parts treg does not do.

It follows the /people-search design: the rotating agent in the H1 (the
served HTML names Claude Code), a self-playing demo that replays the
lead-list run and is captioned as abbreviated and illustrative, and
reveal-on-scroll visuals. Adds a gtm_chapter_view event per chapter.
Shorter title that fits a search result.

Docs: docs/context/interface/seo.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:03 +10:00
Jason ZhouandClaude Opus 5.5 1d75291f44 feat(web): /gtm-engineering hub for the GTM job pages
A hand-written hub framed as GTM engineering with Claude Code. It leads with a
recorded workflow run and the work-email bench, links the seven existing job
pages (no new routes behind it), lists the six workflows with their receipts,
and a dated table of GitHub skills next to the catalog job that covers each
one's data step, stating that those skills have not yet been run with treg.

Hand-written rather than _page()-rendered so it loads sitetrack.js and its
reading can be measured. Value events fire only on real outcomes (a clipboard
write that succeeded) and are flushed once PostHog initialises. Hosted-only
like /workflows: 404 and absent from the sitemap on a self-hosted registry.
Linked from the _page() footer, the launch-page footers, the use-case
landers' next steps and /resources.

Docs: docs/context/interface/seo.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 22:24:03 +10:00
SToneX 0fd8c70c67 feat(search): serve the job-first answer to agents (search_experiment v2)
MCP catalog_search answered from the lexical ranker with the v1 judge's say
over it. The judge recovered a quarter of the searches the lexical gate
admitted nothing for and barely moved the rest: its page was the lexical
candidates filtered and reordered, with no vendor the words did not reach, and
an empty page told the agent to try other words, so half of all searches were
followed by another search. The engine behind /catalog/find (recall by job and
by meaning, one judge request, every vendor of a fitting job, a verdict that
can say the catalog lacks it) now answers agents in a new experiment mode.

- search_experiment gains mode v2: arm_for deals v2 as the majority arm, the
  same two holdouts keep the pure lexical and the pure v1 judged page, so v2
  is read against what it replaces. Arms are dealt by team and email where the
  search resolved them (an OAuth token rotates hourly), else by token.
- catalog_find: decide takes the verdict for its rule 8 (a person gets none,
  an agent keyword: its input always means something), an abstain keeps the
  judge's reason, expand returns its groups (expand_groups) and answer_v2
  splits into judge_and_decide and lay_out, so a holdout that only records
  lays out nothing. store gains routed_discovery_on and routed_parent, the
  one reader of each where four were.
- application/catalog_search lays the answer out for an agent (agent_page):
  the best limit // 2 jobs, rows dealt round-robin and laid out job by job,
  members the judge rated on their own leading their job, a routed parent
  leading a strong or closest job's group, a listed hub tool joining its job
  with no lexical gate, and per job the vendors the page left out.
- The verdicts an agent sees: strong, closest, name, and none only for a
  catalog gap (an empty page that says so, with catalog_request and no near
  misses). Not-a-task, an abstaining judge, a failure or a caller past the new
  per-caller cap (search_judge_max_per_caller_hour, the only bound on what an
  unmetered search can spend; it guards every dealt caller before any judge)
  serve the lexical page under keyword.
- The response adds verdict, reason, jobs, and per row job and more_providers;
  score is null on a judged answer (no probability reaches an agent); the hint
  says only what to do next. Both MCP surfaces answer alike.
- SearchLog: every served row records its job in every mode, so the report
  credits a call to any vendor of a job the page showed (a v2 hint sends the
  agent there); v2 rows carry find's readings and the verdict; the lexical
  holdout still has v2 judged and recorded, the counterfactual on a false
  none. The report gains conversion by job per arm and verdict, calls after a
  none, and the re-query rate per verdict.
- scripts/search_agent_bench.py scores the v2 answer offline against searches
  agents made and what they called next (job-hit, hit@limit, false-none by
  arm), beside the pages the log served, on find_bench's harness (the judge
  and the query vectors cached on disk, the card vectors warmed once).
- find_index: stored vectors decode as arrays, the index builds off the loop.

The HTTP route and the CLI still answer from the lexical ranker; they follow
once the route's hub read holds no session through a judge call and the CLI
sends its token.

Fragments: search-experiment.md, find.md, catalog.md, mcp-oauth.md; llms.txt
and skill.md (and the generated SKILL.md copies) say what the verdict means.
2026-10-01 20:19:52 +08:00
SToneX 3cc48f2851 feat(find): record the verdict on SearchLog; the report runs on Postgres
A v2 answer's verdict was only implied by its SearchLog row (an empty shown for
none, owner judged for both strong and closest), and the experiment report
stratifies by it. Migration 0056 adds searchlog.verdict, the reason after a
colon (none:gap), written by v2 finds.

scripts/search_experiment_report.sql: the interleaving-credit block never ran
on Postgres (round(double precision, int) does not exist; cast to numeric), the
mode its arms are read from is a psql variable, and latency percentiles leave
out answers served from the judge's in-process cache (0 ms).

Fragments: find.md, data-model.md, search-experiment.md.
2026-10-01 15:57:36 +08:00
SToneX 19c2e26779 feat(find): build the card vectors at startup
The semantic channel's vectors were built by the first v2 find in each process,
so after every deploy each worker answered lexically (embed_error not_ready)
until a find happened to arrive and the build finished. The lifespan now starts
the build as a background task on every role (parsing the catalog off the event
loop, building on the one cached catalog object), cancelled with the other
workers on shutdown; the first find still starts it where that did not finish.

Fragments: find.md (semantic channel), composition.md.
2026-10-01 15:57:13 +08:00
SToneX 78b3155d02 refactor(search): move the MCP search page into application/catalog_search
The orchestration behind MCP catalog_search (band, evidence rerank, hub merge,
routed groups, the discovery experiment and the records) lived in mcp.py; the
layering table says a router holds no query orchestration. It now lives in
application/catalog_search.py and the MCP layer only resolves who is asking and
shapes the rows. No behaviour change: both MCP surfaces answer byte-identically
over a fixed query set with routed discovery on and off (verified locally with a
snapshot, not committed because catalog ids change weekly).

Fragment: search-experiment.md names the use case.
2026-10-01 15:34:01 +08:00
SToneX 4bb8003c87 Merge pull request #771 from superdesigndev/data/stats-only-adapters
feat(catalog): hit/miss adapters for apollo.people.search and crustdata.companies.identify
2026-10-01 11:56:07 +08:00
SToneX f73dec12d1 Merge pull request #770 from superdesigndev/feat/stats-only-adapters
feat(routing): adapters can opt out of routing with route: false
2026-10-01 11:55:56 +08:00
SToneX adda1f2648 feat(catalog): hit/miss adapters for apollo.people.search and crustdata.companies.identify 2026-10-01 11:49:55 +08:00
SToneX f4e6d0189d test(routing): slim the route: false test to a hand-built contract 2026-10-01 11:49:53 +08:00
SToneX 3085be66a3 feat(routing): adapters can opt out of routing with route: false 2026-10-01 11:44:45 +08:00
Taus cae7709870 Merge pull request #764 from superdesigndev/fix/balance-collectors
fix(capacity): give every platform key a balance decision
2026-10-01 04:30:07 +06:00
shehjad-dev 4cfdcc4b92 fix(capacity): give every platform key a balance decision
A sweep reading below EMPTY_BELOW refuses the provider's shared-key calls, so
three collectors could refuse calls the account still serves:

- fiber_ai read only the first credit pool, a spent trial pool, and reported 0
  while a paid pool held credits; it now sums every pool.
- spyfu reported a negative balance once the monthly allowance was spent, but
  further units bill as overage; a spent allowance is now informational.
- getleadsio raised on Unlimited plans, which answer fair-use windows instead
  of credits_remaining; it now reports the tightest window.

A new test fails any platform-key slot that is on neither BALANCE_ROUTES nor
NO_BALANCE_API. It found seven undecided slots and a NO_BALANCE_API entry
spelled google-ai that never matched its google_ai slot. AnyAPI, cloro, reAPI
and PiAPI get collectors from their documented free balance routes; MiniMax,
OpenRouter and Replicate are recorded as having none usable with our key.
2026-10-01 04:19:36 +06:00
Taus f4e0fbd3bd Merge pull request #763 from superdesigndev/codex/enrichlayer-provider
feat(catalog): add Enrichlayer BYOK and platform tools
2026-10-01 04:08:04 +06:00
shehjad-dev 1b7c227a27 feat(catalog): route Enrichlayer searches and profiles 2026-10-01 03:57:24 +06:00
shehjad-dev dbf6fb8cb9 test(call): isolate optional key writer in pool assertion 2026-10-01 03:27:45 +06:00
shehjad-dev 060fa3ba70 fix(catalog): bound and verify Enrichlayer pricing 2026-10-01 03:27:40 +06:00
shehjad-dev ca1873bf79 feat(catalog): add Enrichlayer BYOK and platform tools 2026-10-01 02:24:03 +06:00
Taus 5a0ffe26f2 Merge pull request #762 from superdesigndev/codex/search1api-live-fixes
fix(catalog): make Search1API live calls usable
2026-10-01 01:07:39 +06:00
shehjad-dev 06e7d17bff fix(catalog): make Search1API live calls usable 2026-10-01 00:59:02 +06:00
Taus cbb69f842a Merge pull request #761 from superdesigndev/codex/search1api-provider
feat(catalog): add Search1API web tools and routing
2026-10-01 00:29:14 +06:00
shehjad-dev 69bc45d2a5 fix(capacity): pace Search1API at general request limit 2026-10-01 00:22:16 +06:00
shehjad-dev badc46db8d fix(catalog): use Search1API mark and isolate its test key 2026-10-01 00:02:36 +06:00
shehjad-dev b4ffebf422 feat(catalog): route Search1API web search extract and map 2026-09-30 23:50:08 +06:00
shehjad-dev 2293afb736 feat(catalog): add Search1API provider 2026-09-30 23:50:08 +06:00
SToneX 5ee1ca4a38 Merge pull request #758 from superdesigndev/feat/find-v2
feat(find): job-first /catalog/find (V2)
2026-10-01 00:40:23 +08:00
SToneX 3e6349dd1e Merge branch 'main' into feat/find-v2
The find log migration moves to 0055: main's 0054 (callrecord org_id, id index) landed first.
2026-10-01 00:16:55 +08:00
SToneX 9cae6ce47e fix(dashboard): a find result opens its job, and Back returns to the Catalog page
- Back into the first entry of a page opened at /catalog (or
  /catalog/<slug>) showed "Your own tools": that entry has no history
  state and no hash, and popstate fell back to the tools view. It now
  resolves the view from the path, as the first load does. A bug on main
  that the find flow walks into.
- A result opened from the Catalog page landed on its platform's shelf
  with the shelf's box filled with the row's name, which the always-ask
  box then filtered and searched by. It now opens the job: its
  comparison once the shelf has loaded (replacing the shelf's history
  entry, so Back returns to the answer), else the tool in the drawer; the
  shelf's box is never prefilled. openPlatform returns its load.
2026-09-30 23:47:39 +08:00
SToneX 07ade3018d fix(find): a typed prefix of one platform's own name still names it
With the uniqueness rule, "instagra" matched Instagram and Meta Ads
(whose label mentions Instagram) and so named no platform, falling to
the Instagram provider page. A query that starts exactly one platform's
own name without being a whole word of any matched label is that
platform; a whole word several platforms share ("video", "search")
still names none.
2026-09-30 23:20:25 +08:00
SToneX 57aea67e1e fix(find): the first candidates event does not wait for the query's vector
_stream_v2 awaited the query embedding (up to its timeout) before its
first event, so the page's reading animation started late. The lexical
recall now goes out first; the embedding, the fused recall and, when the
meaning changed the units, a second candidates event follow, then the
judged event. The pages already replace the candidate list on each
candidates event.
2026-09-30 23:17:15 +08:00
SToneX 12f2c3eff2 fix(find): a folded job counts providers, not rows
children_hidden counted the endpoints cut from a job under the strong
cut, but the pages say "N of M providers": a provider with two rows
made the total disagree with the vendor count the judge was shown. A
folded job now shows the first row of each of its first five providers,
and children_hidden counts the job's providers not on the page.
2026-09-30 23:16:13 +08:00
SToneX abbea14e49 fix(find): shadow mode files one search miss, from the served engine
In shadow mode v1 and v2 both wrote a SearchMiss for the same empty
query, and searchmiss had no column to tell them apart, so every
double miss counted twice in the demand report. Only the served engine
now files a miss (v2's empty answers stay visible in its SearchLog
row), and searchmiss gains `engine` in migration 0054, which no
deployment has applied yet.
2026-09-30 23:15:13 +08:00
SToneX df292b1669 fix(find): an empty recall still asks whether it is a name or a gap
With no units the judge sent nothing and returned no extra answers, so
decide fell through to not_task: an out-of-catalog word like "weather"
hid the request link and filed its miss as not_task. The judge now asks
the extra questions alone when there are no candidates, and v2 passes
its name Noul and platform Choice, so an empty recall can be a gap. v1
still sends no request for an empty recall.

The bench records each case's embedding error, so a query that ran
without the semantic channel is visible in the run file.
2026-09-30 23:14:30 +08:00
SToneX 2929951ff3 fix(find): platforms named by whole phrase; a shared word names none
- A platform is named by its whole name as a token sequence (slug words
  and the label's short form, longest first), not by a stem of its
  hyphenated slug, which no query token could equal: "tiktok ads library"
  named TikTok, and "meta ads library" and "search console clicks" named
  nothing, so multi-word platforms never got the double weight or the
  reserved seats.
- The name table no longer lets a short or shared word stand for a name:
  a platform or provider prefix needs four letters and must match exactly
  one of them. "video", "search", "ads", "ai" matched several platforms
  and rule 3 replaced a judged answer with a platform list.
- Fusion admits the semantic channel's most similar units whatever the
  sign of the similarity, so a query with no word on any card fills its
  seats by meaning; only the lexical channel requires a hit.
2026-09-30 23:13:12 +08:00
SToneX 7037579115 feat(find): a shelf's empty answer offers the whole catalog, never a gap
A shelf's find reads that shelf's units only, so it cannot know whether a
tool exists elsewhere, yet an empty one showed the catalog-gap copy and
offered "Request it". On a shelf an empty answer (none, or closest with no
rows) is now one line, "Nothing in <platform> for ...", with "Search all
tools", which moves to the Catalog page with the box prefilled and asks
the same words unscoped. The closest page on a shelf offers the same
instead of a request. "Request it" stays on unscoped answers.

On the server a shelf's none is reason `scope`, never `gap`, so its
SearchMiss row does not claim a catalog gap.
2026-09-30 22:56:06 +08:00
SToneX f02072a711 refactor(find): one copy of each rule, per-query work moved to the index
- find_recall: the tokenizer cut is store's; platform names, job counts
  and each platform's providers are built with the index instead of per
  query (a shelf's name lookup scanned every unit per provider); the
  index no longer caches whichever provider-display function its first
  caller passed - name_of takes it; Candidate drops an unread score.
- catalog_find: v1 and v2 name pages share one jobs-first order, and v2
  keeps name_of's platform order instead of re-sorting it; the label cut
  is find_recall's; the two name rules are one branch; the shadow drains
  the v2 stream instead of repeating it, and the log reuses the reached
  endpoints the stream built.
- judge: one question builder for both unit kinds, one cache key.
- object_store: named objects share the put and get bodies, and the name
  shape is generic, the find-vectors prefix owned by find_index.
- find_index: a typed store, one build path for the background and the
  bench, cards hashed once.
- find.js: engine and platform are read from the event for analytics,
  not stored.
2026-09-30 22:37:24 +08:00
SToneX 4c1d59f7f4 fix(find): an alias is vocabulary, not a product name
The product-name layer admitted "tts", "t2v" and "i2v": model names
share them, and names elsewhere rarely do. They are keys of aliases.yaml -
ways of saying a job - so a bare "tts" got a two-row product page instead
of the voice-generation jobs. Alias keys are now left out of the layer.
2026-09-30 22:31:54 +08:00
SToneX 8711dd1dc5 feat(dashboard): a typing pause always asks the finder
Now that an auto answer sits above the still-filtered shelves instead of
replacing them, the gates on it only withheld good answers: a bare name
being typed, a model or platform word, a short job word each get a v2
answer worth showing. One rule remains: a 700 ms pause on two characters
or more asks. The name filter keeps filtering while its answer shows
above; Enter still asks for the full answer.
2026-09-30 22:31:14 +08:00
SToneX b4d8aefd62 test(dashboard): the auto-find spec types after boot and says which step failed
The case for an auto answer above the shelves failed on a reviewer's
machine and passed locally under every load tried. The spec typed as soon
as the first shelf painted, while the session check and the connections
request could still be in flight, and typed key by key, so every partial
state passed through the gate. It now waits for the network to settle,
enters the sentence in one input event, checks the box holds it, and
waits for the /catalog/find request with its own message before
checking the drawn answer. The mocked row carries a real cost shape.
2026-09-30 22:21:20 +08:00