Files
SToneX 0fd8c70c67 feat(search): serve the job-first answer to agents (search_experiment v2)
MCP catalog_search answered from the lexical ranker with the v1 judge's say
over it. The judge recovered a quarter of the searches the lexical gate
admitted nothing for and barely moved the rest: its page was the lexical
candidates filtered and reordered, with no vendor the words did not reach, and
an empty page told the agent to try other words, so half of all searches were
followed by another search. The engine behind /catalog/find (recall by job and
by meaning, one judge request, every vendor of a fitting job, a verdict that
can say the catalog lacks it) now answers agents in a new experiment mode.

- search_experiment gains mode v2: arm_for deals v2 as the majority arm, the
  same two holdouts keep the pure lexical and the pure v1 judged page, so v2
  is read against what it replaces. Arms are dealt by team and email where the
  search resolved them (an OAuth token rotates hourly), else by token.
- catalog_find: decide takes the verdict for its rule 8 (a person gets none,
  an agent keyword: its input always means something), an abstain keeps the
  judge's reason, expand returns its groups (expand_groups) and answer_v2
  splits into judge_and_decide and lay_out, so a holdout that only records
  lays out nothing. store gains routed_discovery_on and routed_parent, the
  one reader of each where four were.
- application/catalog_search lays the answer out for an agent (agent_page):
  the best limit // 2 jobs, rows dealt round-robin and laid out job by job,
  members the judge rated on their own leading their job, a routed parent
  leading a strong or closest job's group, a listed hub tool joining its job
  with no lexical gate, and per job the vendors the page left out.
- The verdicts an agent sees: strong, closest, name, and none only for a
  catalog gap (an empty page that says so, with catalog_request and no near
  misses). Not-a-task, an abstaining judge, a failure or a caller past the new
  per-caller cap (search_judge_max_per_caller_hour, the only bound on what an
  unmetered search can spend; it guards every dealt caller before any judge)
  serve the lexical page under keyword.
- The response adds verdict, reason, jobs, and per row job and more_providers;
  score is null on a judged answer (no probability reaches an agent); the hint
  says only what to do next. Both MCP surfaces answer alike.
- SearchLog: every served row records its job in every mode, so the report
  credits a call to any vendor of a job the page showed (a v2 hint sends the
  agent there); v2 rows carry find's readings and the verdict; the lexical
  holdout still has v2 judged and recorded, the counterfactual on a false
  none. The report gains conversion by job per arm and verdict, calls after a
  none, and the re-query rate per verdict.
- scripts/search_agent_bench.py scores the v2 answer offline against searches
  agents made and what they called next (job-hit, hit@limit, false-none by
  arm), beside the pages the log served, on find_bench's harness (the judge
  and the query vectors cached on disk, the card vectors warmed once).
- find_index: stored vectors decode as arrays, the index builds off the loop.

The HTTP route and the CLI still answer from the lexical ranker; they follow
once the route's hub read holds no session through a judge call and the CLI
sends its token.

Fragments: search-experiment.md, find.md, catalog.md, mcp-oauth.md; llms.txt
and skill.md (and the generated SKILL.md copies) say what the verdict means.
2026-10-01 20:19:52 +08:00
..
2026-09-26 09:30:58 +08:00

treg — Codex / ChatGPT plugin

A distribution wrapper, not a second product. It ships the same skill that treg skill bootstrap already installs into ~/.codex/skills/, packaged so people find treg by searching the plugin directory that ChatGPT and Codex share.

One of five shop windows. The Claude Code plugin lives at the repo root (.claude-plugin/ + skills/treg/) and declares no connector in its manifest, so it installs with no token and nothing about it waits on a review queue — its skill wires up the CLI and the MCP tools at first run instead. Cursor is the same bootstrap in Cursor's own layout (plugins/treg/), and DeepSeek Harness is an npm bundle (package.json + dsh/) whose MCP row is disabled until there is a token. All of them are rendered by the same scripts/build_plugin.py from the same source; they differ only in the prepended bootstrap. See docs/CLAUDE-PLUGIN.md, docs/DSH-PLUGIN.md and docs/MINIMAX-PLUGIN.md (skills-only, from plugins/minimax/).

plugin/
├── .codex-plugin/plugin.json     the manifest + the listing copy
├── skills/treg/        GENERATED — do not edit by hand
└── assets/                       icon + logo (▚, black & white)

The assets

logo.png (1024×1024) and icon.png (512×512) are the ▚ mark in pure black and white — white quadrants on a black rounded square. They are rendered from the geometry in assets/brand/twitter/avatar-dark.svg, scaled rather than upscaled, so re-rendering at any size is exact:

viewBox 512  ·  outer rx 112  ·  quadrants 140.5² at (111,111) and (260.5,260.5), rx 20

icon.svg is deliberately the filled variant, not the transparent mark-white.svg: a white mark on transparency vanishes on a light background, and composerIcon renders in a host UI whose backdrop we do not control.

Note interface.brandColor is still clay #e0703f — the product colour on treg.to. That is an accent beside a monochrome mark, not a conflict, but change both together if the brand moves.

The skill is generated

src/treg/web/skill.md is the one source. It is served at /skill.md, written into every agent by treg skill bootstrap, and rendered into this plugin by:

python3 scripts/build_plugin.py            # regenerate BOTH plugins
python3 scripts/build_plugin.py --check    # fail if either is stale (also a test)

Two things differ from the served copy, both because a plugin arrives where the server does not:

  • {BASE} is baked to the public deployment. The server substitutes that placeholder per request; nothing substitutes it inside an installed plugin, so a raw copy would ship the literal.
  • A bootstrap section is prepended. Every other install path implies the CLI already exists — treg skill bootstrap only runs because treg is installed. This is the one path where the skill can land on a machine with no treg at all, and without it a first run ends in command not found.

Never edit skills/treg/SKILL.md. Change src/treg/web/skill.md and regenerate; tests/test_plugin.py fails if the two disagree.

Testing it locally (verified with codex-cli 0.145.0)

No desktop app required — the Codex CLI installs and loads plugins itself.

The marketplace manifest is a user-side file and is deliberately not part of this plugin. It lives at ~/.agents/plugins/marketplace.json, and its path is resolved relative to $HOME, not to the manifest and not as an absolute path:

{
  "name": "superdesign-local",
  "interface": { "displayName": "superdesign (local dev)" },
  "plugins": [
    {
      "name": "treg",
      "source": { "source": "local", "path": "./devs/superdesign/tools-registry-oss/plugin" },
      "policy": { "installation": "AVAILABLE", "authentication": "ON_INSTALL" },
      "category": "Developer Tools"
    }
  ]
}

Then:

codex plugin list                   # treg@superdesign-local — not installed
codex plugin add treg@superdesign-local     # the @marketplace suffix is REQUIRED
codex plugin list                   # installed, enabled, 0.7.1

Installed copies land in ~/.codex/plugins/cache/$MARKETPLACE/$PLUGIN/$VERSION/. Confirm the skill actually loaded — it should appear as treg:tools-registry:

codex exec --skip-git-repo-check "List the names of every skill you have available."

codex plugin marketplace add ./plugin does not work: that command expects a marketplace root, and this directory is a plugin.

The listing copy is the product

category and capabilities are not guesses — they are what OpenAI's own shipped plugins use (github is Developer Tools + ["Interactive", "Write"]; gmail is Communication; openai-templates is Productivity). The published docs show neither list, so read a real installed manifest under ~/.codex/plugins/cache/ before inventing a value.

Before submitting through OpenAI's plugin portal, check the manifest version matches pyproject.toml (a test enforces this), and that the copy still describes what treg does: it compares providers and the agent chooses. treg does not route automatically and does not fail over — the landing page had to be corrected for that claim once already, and a directory listing is harder to correct than a web page. A test guards this too.