v2's recall gains its second channel. infra/embed.py is an
OpenAI-compatible /embeddings client that never raises and caches a
query's vector in-process by model and folded text. application/
find_index.py builds one card matrix per catalog in the background on
the first v2 find: vectors are read from the archive's object store
under find-vectors/<model>/<card sha256>, only missing cards are
embedded, and those are written back. Without a store they live in the
process; without the API the channel stays off and the build is retried
later. The judged event and the searchlog row record the query's
embedding time and error.
The object store gains named objects for these vectors only, each body
carrying its own card hash and size. numpy joins the server extra for the
matrix product. Settings find_embed_api_key (falls back to treg's
OpenRouter key on the OpenRouter URL), find_embed_model, find_embed_url,
find_embed_timeout_s.
The bench builds the vectors before scoring when a key is set, and its
judge cache key now includes the job question's wording.
Brings main's 86 commits (dashboard boot and loading, legacy dashboard removal, test pruning,
overflow and routing fixes) together with the tool hub branch.
Conflicts, both sides kept unless noted:
- the legacy dashboard stays deleted, as on main;
- App.vue and the dashboard state: main's search page and loading states plus the hub pages;
- ci.yml: main's Postgres job, with the hub tests added to its list;
- dev-local.sh: main's server environment plus the hub flag passthrough;
- test_call_application_contract.py, test_marketplace_call.py: main's pruned files plus the
hub branch's sync `settle: usage` test.
Not conflicts: main and the hub branch fixed the same CompanyEnrich empty-page billing; main's
rule runs first, so the hub branch's copy and its test are dropped. The Listing-tab test reads
the Vue source instead of the deleted legacy page.
Brings main's 24 commits since the 2026-09-21 merge: the new Vue dashboard in frontend/ with
the account-based rollout (legacy frozen as src/treg/web/dashboard-legacy/), Olostep and
Keenable, routed web search, TrestleIQ, call_media and resources_list on MCP, provider
resources, the search log.
Conflicts, every one "both sides added": imports in call/service.py and routers/web.py; the
error-owner table in call/types.py; dev-local.sh keeps the TREG_HUB_ENABLED passthrough in
front of main's new SERVER_ENV; the legacy dashboard keeps both the Resources and the Hub
nav entries and view names; test_mcp.py lists main's two new tools and the hub's three;
test_marketplace_call.py keeps both new test blocks; .gitignore keeps both. MAP.md and
docs/context/README.md regenerated. tests/test_dashboard_markup.py removed, as on main.
The hub's six migrations move from 0041-0046 to 0044-0049 above main's 0043; a fresh database
upgrades to one head, 0049.
SQLModel 0.0.45 rejects naive datetime bind parameters and may return
aware datetimes from the DB, breaking three crons:
- treg-asynctasks-settle: ValueError on WHERE next_check_at <= :now
- treg-catalog-stats: TypeError on created_at >= since comparison
- treg-arena-insights: ValueError on INSERT scan_until bind
Short-term fix: pin sqlmodel>=0.0.22,<0.0.45 → lockfile uses 0.0.44
Long-term prep: NaiveUTC annotations in place for eventual upgrade
Changes:
- Pin sqlmodel>=0.0.22,<0.0.45 in server extra
- Import NaiveDatetime from pydantic, define NaiveUTC type alias
- Update all datetime field annotations in models.py to use NaiveUTC
- Add regression tests for all three affected crons
Fixes production crashes in settle, catalog-stats, and arena-insights.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Jason Zhou <JayZeeDesign@users.noreply.github.com>
main moved 210 commits ahead in five days (413 files: AnyAPI, Prospeo, Dropleads and other catalog
listings, routed-parent adaptation, the archive's own-key and repeat-pricing work, video task
settlement, org slug history). Only ONE file conflicted, the GENERATED MAP.md, which was rebuilt
rather than merged by hand. Nothing in the hub needed a code change this time.
Migration numbers collided again: main took 0039 and 0040, so the six hub revisions move
0039-0044 → 0041-0046 and chain onto main's 0040. Single head 0046; the chain applies clean on a
fresh sqlite with both hub switch columns present. The numbers named in architecture/hub.md follow,
and every file it names exists.
Verified after the merge: our phase 10 work survived intact — `scripts/build_plugin.py` still ships
the hub section unconditionally, and all five plugins still carry the hub (9 mentions each), the
"per-deployment switch" instruction for the off case, and no raw markers. main changed skill.md and
llms.txt; the plugin freshness check passes against the merged agent files.
Green: full suite 4393 passed, 8 skipped; hub + callmatrix + mcp 213; agent pages + seo 228;
surface snapshot + alembic parity 8; import-linter 14/14; plugin check OK. One failure in the
parallel run, test_prospeo_platform_email_settles_one_credit, which passes alone and whose whole
file passes alone (363); dev/hub touches no prospeo or marketplace file.
Closes#557
Root cause: PyPI's 0.19.1 was published before the `treg host` command
landed on main (c8d51c5e), but the skill served from the live server
documented `treg host`. Users installing via `curl | sh` got a CLI
without `host` and a skill that claimed it existed. Worse, because
`host` is also the name of a system DNS tool (/usr/bin/host), the
bare-word fallback (`treg claude` → `treg with claude`) intercepted
`treg host face.jpg` and ran `/usr/bin/host face.jpg` instead,
producing a confusing DNS error.
Changes:
- Bump version to 0.20.0 (0.19.1 is already on PyPI without `host`)
- Add test confirming `host` is a registered subcommand and won't fall
through to the system binary
- Add version note to skill.md (and all plugin copies): "Requires CLI
≥ 0.20.0; run `treg update` if `treg host` is unrecognised."
The existing `treg host` implementation is unchanged — it was already
on main, just never released.
Co-authored-by: Jason Zhou <JayZeeDesign@users.noreply.github.com>
main had moved 150 commits ahead. Nine files conflicted; every one keeps BOTH sides, except the
catalog search hint, which keeps the hub branch and calls main's new `_run_hint` helper for a
catalog row. The app serves the hub routes and main's new media routes; the worker keeps both its
hub and arena subcommands.
Migration numbers collided a second time: main took 0033-0038, so the six hub revisions move to
0039-0044 and chain onto main's 0038. Single head 0044; the chain applies clean on a fresh sqlite
with both hub switch columns present. The numbers named in architecture/hub.md follow.
Two real breakages the merge caused, both fixed here:
1. Money. main made the signup credit verified-only (`claim_signup_promo`), and `POST /users` is
legacy registration, which leaves the user UNVERIFIED. Every hub test buyer therefore started at
zero and each paid run answered 402. New shared helper `conftest.funded_user()` grants the team
1_000_000 micro; used at the 19 call sites that have to pay. The tests that deliberately exercise
the out-of-money path are untouched.
2. Identity. main added `ApiKey | None` to `Caller` for managed keys. The hub's scheduled check
builds a Caller directly and crashed with a TypeError. It now passes `api_key=None`: a scheduled
check is not an API-key call, it runs as the maker's membership.
Green on the merged tree: full suite 4070 passed, 8 skipped; agent pages + seo 228; import-linter
14/14; plugin check OK; alembic parity 4. `hub_enabled` stays False by default, and with the flag
absent every hub route answers 404 with a valid request, the agent files carry no hub text, and
catalog search returns no hub row.
Bring the team's work since 2026-09-09 into the hub branch: the review/feedback system, the arena
insights, the R2 body storage and cache admission for the archive, the shared key-value store, and
the SSRF CGNAT fix. 108 commits, 233 files.
The one real conflict was the migration numbers: both branches used 0026-0030 for different
changes. main's chain is 0026 (callreview) through 0032 (archive body storage); the five hub
migrations are renumbered to 0033-0037 and rechained onto 0032, so there is a single head 0037.
Everything else was additive and kept from both sides: the HubTool/HubRun models beside
CallReview/FeedbackHandling, the `hub` domain and the key-value store both in the AGENTS.md tables,
the hub CLI family beside `treg review`, the hub MCP verbs beside `review` in the tool sets, the
quickjs and obstore/redis dependencies together. mcp_feedback.py was removed on main; the deletion
is kept. The docs fragment's migration numbers were updated.
Full suite 3508 passed; agent pages + SEO 216; boundaries 14/14; the migration chain a single head
0037 on Postgres 16; the plugin staleness guard green.
The CLI now prints one stderr line for X-Treg-Hint: feedback beside the charge line, as it already
did for review invitations, and reads the unified header a 0.19 registry sends (the older
X-Treg-Review: requested still works against a pre-0.19 registry). Plugin manifests, the skill
copies and the lock file carry the same version.
Render Key Value (Redis protocol) at TREG_KV_URL, or a bounded in-process dictionary when unset.
One narrow method, take(key, limit, ttl_s): an atomic INCR + EXPIRE NX window, capped at 100 ms,
failing closed so a store that cannot answer is a spent budget rather than a flood. /admin/kv
(superadmin) reports configured/reachable; startup warns when the configured store is silent.
The first tenant is the review-invitation budget; TREG_REVIEW_BUDGET_PER_HOUR (default 5) is
declared here for it.
A script recipe's run.js now runs end to end (docs/HUB-DECISIONS.md round 2).
- treg/hub_sandbox.py, the CHILD: QuickJS (the `quickjs` package, hand-added to the server
extra and uv.lock; cp312 wheels for production, sdist locally) in a separate short-lived
process with no network, no file system, no require, no process, no timers. A JSON-lines
bridge over stdin/stdout: ctx.call -> parent -> result | refused; ctx.log; done | error.
64 MB engine heap; a memory bomb reads `memory`, an endless loop `timeout`, a throw `script`.
- application/hub/sandbox.py, the PARENT: runner.py's discipline - scrubbed env (never the
server's), private temp HOME, own process group, rlimits (CPU, fsize, no core, RLIMIT_AS on
Linux), a wall-clock kill of the whole group, cleanup on every exit. Enforces what the engine
cannot: the 20-call cap, the log caps, the output shape.
- runner.py, the script road: every ctx.call is the same child call the JSON road makes -
`uses` enforced per call (a URL, an unknown id, or a tool outside the list refuses and stops
the run), the run ceiling, {run}:s{n} holds, a catalog call as the CALLER, an own-tool call as
the MAKER. The returned object must carry output.fields (424 hub_output_invalid); a failed run
is 424 hub_run_failed {hub_script_failed, kind, message, trace, log, charged_micro}.
Two engine facts learned and documented: quickjs cannot call into Python while its own time
limit is set, so the wall clock is the parent's kill; ctx.call is synchronous underneath, so
Promise.all serializes in version one (the JSON road has the real parallelism).
Tests: 13 hostile sandbox cases (no road out, refusal stops the run, memory bomb, endless loop,
throw, missing export, non-object return, log cap, call cap, scrubbed env, parallel runs do not
mix) + 4 end-to-end script runs on the real call path with the four books. 72 hub tests; suite
2900; hub + serial list on real Postgres 16.
Security review of the sandbox is scheduled as its own pass before release (the owner's note).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add playwright for mobile viewport testing. Also add tests/screenshots/
to .gitignore as these are regenerated on demand.
Co-authored-by: Jason Zhou <JayZeeDesign@users.noreply.github.com>
PKCS#7 EnvelopedData Bleichenbacher oracle (GHSA, fixed in 50.0.0). treg never
touches PKCS#7 — Fernet + X.509 CA only — and unpinned pip/Render installs
already resolve 50.x; this aligns the dev lock. Targeted `uv lock
--upgrade-package cryptography`, no format churn (also corrects the lock's
stale project version 0.7.0 → 0.11.0). Full suite green against 50.0.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An agent inside ChatGPT or a Codex plugin has no terminal, and a skill that says "first install this
CLI" is where the visitor leaves. This is the other door: `/mcp`, streamable HTTP, five tools.
FIVE tools, not 2,600. The catalog stays DATA — catalog_search, catalog_get, call, balance,
my_tools. A tool per endpoint would bury the model's context in 2,600 schemas and make the catalog
unusable, which is the opposite of the point.
Everything touching money, tenancy or credentials goes back through treg's OWN HTTP API in-process
(httpx ASGITransport) rather than reaching into the internals a second time. `/call/` already
enforces the member ACL, deny rules, both daily caps, the balance reserve and the settle; a second
entrance that re-implemented any of that is how one copy quietly stops being enforced. Only the
public catalog is read directly, because it is already in memory and needs no identity.
Three traps found by building it rather than reading about it:
* `app.mount()` does NOT run a mounted app's lifespan, and the transport builds its task group
there — every request 500s with "Task group is not initialized". api.py now COMPOSES the two
lifespans. It caught me twice: once in the spike, once in my own test fixture.
* The SDK ships DNS-rebinding protection ON with an EMPTY allow-list, which 421s EVERYTHING. The
deploy would have looked healthy until the first tool call. `_allowed_hosts()` builds the list
from this deployment's `public_url` plus loopback, with TREG_MCP_ALLOWED_HOSTS to extend it, and
a test asserts both directions — the real host works, an unknown one is refused.
* `StreamableHTTPSessionManager.run()` may be called once per instance, so `build_mcp_app()` is a
factory; tests get a clean transport each while still exercising the real tools and the real
enforcement path.
Proven from Codex against the dev box, not just in tests. Asked for work-email options with no API
key: it searched, read two endpoints, and compared hunter ($0.0245/success), leadmagic ($0.025) and
thecompaniesapi ($0.0019) — noting the third returns a PATTERN rather than a confirmed address, then
declining to spend because I said not to. balance reported $1.00 (the new-org grant) and my_tools 0.
`call` went through the full path and returned the correct refusal with `whose_error: treg` — the dev
box holds no provider keys, so a SUCCESSFUL metered call stays unproven until this reaches prod,
where they live.
Import is optional and mounted last: a deploy that has not picked up the `[server]` extra's `mcp`
dependency still serves everything else. The lock was regenerated with a modern uv (0.12) rather than
this Mac's 0.5.2 — format preserved at revision 3, additions only, and it restored `provides-extras`
= ["server", "proxy"], which the old uv had dropped.
1198 tests pass.
Replaces PR #40, which sat parked while main moved 85 commits ahead of it. Its own changes were only
201 lines, so re-applying them onto today's main was cleaner than rebasing 85 commits — and they
merged without conflict.
What ships, a week of work rather than one feature:
* the LOCAL PROXY — `treg claude`, `treg node app.js`, `treg shell --proxy`, `treg serve`: put treg
in front of any command and its calls to registered hosts are credentialed server-side;
* the CATALOG as the front door — `treg catalog search/get`, `treg call <endpoint-id>`,
`treg balance`, `treg topup`, the reordered help shelf, and onboarding that makes a real call
instead of seeding a fake echo tool;
* capability choice — observed success rate, latency and last-answered per endpoint, the decision
procedure in llms.txt and skill.md, and `treg org pin/pins/unpin`;
* Jason's platform keys + prepaid balance, the endpoint catalog and agent identity.
The install fix is the reason this is not just a version bump. `pip install "tools-registry[proxy]"`
is right for exactly ONE of the four ways treg is installed, and install.sh's own way is not it — a
uv-tool or Homebrew venv is not on the ambient pip's path, so that advice silently does nothing. Now
install.sh pulls the extra by default, and when the certificate library is missing treg offers to
install it correctly for THAT install and carries on.
Release checks: wheel + sdist built, light install verified (no fastapi/uvicorn/sqlmodel/asyncpg/
stripe, and cryptography correctly absent from the base since it is the [proxy] extra), twine check
PASSED on both artifacts, sdist carries no .env/db/.claude. 1184 tests pass.
Every new org gets a $1 promo balance. A /call/<endpoint-id> with no org
credential now falls through to treg's own provider key (tier 4 of the
credential ladder) for platform-eligible endpoints, metered reserve→settle
against the org's prepaid balance. Stripe (Checkout + off-session
PaymentIntents) is the payment rail only; the balance itself lives in an
append-only micro-USD ledger.
- ledger.py: credit blocks (promo/purchased), atomic conditional-UPDATE
reserve, promo-first consumption, lazy hold reaper; the only module that
moves money
- billing.py: Stripe top-ups ($5 min), card-on-file, auto top-up with
mandate consent, idempotency, monthly cap, cooldown, and
authentication_required recovery; webhook on /billing/stripe/webhook
- tier 4 in api.py: platform_setting virtual bindings, fail-closed daily
spend cap, 402s an agent can act on, allow-list kill switch
(TREG_PLATFORM_PROVIDERS, ships off); metered calls buffer and force
identity encoding so provider-reported costs always parse
- cost schema: per/unit/source/checked/confidence provenance on all 2,617
endpoints; fx unit rates; Catalog.platform_eligible as the single
billing-eligibility predicate; validator rules
- telemetry: CallRecord gains endpoint/provider/tier, estimated/observed/
charged cost, duration, bytes, params_hash; refused and malformed calls
audit too; Activity shows real charged cost
- reconcile.py: price-drift, provider-spend and repeat-rate reports
(super-admin); scripts/provider_balances.py
- dashboard: Team → Billing tab (balance, add funds, auto top-up); exact
micro-USD money formatting; #billing deep link; light-mode empty-state fix
- security: user bindings carrying platform_setting are rejected (key
exfiltration via relay); platform keys never resolvable in local-run
- fixes found by live verification: dataforseo extended paths double-/v3
(216 endpoints), pasted-secret trailing newline, CLI --data implies POST,
og:title after rebrand
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HiMBj2iKr88cQxmQ2ZogLb
treg only works today when the agent is TOLD to use it. An agent that writes its own script
calling api.stripe.com directly cannot see treg and has no key. The local proxy closes that
gap: catch the call on the way out with HTTPS_PROXY plus a certificate authority generated on
this machine, and let the SERVER add the credential (docs/LOCAL-PROXY-PLAN.md). We take
oneCLI's capture mechanism and reject its injection half, which decrypts real keys on the
laptop — the exact thing treg exists to prevent.
This lands the first two phases in src/treg/localproxy.py (named for localrun.py; NOT proxy.py,
which is the server relay this feeds).
P0 — the proxy. Binds 127.0.0.1 only, authenticates every connection against a per-session
token (Proxy-Authorization, compare_digest), and blind-tunnels CONNECT both ways. Plain http://
is forwarded too, since HTTP_PROXY is set alongside HTTPS_PROXY and an http caller must not
simply break; the forward strips Proxy-Authorization so our session token never reaches an
upstream. A loop guard refuses to tunnel to our own address, and one bad connection can never
take the listener down.
P1 — the certificates. ensure_ca() generates the machine's own authority (ECDSA P-256, 2 years,
key created 0600 BEFORE any bytes go in, regenerated inside its last 30 days, and regenerated
rather than dying if the files are corrupt). build_bundle() writes the system roots PLUS ours —
appending matters, because SSL_CERT_FILE REPLACES the trust list, so a bundle holding only our
CA would leave the agent unable to verify the real internet. context_for(host) signs a leaf and
caches it per host; Python's load_cert_chain only reads from disk, so the leaf key exists in a
0600 temp file for microseconds and is deleted immediately. proxy_env() returns the exact
environment P4 will publish, including NODE_USE_ENV_PROXY=1 — since Node 18 the built-in fetch
silently ignores HTTPS_PROXY, which is the failure that otherwise costs a day of debugging.
Interception is P2, so today the proxy is provably neutral: it never signs a leaf for a host it
tunnels, and a test asserts exactly that.
Decisions settled with Unclecode and recorded in the plan: the proxy lives inside `treg shell`
only (no standalone daemon, so the session token never needs to be written to disk); a 2-year
CA; port 18791; and a new [proxy] extra carrying cryptography alone, because it is compiled and
`pip install tools-registry` must stay the light CLI.
Verified with a real curl (OpenSSL — a different TLS stack from Python's): example.com and
api.github.com tunnel through with certificate verification intact, no token gives a 407, a leaf
we signed is accepted when curl is pointed at ca-bundle.pem, and REJECTED with the system roots
alone. That last one is the evidence for the non-negotiable we care most about — the system
trust store is never touched, so our certificates are trusted only where we deliberately put the
bundle.
34 new tests; suite 944 passing (one unrelated pre-existing failure on main, test_detail_pages).
- docs/context/architecture/catalog.md: the schema, the curation process, the
four endpoint states, and the security rules — including the PII policy
learned from a contact-lookup route that returned a named private
individual's personal email (such endpoints stay catalogued, carry no test
request, and store no example)
- api.md / cli.md / dashboard.md updated for the new surfaces
- tests for the catalog API, the CLI's browse/search/get, the marketplace
markup and the proxy path
- .gitignore: working plans, audits and screenshots stay out of the repo
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
No package added, removed or re-hashed — a newer uv writes revision = 3 and an
upload-time per artifact. Committed on its own so it never hides a real
dependency change inside a large diff.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The base package is now the CLI only (httpx + questionary — pure Python, ~13
deps). The FastAPI/DB/crypto stack moves behind a `[server]` extra:
pip install tools-registry # just the treg CLI
pip install "tools-registry[server]" # to host a registry
- localrun.py: the DB/crypto/oauth imports (used only in the server-side
render_grant) are now lazy / TYPE_CHECKING, so the module imports with the
light deps — the one place that leaked the server stack into CLI paths.
- pyproject: base deps trimmed; [project.optional-dependencies] server = [...];
the dev group self-refs tools-registry[server] so `uv sync` + CI still get it.
- render.yaml: buildCommand -> pip install ".[server]" (server deploy needs it).
- Bump 0.1.0 -> 0.2.0; add Unclecode as author + maintainer.
Verified: base install = 13 light packages, no fastapi/asyncpg/cryptography;
full server + 521 tests pass; render_grant works live.
A remote registry that turns a team's skills into shareable, callable tools:
any member's agent or a human calls a tool without owning its credentials —
a proxy injects the secret server-side.
This is the curated public tree (internal handoffs, plans, meeting notes,
and dev journal are kept in the private archive). Still WIP before going
public: see OSS-PREP-NOTES.md for the remaining genericization + the LICENSE,
CONTRIBUTING, AGENTS, and SECURITY items to finish.