Commit Graph
19 Commits
Author SHA1 Message Date
Jason ZhouandClaude Fable 5 8ca54c0c6f feat: treg.to is the canonical domain — treg.superdesign.dev becomes the legacy alias
public_url/email_from defaults, render.yaml, CLI fallbacks, packaging (npm/plugin/pyproject),
web pages (+canonical tag), docs and context fragments all move to https://treg.to.

The legacy host keeps serving the FULL API forever — installed CLIs, skill.md files and
.mcp.json configs in the wild hold Bearer tokens pointed at it, and HTTP clients strip
Authorization on cross-host redirects. Only browser-facing marketing pages 301 to treg.to
(new middleware + tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 11:56:20 +10:00
Jason ZhouandClaude Fable 5 ce47d44213 fix: gzip /mcp/ responses at the origin — the edge ignores no-transform
Issue #100: Render's edge kept Brotli-compressing large /mcp/ responses
despite Cache-Control: no-transform, and a real client stack
(httpx + brotlicffi) dies mid-decode and hangs to its own timeout.

An edge only compresses what arrives uncompressed, so the origin now
gzips its own responses (GZipMiddleware inside build_mcp_app, >=1KB,
respecting Accept-Encoding): the edge passes them through, and gzip
decodes via zlib on every mainstream client — the brotli decoder never
runs. no-transform stays: correct and free, just not sufficient.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 10:28:34 +10:00
Jason ZhouandClaude Fable 5 8697489c2d fix: catalog_get.example_response accepts any shape — brightdata examples are arrays
Found live: catalog_get(brightdata.linkedin.user.profile) failed the
server's own outbound validation because its example_response is a list
of records, not a dict. Typed Any like call.body; the e2e schema test
now covers that exact endpoint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 09:54:31 +10:00
Jason ZhouandClaude Fable 5 c690a20c81 fix: MCP responses validate against their advertised schemas, and forbid edge re-compression
Two field reports arrived the same day (#93, and one on X): strict MCP
clients refused every tool response with -32602. The SDK serializes each
response through the TypedDict-derived model, which fills absent
total=False keys in as null in structuredContent — while the advertised
schema said 'type: string'. Every output field is now nullable, and the
suite validates real structuredContent against the advertised schema.

The other half of #93: Render's Cloudflare edge Brotli-compresses large
/mcp/ responses, and at least one client stack (httpx + brotlicffi) dies
mid-decode and hangs to its own timeout. Every MCP response now carries
Cache-Control: no-store, no-transform, stamped by an outermost ASGI
wrapper so 401 challenges carry it too.

Closes #93

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 09:28:46 +10:00
UncleCode 0403f8fdd6 feat(mcp): call takes an idempotency_key — the feature reaches the surface it was built for
The whole thing came from "agents need idempotent billing", MCP is the agent path, and the tool could
not send a key. Built for agents and left unreachable from the agent surface.

The key is the CALLER's, and I nearly got this wrong. I first proposed deriving it server-side from
the endpoint and params, so a model would not have to manage labels — which contradicts a warning I
had written into the plan myself: "a retry guard, not a response cache. It must never serve a request
that did not carry the same key, or it stops being a correctness feature and starts being stale
data." Unclecode caught it. Two identical searches an hour apart are new work, and treg cannot tell
that from a retry; only the caller knows.

So the description tells the model WHEN, because that distinction is where a mistake costs something:
same key when repeating a call whose answer never arrived, a new key (or none) for genuinely new work
even with identical parameters. Reuse for a DIFFERENT request is refused rather than answered, so the
failure mode is loud.

The result carries `replayed: true` and a hint that nothing was charged, so an agent reporting cost to
a human can say so accurately.

The end-to-end test earned itself immediately: my first attempt at the header edit did not apply at
all — the pattern did not match — and the test is what caught it rather than a passing suite.

3 new tests, 1327 pass.
2026-08-12 15:47:23 +08:00
Jason ZhouandClaude Fable 5 203e7832c1 fix(mcp): no_key_needed now requires an actually-enabled platform key
catalog_search advertised no_key_needed from platform_eligible alone (the
price side), so an eligible endpoint whose provider had no configured key
promised keyless service and then refused at call time. AND it with
platform_key_for — the same key-AND-allow-list check the call path uses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 17:18:11 +10:00
Jason ZhouandClaude Fable 5 6947725b62 fix(mcp): eager auth — challenge initialize, not just tool calls
Every treg tool needs auth, so there's nothing to browse anonymously —
but leaving initialize/tools-list open meant a client (Claude Code,
Cursor) got a 200 on connect, showed "Connected", and never prompted:
connected-but-unusable, the sign-in surfacing only later as prose in a
tool result. The MCP spec's canonical flow challenges the FIRST request
so OAuth runs before the session proceeds (Stripe/Subframe/AuthKit do
this; FastMCP tracks a 401-free initialize as a bug).

Now every id-bearing JSON-RPC request (initialize, tools/list,
tools/call) without a valid credential gets 401 + WWW-Authenticate;
only notifications and ping pass. The challenge also carries `scope`
(spec SHOULD) for least-privilege up front. The transport's host (421)
/ origin (403) guards sit behind this, so a credentialed bad-host
request is still refused by them — auth doesn't mask the guard.

Drops the now-unused _NEEDS_AUTH set; gate is by method, not tool name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
2026-08-11 10:36:35 +10:00
Jason ZhouandClaude Fable 5 a1d4457ddf fix(mcp): a dead access token gets 401 invalid_token — the refresh cue
An expired OAuth access token presented on a tools/call sailed through
the middleware (which only challenged the missing-header case) to the
tool's friendly prose in a 200. RFC 6750 has the resource answer 401
with WWW-Authenticate: Bearer error="invalid_token" — the machine-
readable cue on which a client silently runs its refresh grant. Told
nothing, Claude Code gave up with "requires re-authorization (token
expired)" every hour instead of refreshing.

Only bearers that CLAIM to be our OAuth access tokens are judged
(looks_like_access_token reads the unverified payload's typ); a
per-org/identity token (the Codex env-var path) passes through for the
API to validate downstream — its holder has no refresh grant to run.
Access-token validation is stateless (HMAC + expiry), so the transport
can afford it; nothing needing the database moved into the middleware.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U72hGQwYvVSyzh7XGqNra9
2026-08-11 10:09:19 +10:00
UncleCode 3d4aa21497 fix(mcp): actually strip the purchase pointer from a 402 — the first fix did not
Shipped an hour ago and verified on production afterwards: the link was still there. Two mistakes,
both worth writing down because they are the same mistake in different clothes.

**1. I stripped the wrong thing.** The fix popped a top-level `topup_url`, but the real body nests
everything under `detail` AND repeats the URL inside a prose `message`:

    "message": "... would cost ~$0.875 ... \n  add funds:      https://…/app#billing\n  or use your
                own key: treg connections connect --provider akta"

`_without_purchase_pointers` now walks the structure: drops the key wherever it appears and removes
any line containing a URL. The diagnosis survives — the amounts, and the "use your own key"
alternative — only the invitation to pay goes.

**2. The test checked the CODE, not the RESPONSE.** It asserted on the source text of mcp.py, so it
passed while production returned the link. That is exactly the failure this codebase keeps catching
in other places, and I wrote one. Worse, the replacement tests exercised the helper directly and
STILL passed with the strip deleted from `call` — the helper worked and nothing connected it to what
a user sees. Only the third attempt, which drains a balance and drives the real MCP tool, fails when
the strip is removed. Verified by removing it.

One more thing the first version got wrong: it matched `public_url`, which differs per environment,
so it stripped nothing anywhere except production while the local test passed. It now matches any
URL, because "no link out" does not depend on which host we happen to be running as.

Cost of finding this honestly: a $0.38 PDL call against the reviewer balance while probing for a
402, and a $0.875 refusal that (correctly) charged nothing.

3 new tests, 1295 pass.
2026-08-10 20:41:24 +08:00
UncleCode 50e2277817 fix(mcp): the 402 states the fact and stops — no purchase pointer
ChatGPT's submission form asks whether a plugin "links or directs users out of ChatGPT to make
purchases", and states that only PHYSICAL goods can be supported. treg sells prepaid API credit,
which is a digital good — so the top-up link made the honest answer a yes, in the one category they
cannot support. Unclecode's call: drop the link.

The hint now reads "the team's prepaid balance is not enough for this call". The diagnosis survives —
an agent still learns it ran out of money rather than hitting a broken endpoint — it is simply not
handed a payment page.

`topup_url` is stripped from the RELAYED BODY as well, and that is the part worth the test. The
API's 402 carries the field, so removing only the sentence would have left the link sitting in the
JSON while the code looked clean. The test asserts on the whole response for exactly that reason.

Scoped to the MCP path deliberately. `/call/`'s 402 still carries `topup_url` for the CLI and the
dashboard, where no such policy applies and the shortcut is genuinely useful — llms.txt line 112
still documents it for those callers, correctly.

Also worth recording: the plugin manifest is ALREADY 0.7.1 (Jason's release commit), matching
pyproject. I had said it was 0.7.0 — wrong. Only the OpenAI form draft still shows the old number,
which is a field in their UI, not a file here.

2 new tests, 1293 pass.
2026-08-10 20:27:42 +08:00
UncleCode 8c6abffefa fix(mcp): call accepts an ARRAY body — DataForSEO's 217 endpoints were uncallable
Found while looking for a demo endpoint, by trying a real call rather than reading the signature.

DataForSEO takes an ARRAY of task objects on every one of its `live` POST routes — that is its API
shape, not an oddity — and `call` typed `params` as `dict | None`. Every one of those endpoints
answered with a pydantic type error, which reads as "you passed it wrong" when the truth was "this
tool cannot express that". DataForSEO is the largest provider in the catalog at 217 endpoints, so a
whole family was silently unreachable through MCP while working fine through the CLI.

`params` is now `dict | list | None`. A list on a GET is still refused, because a query string is
key/value pairs and a list has no meaning there — but it says so, rather than surfacing a validation
trace that blames the caller for the tool's own limitation. The tool description now states both
shapes, since the model reads that before deciding what to send.

Two tests: an array body reaching the upstream, and the GET refusal explaining itself. The first one
caught a mistake of my own — the tool was registered in one org and called from another, which is
correctly invisible; the test now registers it as the calling token's org.

1290 pass.
2026-08-10 19:20:23 +08:00
UncleCode 768314d83e fix(mcp): every tool needs a credential, catalog included
Unclecode's call, and the simpler contract. The previous split — catalog tools open, the rest
protected — meant "some tools need auth, some do not", which every client has to learn by trying.
Now it is one rule: connect, then use treg.

Worth being precise about what this buys, because it is easy to overstate. `/catalog/search` is still
public on the WEBSITE — the landing page and `treg catalog search` both rely on it — so this is about
giving MCP clients one predictable rule, not about hiding the catalog. Anyone wanting that data
unauthenticated can still take the website route, and closing that is a separate decision with
different consequences. Unclecode was explicit that this change is for MCP only.

`initialize` and `tools/list` stay open, deliberately: that is how the flow STARTS. A client connects,
lists what exists, calls a tool, gets the 401 with WWW-Authenticate, and only then authorizes. If
discovery itself challenged, ChatGPT could not scan the server to find the tools at all and the
connector page would show nothing.

Also corrected two comments that still described the old behaviour, and the connect-demo copy that
told the reader two tools were public. The catalog is still read straight from memory rather than
through the API — that stays, but it is a SPEED choice, not a permission one, and the docstring now
says so rather than implying identity is unnecessary.

Three tests that assumed anonymous catalog access were given tokens (their subject — priced results,
the no-match hint, price-before-spend — is unchanged), and the one asserting public access was
inverted to assert the challenge.

1288 pass, verified on a live server: 401 for all five tools, 200 for initialize and tools/list.
2026-08-10 19:03:55 +08:00
UncleCode 9f3f97fb83 fix(mcp): protected tools answer 401 with WWW-Authenticate, as the spec requires
Unclecode asked why /mcp/ never 401s. He was right to. The answer split in two.

The catalog staying PUBLIC is deliberate and unchanged: someone evaluating treg should see what is in
it and what it costs before creating an account, and that is the same data /catalog/search already
serves on the website. Nothing tenant-specific and nothing spendable was ever exposed.

The gap was the SHAPE of the refusal. `balance` without a token returned HTTP 200 carrying
`{"error": "not authenticated"}` and no `WWW-Authenticate` header. The MCP spec has a protected
resource reply 401 with `WWW-Authenticate: Bearer resource_metadata="…"`, because that header is how
a client DISCOVERS it must authenticate and where to begin. A friendly English sentence inside a 200
tells a human what happened and tells a program nothing — a client that connects first and discovers
auth lazily, which the spec allows, would read 200 as success and surface our prose with no way to
act on it.

It passed unnoticed because ChatGPT authenticates up front. I had actually SEEN it: an RFC 9728 probe
earlier returned `HTTP/2 200` and I moved on because the client in front of me did not need it —
the same reasoning that nearly cost us dynamic client registration.

`RequireAuthForProtectedTools` sits in front of the transport, because status codes and headers are
an HTTP concern and a tool function can return neither. It challenges only when there is NO
credential: deciding whether a token is VALID needs the database, and doing that in middleware would
put a second authentication implementation in front of the first.

One thing worth recording about the first attempt: I wrapped only the module-level app, and the tests
— which build a fresh transport per test — silently exercised an unprotected server and passed. The
wrapping moved inside `build_mcp_app()` so a caller cannot get a different server from the one
production runs.

7 new tests; two older ones updated to the new contract while keeping their original job. Verified on
a live server: 401 + header for the three protected tools, 200 for the public two, 200 for
initialize and tools/list. 1286 pass.
2026-08-10 18:47:51 +08:00
UncleCode 1d806d239c feat(mcp): output schemas for all five tools — optional AND nullable
ChatGPT's connector page flagged every tool with OUTPUT SCHEMA RECOMMENDED. A model that has to guess
at field names guesses, so each tool now declares its shape.

Getting there turned up two traps, both of which would have shipped as regressions:

**A strict schema breaks every error path.** Schemas are validated on the way OUT, so the first
`{"error": "not authenticated"}` raises instead of returning — converting a message written to tell
an agent how to recover into an opaque tool failure. Every field is therefore optional. A schema is a
hint to the model, not a gate on our own error handling, and a test asserts no tool has required
output fields.

**Optional is not the same as nullable.** `total=False` says a key may be ABSENT; it says nothing
about the value being null. Real rows carry nulls — a registered tool with no description, an
endpoint with no published price — and typing those as plain `str` made `my_tools` return a schema
error INSTEAD OF the team's tools.

That one was caught by an assertion written days ago for a different reason: the isolation test
insists org A must genuinely SEE its tool, so that "org B cannot see it" can never pass for free. It
was there to prevent a vacuous pass and it caught a real regression. A new test now pins the
nullability rule so the next field added is held to it.

Also confirms the annotations from the earlier pass are being read: ChatGPT's UI shows `call` as
PUBLIC WRITE / OPEN WORLD / DESTRUCTIVE, which is the honest labelling rather than the comfortable
one.

3 new tests, 1277 pass.
2026-08-10 18:15:19 +08:00
UncleCode 9da493d9c3 feat(oauth): a connect demo page — and the two browser bugs it found
Unclecode's idea, and it earned itself twice over before ChatGPT ever saw the flow. `/connect-demo`
pretends to be someone else's app and does exactly what an MCP client does: dynamic registration, a
popup to treg for approval, code exchange, then tool calls. Only public endpoints — nothing about
being served from treg's own domain makes it privileged.

**Bug 1: every browser was refused.** Tool calls answered "Invalid Origin header". In this SDK `"*"`
is NOT a wildcard — origins are compared literally, and only a `:*` port suffix is special — so
`allowed_origins=["*"]` permitted exactly one origin: the literal string "*". Nothing caught it
because the suite and every CLI client send NO Origin header, so the check never ran until a real
page called /mcp/. `_allowed_origins()` now builds a real list, with TREG_MCP_ALLOWED_ORIGINS to
extend it, and two tests send a browser Origin deliberately — one that must be served, one that must
still be refused. This is the second time the same default has bitten in the same shape as the host
allow-list, and ChatGPT's web client is a browser, so it would have hit that too.

**Bug 2: approval failed intermittently.** Unclecode saw "cross-origin authorization rejected"; the
log showed one 403 then a 302 from the same IP. `Origin: null` is a browser reporting an OPAQUE
origin, which is what happens after certain redirect chains — a consent page reached by way of a
sign-in bounce. It is not evidence of a cross-site request. `_same_origin` now accepts `null` only
when `Sec-Fetch-Site` corroborates it, which script cannot forge, so the guard is narrowed rather
than removed. The refusal also now names the values it saw: the guesswork that cost us here was
avoidable.

The `call` button was left out at first because it is the only one that spends, and a demo that
charges on each click is a trap — but omitting it silently read as an oversight, which is fair. It
now does what skill.md asks an agent to do: read the price, say it ("will spend $0.0245 from your
team's balance"), and ask before spending.

Testing that turned up a third, smaller thing: the price sits on `endpoint.cost.usd`, not the top
level, and a bare try/catch swallowed the mistake so the button just stopped. A demo that fails
silently teaches the reader that the FEATURE is broken rather than the code. Every handler now
reports its own failure.

3 new tests, 1271 pass.
2026-08-10 17:38:28 +08:00
UncleCode 0bf9194dfc feat(mcp): tool annotations + the domain-verification endpoint — what review actually checks
Reading the submission portal's own requirements turned up two hard blockers that would have failed
at "Scan Tools", not at submit.

**Every tool must declare readOnlyHint / openWorldHint / destructiveHint**, and review validates them
against real behaviour. The four reading tools change nothing anywhere. `call` is the honest
exception and is marked in the strongest direction: it relays whatever the caller asks to whichever
upstream the endpoint names, so it can be a POST that publishes, an email that sends, or a DELETE.
treg deliberately does not model the upstream — that is the founding rule — so it cannot promise the
call is safe, and a comfortable annotation here would be a false assurance in the exact place a model
consults before acting. It also spends real money.

**Domain verification** needs a token at /.well-known/openai-apps-challenge on the MCP host. The
portal issues it; the endpoint must return THAT token as plain text and nothing else — the docs are
explicit that JSON, a list, or several tokens all fail. It is driven by a new
`openai_apps_challenge` setting and 404s while empty, which is right for every deployment that is not
ours: an empty 200 would read as a verification that never completes, and that is harder to debug
than a plain absence.

1214 tests pass.
2026-08-09 20:08:59 +08:00
UncleCode a7ac524dd9 fix(mcp): call can reach the team's OWN tools, not only the catalog
Found by using it on production rather than by reading it: `my_tools` listed the seven tools this
team has registered, and `call` refused every one of them with "unknown endpoint". The tool that
lists things an agent can use was paired with a tool that could not use them.

The cause was `call` pre-checking `catalog_store.by_id` and refusing anything absent. `/call/{rest}`
already resolves a team's own tool FIRST and falls back to a catalog id — that ordering is the
product's "your own key always wins over treg's" rule — so the fix is to stop second-guessing it and
pass the target through.

`call` now takes either shape: a catalog id from `catalog_search`, or `<tool-name>/<path>` from
`my_tools` (e.g. `render/v1/services`). The catalog lookup survives only to pick the right HTTP
method and to give a good error for a bare name that matches nothing; an optional `method` argument
covers a team tool that needs POST.

The description says both shapes plainly, because it is what the model reads when deciding what is
possible.

1209 tests pass.
2026-08-09 19:29:25 +08:00
UncleCode db8ecddbe6 fix(mcp): resolve the team for an IDENTITY token, not just a per-org one
Production found this within a minute of the deploy, which is the argument for testing on the real
thing: `balance` answered "could not resolve the team for this token" for a perfectly valid token.

There are two kinds, and they answer differently. A PER-ORG token (`treg org agent-new`) has its team
baked in, and `/auth/me` reports it. An IDENTITY token (`treg login` — what most people actually
hold) belongs to a PERSON who may be in several teams, so `/auth/me` reports no org at all and every
`/orgs/{id}/…` route has to be told which one. I had handled only the first kind, so the commonest
token there is got a dead end.

`_resolve_org` now does what `cli._active_org_id` does: ask `/auth/me`, fall back to `/orgs` and take
the active team, and pass `X-Treg-Org` on the routes that need it. When a person is in several teams
and none is active it NAMES them and asks rather than picking one — reading the wrong team's balance
is merely confusing, and spending from it would not be.

`balance` and `my_tools` now report which team they answered for, so a multi-team human can see at a
glance that it was the right one.
2026-08-09 19:07:41 +08:00
UncleCode 9ce587fbe6 feat(mcp): the MCP front door — five tools, ~1.5ms, and one implementation of the rules
An agent inside ChatGPT or a Codex plugin has no terminal, and a skill that says "first install this
CLI" is where the visitor leaves. This is the other door: `/mcp`, streamable HTTP, five tools.

FIVE tools, not 2,600. The catalog stays DATA — catalog_search, catalog_get, call, balance,
my_tools. A tool per endpoint would bury the model's context in 2,600 schemas and make the catalog
unusable, which is the opposite of the point.

Everything touching money, tenancy or credentials goes back through treg's OWN HTTP API in-process
(httpx ASGITransport) rather than reaching into the internals a second time. `/call/` already
enforces the member ACL, deny rules, both daily caps, the balance reserve and the settle; a second
entrance that re-implemented any of that is how one copy quietly stops being enforced. Only the
public catalog is read directly, because it is already in memory and needs no identity.

Three traps found by building it rather than reading about it:

  * `app.mount()` does NOT run a mounted app's lifespan, and the transport builds its task group
    there — every request 500s with "Task group is not initialized". api.py now COMPOSES the two
    lifespans. It caught me twice: once in the spike, once in my own test fixture.
  * The SDK ships DNS-rebinding protection ON with an EMPTY allow-list, which 421s EVERYTHING. The
    deploy would have looked healthy until the first tool call. `_allowed_hosts()` builds the list
    from this deployment's `public_url` plus loopback, with TREG_MCP_ALLOWED_HOSTS to extend it, and
    a test asserts both directions — the real host works, an unknown one is refused.
  * `StreamableHTTPSessionManager.run()` may be called once per instance, so `build_mcp_app()` is a
    factory; tests get a clean transport each while still exercising the real tools and the real
    enforcement path.

Proven from Codex against the dev box, not just in tests. Asked for work-email options with no API
key: it searched, read two endpoints, and compared hunter ($0.0245/success), leadmagic ($0.025) and
thecompaniesapi ($0.0019) — noting the third returns a PATTERN rather than a confirmed address, then
declining to spend because I said not to. balance reported $1.00 (the new-org grant) and my_tools 0.
`call` went through the full path and returned the correct refusal with `whose_error: treg` — the dev
box holds no provider keys, so a SUCCESSFUL metered call stays unproven until this reaches prod,
where they live.

Import is optional and mounted last: a deploy that has not picked up the `[server]` extra's `mcp`
dependency still serves everything else. The lock was regenerated with a modern uv (0.12) rather than
this Mac's 0.5.2 — format preserved at revision 3, additions only, and it restored `provides-extras`
= ["server", "proxy"], which the old uv had dropped.

1198 tests pass.
2026-08-09 18:55:42 +08:00