49 KiB
title, status, sources, related
| title | status | sources | related | |||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| The proxy — faithful credential-injecting relay + tool resolution | shipped |
|
|
Proxy and call execution
application.call.service orchestrates resolution, authorization, reservation, relay and
finalization. application.call.resolve selects the target and credential;
infra.upstream.relay.relay only injects credentials and forwards bytes.
Named catalog calls with authorization metadata select the provider and grant method before
comparing hosts. This separates Facebook and Instagram tools sharing graph.facebook.com;
the resulting tool enters the same relay without provider-specific relay logic.
A live-verified free GET catalog row can declare platform_auth: anonymous. The normal team tool
and team credential still win. If neither exists, _anonymous_offer requires the deployment's
provider allow-list and creates an unmetered virtual tool with no bindings. relay() therefore
forwards the request without injecting any provider credential. The field is generic catalog
metadata; the resolver and relay contain no provider-specific anonymous path rules.
Cache experiment metadata is attached to the existing tool_called event by the call-service
capture funnel: outcome/reason, comparison and TTL policy, rollout percentage, lookup duration,
and candidate age/window, plus cache_price (full | repeat | free) on a hit, and, like
every server event, the build and archive_config fingerprints from analytics.py. It contains
no response/request content or cache key. The stable team/endpoint rollout runs before archive
DB lookup (open to every endpoint and team by default); unselected calls retain
the normal relay and money path. Own-key catalog calls take part too: a storable own-key 2xx is
read whole when it fits the archive's cap and is asked for identity encoding, otherwise it
streams untouched; own-tool calls never touch the archive. See
archive and its pricing section for controls and metric
denominators. This does not remove authorization/reserve/settle DB work.
The faithful-relay contract
relay() alters only three things; everything else is verbatim (method, path, all query params
incl. duplicates, headers, cookies, body bytes):
- hop-by-hop transport headers -
_HOP_BY_HOP(host, content-length, connection, keep-alive, te, trailers, transfer-encoding, upgrade, proxy-*); re-derived per hop or the stream corrupts. One value is carried, not re-derived: when the caller declared aContent-Lengthfor a body, the relay sets the same value on the upstream request. The bytes are the caller's, unaltered, so the length is exact, and without it httpx frames the streamed bodyTransfer-Encoding: chunked, which Meta's Graph API edge does not read (every POST body vanished; an ad creative failed "Ad incomplete" until the spec was moved into the query string). A caller who streamed chunked stays chunked. - treg's control/infra + edge forwarding headers -
_CONTROL(x-treg-token,x-treg-org,ngrok-skip-browser-warning,x-forwarded-*,x-real-ip,forwarded,via), dropped via_DROP_REQUEST = _HOP_BY_HOP | _CONTROL, so none leaks upstream._scrub_treg_cookiesalso strips treg's own cookies (treg_session,treg_oauth_state) from the Cookie header - the dashboard'scredentials:'include'Try-it would otherwise leak our session token - while keeping other cookies. - the injected credential(s) - each binding overwrites only its target header, query parameter, or top-level JSON field. A JSON binding is an explicit provider contract and is the only case that buffers and reserializes the caller body; all other bodies retain the streamed byte path. This is semantic JSON relay, not byte-faithful relay: formatting changes, non-ASCII text is emitted as UTF-8, and duplicate keys fail closed. Never use a JSON binding where the upstream signs or hashes the caller's raw body bytes.
What treg keeps from a call. Successes retain no content: the relay forwards bytes and the audit row records status, size and timing. A failed relayed call - platform, own-key, or plain own-tool
- is the exception:
CallRecord.error_request/error_responseretain a redacted, truncated copy of what the caller sent and what the provider (or treg-side 502) answered. Without it a failure is a bare status code:pathholds the catalog URL rather than the caller's parameters andparams_hashis one-way.application.call.settlebuffers metered responses with_buffer_response, while_peek_stream_headreads only the first 8 KiB of a failed unmetered response and replays every consumed byte before the rest of the original iterator, preserving status, raw headers, streaming, and the upstream-close task. Caller bodies on unmetered paths are cached only whenContent-Lengthis declared and at most 64 KiB; large/chunked uploads stay streaming and retain only their query-param half. See data-model for the redaction order, admin-only access, and retention.
Faithfulness mechanics inside relay():
- request headers rebuilt from
UpstreamRequest.raw_headersinto anhttpx.Headersmultidict (preserves duplicate headers / cookies); injection (headers[name] = v) overwrites only the named one. - on the platform tier only,
servicefirst passes the raw headers throughrelay.scope_shared_idempotency_key, which replaces the caller'sIdempotency-Keywith a digest of (org, label). Every org shares one provider account on treg's key, and a provider that honors the header (LeadsForge does) would otherwise return org A's job to org B under the same label — andresource_ownership.produceswould then record A's job id as B's. Treg's own idempotency table already replays a caller's answer for the same label, so the caller loses nothing. A team's own key relays the header verbatim: that account is theirs. - query as the router-captured ordered pairs in
UpstreamRequest.query_items(keeps duplicate keys like?tag=a&tag=b), merged onto the upstream URL withcopy_add_paramrather than passed asparams=: httpx replaces a URL's existing query wheneverparamsis given, even empty, which silently stripped a catalog path's own query (/rest/images?action=initializeUpload) until 2026-09-19.tests/test_relay_path_query.pypins it. - path rebuilt from
request.scope["raw_path"](incall_tool), not Starlette's URL-decoded path param - percent-encoding survives to the upstream (npm's scoped publishPUT /@scope%2fname404s if%2fis decoded to a literal slash). - body streamed via
content=request.stream()(stream, never buffer). The caller'sContent-Length, when declared, rides along so the stream is not re-framed chunked (see the contract above). Exception: a caller may base64/gzip-encode the body withX-Treg-Body-Encodingto slip SQL/HTML past a hosting-edge WAF;_BodyDecodeMiddleware(in api.py) then buffers + decodes it beforerelay()runs, so the relay still forwards the real plaintext bytes verbatim upstream. See api. - upstream call uses the shared
client(the long-livedhttpx.AsyncClientatapp.state.http, created inlifespan- keepalive is the biggest latency win). - the infra relay returns framework-neutral
UpstreamResponse(status, raw_headers, body_stream, close). The router wraps it inStreamingResponseand copies every upstream response header (incl. multipleSet-Cookie) minus_DROP_RESPONSE. Its body wrapper and background task share the same idempotent close operation, so full reads, partial disconnects, stream errors, and cancellation close the upstream response exactly once.
A request may carry several credentials: relay() loops tool.bindings and calls
injectors.inject(headers, params, binding, crypto.decrypt(secret.value), json_body=...) per binding.
Bindings can also stamp provider protocol constants: a format with no {secret} renders literally
(Crustdata's required API-version header is the first registry use). It still carries the same secret
reference for binding validation and lifecycle, and the assignment overwrites a caller-supplied value.
This is generic binding behavior, not an upstream-specific branch in the relay.
An anonymous catalog fallback uses the same relay with an empty binding list. It does not strip
caller headers or rewrite the request; it only omits a credential that treg would otherwise inject.
The catalog price is free only when the caller does not supply a provider credential header.
Platform bindings - injecting treg's OWN credential. A binding with a platform_setting key (instead
of a secret_id) injects one of treg's own credentials read from get_settings() - the Google Ads
developer token is the case that exists. The value never lives in the org's secret store, so a tenant
can't read it or extract it through a local run; a missing setting is a clean 502
(this server has no <setting> configured). Used by the OAuth-marketplace auto-provisioner for a provider
that needs a second credential treg holds centrally, and by tier-4 catalog calls. Tier 4 also copies
the provider's constant required_headers bindings, so Crustdata's x-api-version: 2025-11-01 pin is
identical on BYOK and platform-key calls (see api).
A separate case that looks similar but is NOT a platform binding: the Google Ads conversion
uploader (adsconv.py) also spends treg's own platform connection, but it is not a caller-issued
/call/ request at all, so it never reaches relay() or infra/upstream/injectors.py - it reads the platform org's
stored OAuth secret directly and builds its own headers. See ads-conversions.
Accept-Encoding is normalized to identity when the caller sent none. relay() streams the upstream
body raw (aiter_raw), so if the caller doesn't ask for compression httpx would otherwise add its own
Accept-Encoding: gzip and hand a plain HTTP client / agent compressed bytes it never requested. Asking
for identity keeps what the caller receives matching what the caller requested.
Connection discipline: a call in flight holds no DB connection
Resolution, authorization and reservation own short sessions. service._execute_call commits
the request's secret-loading session before opening reservation, and again before relay or the
live demo's network call. The invariant is zero checked-out request connections during upstream
I/O, including rate-smoothing waits.
Settlement, first-call recording and idempotency storage use their own sessions. A commit does not
invalidate loaded objects because the session makers use expire_on_commit=False.
OAuth refresh likewise separates DB reads/writes from token-endpoint network I/O.
Keeping the request transaction open through relay can deadlock a saturated pool: each request
holds a slot while its settlement waits for another. tests/test_call_pool_discipline.py checks
pool occupancy at the network boundary and covers concurrent calls. Pool sizing, timeouts and
the separate API/admin/background pools are specified in deploy.
Tool resolution (application.call.resolve)
* /call/{rest:path} → routers.call.call_tool() → routers.call.run_call_surface() (shared with
/catalog/call/ and /table/, which differ only in the finish that turns the answer into the
response: see table) → application.call.service.execute_call()
→ resolve_call_target(...) returns a framework-neutral
ResolvedTarget(tool, upstream). Each resolution use case owns and closes its read session.
A named miss that is not an org tool falls through, in order: a catalog endpoint, then — only for
a top-level call (request.context.input.child_of is None) — a hub tool (<team-slug>.<name>,
hub_app.tool_for). An own tool or a catalog id always wins; a hub tool never shadows either. See
hub for what a hub run is; service._execute_call runs it under a per-team
hub_limits.slot(), refusing a new run with 429 hub_busy (naming the active count and a
retry_after_s) rather than queuing past hub_limits.MAX_RUNS_PER_TEAM.
Both shapes are scoped to the caller's org (Tool.org_id == org_id), so two
orgs resolve independently and may reuse a tool name or upstream host; the use case then loads only
same-org secrets. After resolution application.call.authorize runs tool/project ACL, deny, member-cap,
and public-demo gates in that order, with no money hold or upstream access. Its short session closes before
the reserve stage; -1/default member caps add no query. Two resolution shapes:
* /catalog/call/{rest:path} is the narrower entrance used by catalog-only MCP surfaces.
routers.call.call_catalog_endpoint sets request.state.catalog_only (gated on
claude_connector_enabled) and then enters the same call_tool handler; the flag travels on
CallInput into execute_call. Resolution accepts only an exact catalog endpoint id and never calls
resolve_call_target, so a private team tool or arbitrary passthrough path cannot shadow the catalog
entry. Everything after catalog resolution stays shared: credentials, ACLs, deny rules, caps,
cancellation cleanup, metering, audit, idempotency, and faithful relay.
- URL-passthrough (agent-native):
restis the real upstream URL (/call/https://api.intercom.io/me)._normalize_scheme()restores thehttps://a path param collapses tohttps:/. The tool is resolved by host (_host_of()=urlsplit(...).netloc, matched against the indexedTool.host) then the longestbase_urlprefix; a tie →409, no match →404(or403when the caller's ACL is the only thing that removed the match - see below). - Named:
rest = "<tool>/<path>"(rest.partition("/")), looked up byTool.name; upstream URL =base_url + path. No path → the base URL itself, without a trailing slash - a tool pinned to a full resource (.../v1/charges) must relay as-is, since Stripe404s/v1/charges/.
A named miss whose <tool> is an exact catalog id with a path behind it (reapi.tasks.get/tasks/1)
is the own-tool shape applied to the catalog half: it answers 400 naming the endpoint's parameter
slots and the --query form, before any hint below runs. Named misses also inspect the org's caller-usable own tools on the error path. When a dotted operation
name shares its provider/first segment with one (for example google-analytics.report beside the
connected google-analytics tool), the 404 carries hint plus did_you_mean and points at
/call/google-analytics/<path>. If that dotted name is a real catalog endpoint, the hint follows the
catalog fall-through and is attached only if the marketplace credential ladder also dead-ends. Catalog
near-id matching remains provider-local and takes precedence for genuine misspellings.
If both shapes miss with 404, a dotted target gets one final lookup in the endpoint catalog. A live
row enters _resolve_marketplace_call and its credential ladder. _marketplace_upstream fills catalog
path placeholders by percent-encoding raw values, but preserves a value containing a valid %HH
escape. This prevents an already encoded Search Console property id such as
sc-domain%3Aexample.com becoming double-encoded as %253A. Raw @ remains literal because it is
a legal path-segment character. This also supports email-path APIs such as Tomba's verifier, which
rejects %40 before decoding. Slashes, query/fragment delimiters and invalid percent signs remain
escaped; URL-passthrough bytes are unchanged. Before building that URL, an optional catalog host
on a provider that opted in to catalog_targets must resolve through
OAuthProvider.profile_for_catalog_host to an exact approved HTTPS base URL. Providers without
that opt-in keep resolving catalog paths against their primary base URL.
The approved root keeps its path prefix when _marketplace_upstream appends the endpoint path, and
the selected provider profile supplies the correct credential binding. An unapproved or malformed
target fails as a treg-owned 502 before reserve and relay; catalog data cannot redirect an injected
credential to a host of its choice. A retired/broken tombstone is
instead refused with 410, its status_note, and its optional superseded_by, before credentials are
selected or the relay can run; the refusal is audited as refused_by=retired. This ordering is
deliberate: an org's own tool named exactly like the old catalog id already resolved above and is not
shadowed, while URL passthrough has no catalog-id shape to catch accidentally.
Capability alternatives. A 410 tombstone without superseded_by, or a credential-ladder
404, includes _capability_alternatives(ep): curated same-capability endpoints, cheapest first,
marked as callable on treg's key or requiring an own credential. Marked catalog rows are excluded.
The helper is synchronous and performs no DB reads; measured reliability belongs to catalog
inspection. It suggests alternatives without substituting one. Endpoints with no capability
cannot participate.
ACL-filtered candidates. _resolve_call takes the caller and filters passthrough candidates by
_tool_usable (project scope AND the per-tool list) before the longest-prefix tiebreak. A same-host
tool the caller cannot use must not be able to cause a 409 - or win the tiebreak - for someone who
cannot even see it in list_tools. This only NARROWS the candidate set, so it can never grant access:
whatever resolves still passes _require_tool_use. The named shape needs no filter (it resolves one
tool, then the gate runs).
ACL-only misses return 403. Resolution retains unfiltered host matches to distinguish an absent tool (404) from candidates removed solely by ACL (403). The latter names only the host the caller supplied, never an inaccessible tool.
Policy deny (_enforce_deny, _deny_match). After resolution and the tool ACL, the resolved
upstream is matched against the org's DenyRule rows (org-wide + the ones aimed at this caller) →
403 naming the rule. Evaluating the resolved upstream is what makes both call shapes equally
gated - a caller cannot dodge a rule by switching to URL-passthrough - and the relay does not follow
redirects, so a blocked host is not reachable via a 3xx bounce. The path match is anchored at a
segment boundary (/v1/charges must not match /v1/chargesX), the same trap _resolve_call guards.
It applies to every role including owner (a guardrail, not a permission tier) and to both run
tiers, where the tool's own base_url host stands in for the request path. _deny_match is pure, so
it unit-tests without a DB - mirroring localrun.check_deny, which is the same idea one layer down
(argv instead of URL). Zero rules = one indexed query and no behavior change. A rule may also carry a
project_id: it then fires only on calls through that project's tools (every enforcement point has a
resolved Tool by then, so _enforce_deny takes tool.project_id); an org-wide-tool call is never
caught by a project rule. The three scope axes - host/path/method, member, project - are ANDed and
each is NULL-means-any.
Responses and diagnostic evidence
Treg refusals on both call surfaces carry X-Treg-Error: 1. Mechanism-keyed application
failures map to caller | treg | upstream | org_connection blame; the HTTP adapter preserves
status and detail. Provider failures remain response data. Header contracts are documented in
the API fragment.
After credential refresh and relay, _audit records the attempt and mirrors it through
_tool_called_props / analytics.capture. Provider identity is the catalog provider or the
own tool's upstream host. Analytics includes outcome, status, timing, cost, call reference,
capacity/cache/smoothing signals and user-agent attribution, never params or bodies.
domain.catalog.results.classify shares verified hit/miss rules between business-hit telemetry
and cache admission. Found means hit=true, explicit empty means false, and errors/unknown
results mean null. Hunter company emails, LeadMagic employee finder, and SE Ranking keyword
volume have additional field checks; other verified adapters retain their miss expressions.
Only endpoints with enabled hit/miss rules adopt result-aware cache behavior; unconfigured or
unverified endpoints keep original cache learning and serving. Classification inspects only
already-buffered bodies and does not change relay bytes or settlement. Existing tool_called
events expose result_state, result_reason, cache_admission and cache_result_policy.
See archive result admission.
Overflow retains both attempt rows under the same call reference, but emits one product event for
the final answer. defer_analytics holds the parent's event until the child succeeds or the
parent's answer stands. Failed-request redaction and retention belong to
data-model.
Exceptional exits are recorded once, guarded by the audit marker:
| Exit | Event |
|---|---|
| Pool timeout before caller identity | call_intake_failed, surface only, no team/target |
| Pool timeout after identity | tool_called, outcome=gateway_failed, failure_kind=db_pool |
Unexpected exception while awaiting execute_call |
tool_called, failure_kind=unexpected_exception; Starlette still returns the bare 500 |
Before target resolution, target/provider identity remains null. Exceptions outside execute_call
and body-stream failures after the handler returns are outside this compensation contract.
Never relabel a stream failure as a new 500 after response headers have already been sent.
Resolution and relay guards
Resolution + error hardening: the URL-passthrough prefix match respects a path-segment boundary
(norm == base or base + "/"), so .../v1 no longer matches .../v10/... and inject the wrong
credential; the longest-prefix tiebreak compares rstripped lengths (a trailing-slash duplicate is a real
409, not a silent winner). When two same-host tools still tie on prefix length, _resolve_call
prefers the registry-provider-backed tool (one whose binding points at a Secret with a provider)
over a hand-registered one that often holds a stale credential - a 409 there would break exactly the
agent-facing URL-passthrough callers who never typed a tool name; only a genuine ambiguity (neither or
both provider-owned) still 409s. That 409 names every caller-usable colliding tool and directs the
caller to the unambiguous /call/<name>/<path> form. Binding validity is checked at registration (_validate_bindings rejects
an unknown injector and a cross-org/dangling secret_id; register_skill runs the same gate), and
call_tool translates a call-time injector ValueError and an upstream httpx.RequestError into a
502 instead of an unhandled 500 (and audits the failed attempt, not just successes). A binding
format is validated to render with only {secret} and name/secret_field to be non-empty strings;
duplicate target names are rejected independently for location:"header", "query", and "json"
(they would silently overwrite each other).
health._probe skips a dangling binding rather than KeyError-ing the whole run.
Relay security + faithfulness (bug-hunt): the response side strips a Set-Cookie for treg's own
cookie names (an upstream must not overwrite treg_session/treg_oauth_state - fixation) and adds
X-Content-Type-Options: nosniff + Content-Security-Policy: sandbox (a browser navigating to /call/…
must not execute upstream HTML/JS under treg's authenticated origin). It keeps Content-Length on a
bodyless reply (HEAD/204/304), only carries a request body when the caller sent one (no bogus chunked
frame on a GET), and honors headers a peer marks hop-by-hop via its Connection header (RFC 7230).
injectors._token_from_json rejects a non-string field value instead of injecting garbage.
Call-time SSRF guard (DNS-rebinding defence). Just before the upstream send, relay()
re-resolves the upstream host (infra.upstream.ssrf.host_is_public, gated by the proxy_ssrf_check setting) and
refuses with a 502 if any resolved address is internal (loopback/private/link-local/reserved/multicast).
This catches the case where a base_url was public at registration but its DNS now points at an
internal target like 169.254.169.254 or localhost - the registration-time check alone can't stop a name
that resolves differently later. Registration itself (infra.upstream.ssrf.safe_webhook_url, re-exported
by health and reused for base_url)
also rejects numeric IP encodings - decimal/hex/octal/short forms like 2130706433 / 0x7f000001 /
127.1 are normalized via inet_aton and re-checked, so they can't sneak past the literal-IP block.
Targets must be globally routable unicast addresses. CGNAT 100.64.0.0/10 is internal
service space: Tailscale, WireGuard overlay deployments, Fly.io, some Kubernetes pod CIDRs,
and Alibaba Cloud's metadata endpoint 100.100.100.200 use addresses in this range. These are
precisely the services a caller-controlled upstream must not reach. NAT64 translation prefixes
mapping non-global IPv4 addresses do not make those targets public; 64:ff9b::/96 remains blocked.
(A narrow resolve-vs-connect race remains; pinning the resolved IP would need a custom transport.)
Why relay instead of modeling the upstream: foundation/charter.md.
Routed endpoints - the resolve stage short-circuit
A catalog row with kind: routed (treg.<capability>, generated - architecture/catalog.md
§ Routing) never reaches the credential ladder itself. service._execute_call hands it to
application/call/route.py, which builds the plan and runs each child endpoint through this same
use case as a child CallContext (call_ref {parent}:r{n}), so every rule below - ladder,
reserve, relay faithfulness, capacity, overflow, audit, cancellation - applies per child
unchanged. Settle does not apply straight: each child's MarketplaceCall.deferred points at
the parent's CallContext.deferred_settles list, so _platform_settle appends a DeferredSettle
(call id, billable, the amount it would have charged, its archive-use marker) and leaves the hold
open instead of closing it. The parent - not each child - decides whether the caller pays, and
closes every child's hold exactly once with settle.close_deferred(charge=…): charge=True settles
each billable child as its own settle would have, charge=False releases all of them, because a
routed call that fails charges the caller nothing. A crash before close_deferred runs leaves the
holds to the reaper, which releases in the caller's favour. The parent only assembles
{output, raw, _treg} and owns the idempotency label.
A routed call whose outcome is still pending (202 with X-Treg-Route-Outcome: pending) leaves
request.state.call_cost_micro as None and skips the X-Treg-Cost-Micro response header rather
than reporting a charge that has not happened yet.
Platform capacity: refuse before reserve (plan step D)
Catalog platform_request checks and provider-specific request guards run only after a platform
offer is selected and before reserve. Exact selectors require evidence such as Tavily Search's
caller-supplied include_usage: true. Openmart requires its explicit bounded record count; Tavily
Map and Crawl require an integer limit from 1 to 20. Missing, Boolean and out-of-range values are
caller errors. BYOK is unchanged because the provider, not treg, bears that account's exposure.
Tier 4 spends treg's own vendor account, and that account can be empty. _resolve_marketplace_call
asks, after _platform_offer says yes: is this call exhausted in the in-process capacity view
(domain.capacity.view, loaded from ratestore on a 60 s TTL by resolve_marketplace_target before
its session opens)? Two sources say so: the sweep's capacity:state:<provider> (a balance API that
read zero) and the call path's own lock capacity:lock:<key> (below). If so it raises
CallFailure("provider_capacity", 503, blame="treg") - before any hold exists - whose body carries
resets_at when known and the same-capability alternatives from _capability_alternatives. treg still
does not choose for the caller (charter): it names the options. The audit row is refused_by="capacity",
X-Treg-Error: 1, cost 0. A stale, empty or "ok" view never refuses; only a confirmed signal does.
The call path's breaker (domain.capacity.marks) opens slowly and closes fast. After a tier-4
answer ≥ 400, settle._note_capacity_signal runs domain.capacity.signatures.classify on the
vendor's status/headers/body. A balance or quota signature is a strike; the second strike
within 10 min, at least 15 s after the first (a burst of concurrent calls hitting the same empty
instant is one strike), with no 2xx in between, locks - the provider for a balance signature, only the
endpoint for a quota one (allowances are per operation). While locked, resolve admits one real
call per process per minute as a probe (MarketplaceCall.probe_lock_id, probe on the
tool_called event); its 2xx clears exactly that lock (settle._note_capacity_recovery, conditional
on the lock id), any other answer leaves it. A guessed hold lasts 1 h, a vendor-stated reset at
most 6 h whatever retry-after said, and
the sweep never writes this namespace. Both writes run on their own short session after the
settle closed the hold, never during flight, and are the dataplane writes this feature adds
(capacity_exhausted_mark in tests/test_call_architecture.py). A burst 429 (retry-after ≤ 60 s)
or an unknown one only logs; step D′ smooths those. An edge_block (the vendor's CDN answered, not
the vendor: cf-mitigated, its HTML block page, or its 1xxx problem-JSON) strikes nothing: one caller's request shape must not take the provider away from every
other team. The kind rides the tool_called event as capacity_signal. Tiers 1/2 resolve earlier
and never consult the view: an org's own key running dry is the org's own answer, relayed
unchanged. The vendor's 402 on THIS call is also relayed unchanged - the protection is for the
next caller.
An unrecorded signal - a 4xx no row matched whose body still names credits/quota/balance -
is neither a strike nor a mark: it logs unrecorded capacity-looking … with the phrase and rides
tool_called as capacity_signal=unrecorded, the tripwire for a vendor whose out-of-credit answer is
not in the table yet (how Apollo's 422 went unseen on 2026-09-01).
Burst smoothing on treg's own keys (plan step D′)
Many callers share one platform key, so tier 4 makes its own bursts: leadsforge 429'd 27% of its
calls, crustdata 34%, with retry-after headers nobody downstream could act on. Two bounded
mechanisms in service._execute_call, both after the DB phase ended and before the relay (the
pool-discipline rule holds through the wait; proven by test), both platform-tier only, neither ever a
refusal:
- Spacer -
infra/upstream/limiter.py: one call perwindow_s / limitper provider (a token bucket of capacity one - a burst oflimitat t=0 is legal for a classic bucket and exactly what a sliding-window provider 429s). A call that would exceed the rate waits ≤ 2 s (DEFAULT_MAX_WAIT_MS), then proceeds regardless; the hold is already placed, so the org pays latency, never money. The limit comes from the capacity view (view.rate_limit: published by the sweep fromCapacityPolicy.rate_limit, with the verified defaults - leadsforge 120/min, leadmagic 300/min, crustdata 30/min, tikhub 30/s - before the first sweep). In-process on purpose: a second replica doubles the effective rate, and therate_pressurealert (step C) is the answer to that, not a shared counter on the request path. - One bounded
retry-afterre-send - on a tier-4 429 classifiedburstwithretry-after ≤ 5 s(SMOOTHING_RETRY_MAX_S), for a body-less GET/HEAD only: close the first response, sleep, send the identicalUpstreamRequestonce more on the same hold, settle on the second answer. A quota-429 (lusha "Daily", hunter "per billing period", anyretry-after> 60 s), an unknown 429, a POST, or a second 429 are relayed as is. The "no retries" rule for 401/402/5xx stands.
Both are visible: X-Treg-Smoothed: wait=<ms> and/or retry=1 on the response (metered exit only).
No audit column yet - smoothed_ms would be an ALTER on the hot callrecord table, a migration-class
change kept out of this behaviour PR.
Overflow - the child cycle (plan step E; off by default)
Overflow = the same vendor endpoint, another account of ours. When a tier-4 call fails on treg's
own key for a treg-side reason - a balance/quota signature, a burst-429 smoothing could not absorb -
and the worker has an enabled OverflowRoute for the endpoint, application.call.overflow. maybe_overflow runs a child cycle after the primary's settle released its hold:
- Route from the in-process route view (
domain.capacity.routes_view, Orthogonal first), skipping an aggregator marked unhealthy (overflow:<name>in the capacity view) or without a key; budget check againstOverflowSpend(overflow_daily_budget_usdper aggregator per day) on a short session. A deployment's live value belongs in its private operational configuration. - Child hold, own id
{call_ref}:overflow, through the ordinary_platform_reserve(tag budgets, daily cap, trial allowance apply; an empty balance is the normal 402). Never the parent's id: release-by-id is a conditional claim and_finish_cancelled_callreleases both ids exactly once. - One aggregator run with no DB open:
infra.upstream.aggregators.<name>.buildwraps the caller's original query + buffered body; the key comes fromSettings.overflow_key_<name>and is never logged. Monid's async runs are polled (bounded). parse→ vendor status + body + the real in-band cost._platform_settle(child, observed_override=cost, overflow_spend=(aggregator, cost − treg's direct price))charges exactly the aggregator's price, 0% markup, and folds the day's spend delta into the same transaction - the one allowlisted overflow write (overflow_spend_in_settle).- The vendor's body goes back as the answer,
X-Treg-Served-Via: overflow:<name>,X-Treg-Cost-Microthe child's charge,X-Treg-Call-Idthe parent's. Two audit rows share thecall_ref: the primary attempt with its real status and the child withcredential_tier="platform-overflow". The MCPcallresult carries the same disclosure asserved_viaplus a hint (an MCP client never sees headers), and/catalog/endpoints/{id}/catalog_getshow the route's price up front asoverflow_price_usd- seearchitecture/money.md§ Overflow money.
When the resolver already knows the account is out (the exhausted view) and a route is on, the
ladder skips the direct attempt entirely (MarketplaceCall.skip_direct): no parent hold, no vendor
402, straight to the child - the plan's tier 4b.
Influencers Club discovery/search and similar-creators routes, and Icypeas people search, have a
verified flat Orthogonal price against our per-row direct price. They are admitted by an absolute
fee ceiling (routes.request_priced, $0.03) instead of the price ratio, and route_for applies
no per-request check, so even a one-row request can overflow. The original paging/filter body is
relayed unchanged and the child settles once at the aggregator's reported price, even if the page
is empty. See ops/capacity.md for verification.
An aggregator failure is data. Its own 401/402/403 or a malformed envelope releases the child
hold and marks overflow:<name> unhealthy for everyone; a relayed vendor answer the signature table
reads as that vendor's own out-of-credit or quota dialect (VENDOR_DRY: a 402, Apollo's 422 through
Orthogonal's dry Apollo account, a period 429) releases the child hold and marks
overflow:<name>:<provider> only - one vendor's cap never takes the others offline. Either mark
lasts 15 minutes, and the caller gets the typed provider_capacity 503 with alternatives; a second aggregator is never tried on the same call.
The aggregator's own per-request refusal (contract: its stricter schema, or Orthogonal's bare
400/422/404 with no vendor data) is request-scoped - it releases the child, charges nothing and marks
nothing; the vendor's own answer stands, and on the skip-direct ladder, where there is none, the caller
gets the typed 503 naming the refusal rather than the aggregator's envelope dressed as the vendor's
answer. malformed is reserved for what is not an envelope at all (non-JSON, a 5xx, a transport
error); one validation 400 read as malformed once took Orthogonal offline for every org for 15 minutes.
Shadow mode (TREG_OVERFLOW_MODE=shadow): the aggregator is called, status / shape / cost logged
and the probe's cost recorded in OverflowSpend (treg pays, budget-bounded) - the caller still gets
the vendor's own error and is charged nothing. This is the week the plan requires before routes serve.
Never on tiers 1/2, a caller-caused 4xx, a 401, a timeout, PUT/PATCH/DELETE, a route the worker has
not enabled, or a team that opted out (Org.platform_overflow_disabled, treg org overflow off) -
checked before any aggregator is contacted, on both entry points.
Control-header isolation
proxy._is_dropped_request_header strips every request header beginning with x-treg-,
including future control headers. _CONTROL lists the non-prefixed infrastructure headers.
This prevents caller metadata, runtime identity and routing controls from reaching providers;
tests include an invented prefix-matching header so the guarantee cannot regress to an enumeration.
Asynchronous submissions on the call path
A catalog endpoint carrying an async descriptor resolves like any other (resolve.py freezes a
settlement_basis with when: terminal and the descriptor on the MarketplaceCall). In
service._execute_call, a metered 2xx from such an endpoint is deferred
(application.asynctasks.defer_submission writes the pending row and leaves the hold open) unless
_submission_rejected says the body is not an accepted submission (not JSON, expect rule failed,
no task id), in which case it settles at zero at once; a persistence failure releases the hold with
an alert. routers/call._attach_async_descriptor adds X-Treg-Async (the effective descriptor) to
the response, also on an idempotent replay, so a retried --await polls the task already running.
The settlement itself is the money fragment's subject.
Tier-4 calls add two checks before reserve and relay that deliberately do not apply to BYOK. First,
a body field used as a table-pricing discriminator and declared as a singleton catalog enum (for
example OpenRouter's model) must equal that fixed value; strict JSON parsing rejects duplicate keys
that could make validation and upstream interpretation disagree. Second, endpoints referenced as an
async descriptor's poll or fetch utility accept only task/result ids present in an
AsyncTaskRecord for the caller's org and provider, with the same frozen endpoint/parameter rule.
Legacy async pairs use catalog resource_ownership metadata and AsyncResourceRecord for the same
check without changing their existing settlement behavior. Formal submissions mirror their
poll/fetch ids into that table too, so removing a live descriptor reference cannot make its utility
fail open; pre-migration pending rows still authorize through their frozen descriptor.
Extended task consumers whose producer provenance is not modeled are catalogued as BYOK-only via
platform_blocked, rather than accepting an unverifiable task id on the shared account.
Unknown and cross-org ids receive the same 403 without contacting the provider. Fetch-mode result ids
are learned from an authorized successful poll or from the worker's terminal response. BYOK keeps its
faithful-relay semantics because those ids belong to the caller's own provider account.
Durable shared-account objects use the separate catalog managed_resource contract and
ProviderResource table. Create relays with no DB connection held, then commits ownership before the
provider id is returned; persistence failure triggers best-effort provider deletion and returns a
treg 502 without exposing the id. Update changes local display state only after upstream success.
Delete authorizes active or tombstoned ownership, treats an owned upstream 404 as deleted, then
tombstones locally, making retries safe. These rules run only on the platform-key tier; own keys keep
the ordinary faithful relay. A managed use declaration may permit provider-public ids through a
bounded GET predicate. The local ownership check runs first; cross-org and tombstoned ids are denied
without upstream I/O, while wholly unassigned ids are verified only after the DB phase closes and
before money is reserved. Managed responses remain under the same 8 MiB complete-body limit.
An owned platform status poll with an explicit free price and zero estimate takes the
MarketplaceCall.free_owned_poll branch. It bypasses a new poll reservation and settlement while
buffering the response for observe_owned_poll, which learns fetch ownership and finalizes the
original task on terminal 2xx evidence. Missing required usage leaves the hold pending; settlement
errors preserve the provider response for cron recovery. Polls read the live provider, do not use or
populate replayable cache, and release any caller-supplied idempotency label so the next poll
can observe a changed status. Successful and failed polls retain diagnostic audit rows with
kind=async_poll and zero charged cost; /calls excludes them before pagination. The original
submission shows the shared finalizer's settlement state and result in Activity. Terminal evidence
is archived under that submission's call id, not the poll's id.
Complete downloads and bounded response evidence
MarketplaceCall.streamable_free_result identifies platform fetch utilities with an explicit free
price, zero estimate, a required fetch ownership rule, and no async submission, owned poll or produced
resource evidence. Only successful GET responses bypass _buffer_response; the normal ownership
check still runs before relay. Their existing zero-amount reserve/settle gates stay in place, while
the full upstream stream and headers (including Range metadata) pass to the router's close-once
lifecycle. The original generation task is not observed or finalized by the download. Audit byte
size is unknown, not a fabricated zero. No body is archived or retained for idempotent replay:
these free reads release the claim, so a retry fetches the provider again.
Other responses needing settlement or ownership evidence have a hard 8 MiB complete-body budget.
_buffer_response raises GatewayFailed(response_buffer_limit) on the first overflowing chunk,
before any success headers, and closes the upstream in finally on EOF, exception or cancellation.
The application releases the hold/claim and returns 502 with zero cost; no partial response reaches
settlement, archive or replay. This is an explicit size limitation, not support for arbitrarily large
metered JSON. The fault is attributed to treg's buffer limit, not to the provider. Own-key streams
remain outside this limit. tests/test_call_response_limits.py exercises both real HTTP hops,
CLI output, boundaries, Range, disconnects, settlement evidence, archive and replay behavior.
Spooled evidence for inline media
Some providers return generated media inside their JSON answer: Gemini's generateContent puts
a base64 image and a multi-megabyte thoughtSignature in the body, about 9 MB at 2K and 23 MB at
4K, with its token meters (usageMetadata) after them. Streaming that answer would mean settling
after the response is sent (no exact X-Treg-Cost-Micro, a second close-once path for disconnects,
routed children and overflow rebuilt); raising the buffer would put tens of megabytes per call in
a web process. The endpoint instead declares spooled_response: true and
_read_evidence hands its metered 2xx to _spool_response:
- The body streams into
tempfile.TemporaryFile(unlinked from creation, so a crashed worker leaves nothing on disk) underspool_max_bytes(64 MiB) and a per-processspool_budget_bytes(512 MiB) shared by concurrent spools. A declaredContent-Lengthis claimed whole before the first read, so concurrent answers are admitted or refused whole rather than all stalling half-read; without one, the claim grows chunk by chunk. Crossing either limit raises the sameresponse_buffer_limitbefore headers (the budget refusal says it is temporary); the upstream closes infinallyas in_buffer_response, temp-file creation included. - The file is parsed once with the stdlib
json(measured on a 23 MB Gemini answer: ~30 ms and ~45 MB peak; about three times the body at worst) in a worker thread, at mostspool_parse_concurrency(2) per process. Only the paths the row's settlement reads survive (_spool_evidence_paths: each usage term's top-level object and theexpectleaf, kept at its original place, e.g.{"candidates": [{"finishReason": "STOP"}]}), re-serialized as thebodyevery later consumer sees: usage settlement, result classification and capacity signatures. A body that is not a JSON object yields empty evidence; ausagebasis then settles at its reserve. - Settlement then runs exactly as for a buffered body, before the response starts, and the router
relays the file in 256 KiB reads with the provider's headers and a recomputed
content-length. The replay's close returns the budget;weakref.finalizedoes so too for a response dropped without closing. - A spooled body is not archived and not stored for idempotent replay; its audit row records the
full size. Non-2xx answers, own-key calls and routed children (whose parent reads the child's
body) keep their existing paths.
tests/test_call_spool.pycovers the lifecycle (oversize, budget, reset, cancellation, garbage collection, parse gate) andtest_call_response_limits.pya 20 MB Gemini-shaped answer over two real HTTP hops.
HarvestAPI integration
Catalog entries can opt into strict_query: _enforce_catalog_query rejects bodies,
undeclared/duplicate query parameters, missing required inputs and unsupported enum values before
credential selection. They can separately opt into body_allowlist: _enforce_catalog_body
rejects undeclared top-level JSON fields, missing required fields, invalid declared scalar values,
and arrays outside their declared cardinality or item enum. An omitted optional array is allowed;
its cardinality applies only when the caller supplies it. Both checks apply to catalog calls on every
tier, leave unmarked entries unchanged and do not constrain arbitrary raw own-tool relays. Legacy
resource_ownership.requires accepts a declared body
parameter as well as path/query parameters, allowing a POST status utility to authorize an opaque
shared-account task id.
Pinned shared-provider reads
_enforce_platform_async_ownership adds pinned_tag_predicates to its org-scoped task and resource
queries. All pins must match the submission snapshot; a pinned unknown/foreign id returns 404 before
relay. BYOK and raw own-tool access keep their existing credential ACLs. The shared-provider
Idempotency-Key rewrite additionally includes the complete enforced pin, so different customers
cannot receive one upstream job through provider deduplication. Unpinned digests are unchanged.
intake.prepare_call_intake includes the same full pin in Treg's membership replay namespace. See
multi-tenancy for the history and ledger scopes.