A metered answer is buffered whole before headers so settlement sees complete evidence, and
anything over 8 MiB fails uncharged. Providers that inline generated media in JSON (Gemini
returns a base64 image plus a large thoughtSignature: ~9 MB at 2K, ~23 MB at 4K) can never fit.
Streaming instead would settle after the response starts, losing the exact X-Treg-Cost-Micro
and adding a second close-once path for disconnects, routed children and overflow; raising the
buffer would hold tens of megabytes per call in RAM.
An endpoint now declares `spooled_response: true`. Its metered 2xx is written to an unlinked
temp file under `spool_max_bytes` (64 MiB) and a per-process `spool_budget_bytes` (512 MiB),
claimed whole up front when the provider declares a Content-Length;
crossing either is the existing uncharged `response_buffer_limit`. The file is parsed once with
stdlib json in a worker thread (a 23 MB answer: ~30 ms, ~45 MB peak) behind
`spool_parse_concurrency`, and only the top-level keys its usage paths start at
(`settlement.usage_roots`) become the body every settlement consumer reads, so the evidence
cannot drift from the price. Settlement runs before headers as before; the router then relays
the file byte for byte, and its close (or garbage collection) returns the budget once. Spooled
bodies are not archived or kept for idempotent replay. Errors, own-key calls and routed
children keep their existing paths. `_buffer_response` and the spool share the refusal text and
the content-length rewrite.
The validator requires settle: usage and refuses the field beside async, resource_ownership or
managed_resource. AGENTS.md non-negotiable 4 records the new path.
Fragments: architecture/proxy-model.md, money.md, archive.md, catalog.md, interface/api.md,
ops/deploy.md.
`settle: usage` read one response path times one unit. A provider that reports token meters
and no charge (Gemini's usageMetadata: prompt, candidates, thoughts, and image tokens inside a
per-modality list) could not be priced from its own answer.
`usage.terms` lists {path, rate} pairs in USD per unit; the settled figure is their sum. A
path segment may select a list item by key (`candidatesTokensDetails[modality=IMAGE]`), since
per-modality entries carry no order. An absent meter counts as zero (proto3 JSON omits zero
fields); a response with none of them is unobserved and settles at the reserve, and a
malformed meter poisons the figure instead of billing less. The validator requires unit `usd`,
positive finite rates and well-formed paths; the CLI price view names the meters.
* fix(billing): quickenrich email miss free, icypeas bulk billed per row, icypeas reads scoped to the team
- quickenrich.people.email.find: QuickEnrich bills a phone-only answer (credits_used=1), but the
adapter calls it an email miss; settle it at zero so a routed miss is free and no longer eats
the route's max-cost budget before the next provider.
- icypeas.bulk.search: the start answer has no rows, so the reserve is the bill; it defaulted to
a 20-row page. Reserve data rows x the task's rate (0.1 credit per verification).
- icypeas.search.results.read / search.files.read listed every search on treg's shared Icypeas
account. Searches and bulk files now record ownership; reads require an owned id and reject the
listing modes. Bulk rows move to icypeas.bulk.results.read.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(icypeas): a single email search reserves one credit, not a 20-row page
Its answer is {item: {_id}} with no rows, so the reserve is the bill; the 20-row default charged
$0.38 for one lookup. Live: $0.38 before, $0.019 after.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(catalog): keep Apify actor starts off treg's shared key
apify.web.scrape.job.start is priced free because the run bills later by the
actor's own pricing, and nothing on the shared-key path meters that run. On
treg's key it let any caller run any actor on treg's Apify account at no
charge. It now needs the team's own Apify key; job.status and job.results
stay open because per-team ownership already scopes them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(catalog): add Weibo, TikTok Shop, Lazada and Google Maps Apify actors
Combines #641-#644 into one catalog change: nine run-sync entries over four
herus13 actors, plus the lazada platform. The six Weibo entries serve on
treg's key at a price observed on it. Lazada, Google Maps and TikTok Shop are
own-key only: their actors bill run compute or a per-GB start fee that a
per-result price cannot meter.
Co-authored-by: Herus13 <bootforge.ai@gmail.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(catalog): require own keys for Weibo actor runs
* test(frontend): isolate landing interactions from continuous WebGL rendering
* feat(call): let platform_request pin query parameters
Some upstreams take their spend bounds as query options rather than body
fields (Apify's maxTotalChargeUsd, memory and timeout run options). A
queryParams pin must be sent exactly once and is compared as the pinned
value's type; own credentials keep the upstream contract.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): bill Apify platform calls by returned dataset rows
run-sync-get-dataset-items answers a bare array, so every Apify per_result
call settled at its estimate whatever it returned. Count the rows, add an
optional per-row-independent call_fee for the actor's start or compute
charge, and require maxItems (1-200) on the platform key so the hold is the
worst case. Apify's usageTotalUsd lags a finished run by minutes, so the
response body is the settlement evidence.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(catalog): serve the herus13 Apify actors on treg's key
Weibo, Google Maps, TikTok Shop and Lazada now settle on the platform key by
counted dataset rows plus a flat call_fee (actor start, or Lazada's run
compute). memory and timeout are pinned so the fee is fixed and a run ends
before Apify's synchronous wait; TikTok Shop is held to keyword search.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): bound Apify platform calls by maxTotalChargeUsd
maxItems does not bind pay-per-event actors whose own input sets the row
count, so a one-row hold could settle thousands of rows. Require the
maxTotalChargeUsd run option Apify enforces (at most $1), hold it plus
call_fee, accept only the run options each once in plain ASCII, and bill the
hold when a run reaches its cap. Pin meta-ads enrichment off and bill the
LinkedIn actor-start event on treg's key.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(catalog): cap herus13 Apify runs by maxTotalChargeUsd on treg's key
Row notes name maxTotalChargeUsd as the enforced spend limit; Lazada's
compute fee is 0.015 under a 180-second timeout pin.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(catalog): name maxTotalChargeUsd as the Apify spend cap
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): treat an Apify run within two rows of its hold as capped
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(call): require a short timeout and a three-row cap on Apify platform runs
Past Apify's 300-second synchronous wait a run answers 408 and keeps billing,
so every Apify per_result platform call now names timeout <= 280. A cap under
call_fee plus three rows would bill an empty answer in full, so it is refused.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(call): keep Apify platform runs inside every wait
treg's upstream read timeout (call_timeout_s, 180 s) and the MCP client's
120 s end the call before a 280 s run finishes, releasing the hold unbilled
while the run keeps billing. Bound timeout to 90 s and 30 s under
call_timeout_s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(catalog): pin herus13 Apify runs to a 90-second timeout on treg's key
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): bill a timed-out Apify platform run at its hold
A run that outlives its own timeout answers 400 run-failed with no rows, but
Apify billed its events up to the caller's cap and the run id in that body
reads the dataset. The caller chose the run's size and timeout, so the hold
settles instead of releasing. The minimum-cap check compares micro-USD.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(catalog): tell Lazada callers to keep platform runs small
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): release timed-out Apify runs; the account must stay Restricted
Billing the whole cap on a TIMED-OUT 400 overcharged callers: a run that
timed out after its start event cost Apify $0.00005 and would have billed the
full cap. The loss treg absorbs stays bounded by the $1 cap, and keeping the
Apify account's resource access Restricted stops anyone reading the unbilled
run's rows by id. Examples now show the 90-second platform timeout.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(money): name disconnects as a bounded Apify loss; call_fee wording
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(catalog): record Lazada's 0.05 minimum cap and 10-product floor
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(asynctasks): keep the Bright Data and CompanyEnrich platform keys the merge dropped
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(catalog): verify every Apify per-result price on the platform key
Each row's test_request ran on 2026-09-26 and Apify's chargedEventCounts
matched its rate card. The TikTok ad library actor returns rows again, so it
is verified with a captured example.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(catalog): route Google Maps and Weibo post detail through Apify
Adapters let treg.google.serp.maps and treg.weibo.post.detail choose the
Apify actors, with the run's spend cap, timeout and memory fixed so the
platform guard accepts the child call.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(catalog): drop Lazada's compute fee and keep account details out of notes
The actor's 2026-09-25 pricing no longer bills run compute to the caller (a
79-second run showed no platform usage), so its call_fee over-charged every
call. Notes describe observations by price tier, not by the account that
made them, and carry no run ids.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(catalog): state tiered Apify observations without the account's tier
Notes for plan-tiered actors record the events billed and that they matched
the rate card for the key's tier, not the dollar figure that would name it;
the repeated Weibo observation and the owner-meter notes go.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(settle): bill LinkedIn's actor-start once per query it runs
The LinkedIn jobs actor bills its actor-start event for every job title x
location searched, so a flat call_fee under-billed any multi-query call.
cost.call_fee_per names the body arrays whose lengths multiply the fee. The
apify.yaml header no longer names a plan or calls the TikTok actor broken.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(money): describe call_fee_per
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(catalog): accept only declared fields on LinkedIn job search
The actor bills an actor-start for every geo id it searches, and geoIds was
undeclared, so it passed through unbilled. The row now takes its declared
filters only (salary, easyApply, under10Applicants and industryIds added);
places go in locations, which call_fee_per counts.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(catalog): leave headroom above the Apify spend cap
A run whose charges land exactly on maxTotalChargeUsd can be aborted by
Apify and answer 400 with no rows, which releases the hold. Lazada's
10-product floor costs exactly its 0.05 minimum cap, so its example now uses
0.06 and the notes say to set the cap above the expected spend.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Herus13 <bootforge.ai@gmail.com>
* feat(auth): scope pinned agent history and async reads
* fix(auth): a routed parent row carries the pin; reviews use the pinned read scope
`_audit_parent` in `application/call/route.py` wrote the successful routed
parent audit row without `_tag_telemetry`, so once history is filtered by
pin a customer's own routed calls vanished from `/calls`, its result 404'd
and `/calls/{ref}` returned `call: null`. The refusal path already carried
the tags; this aligns the success path.
`POST /reviews` resolved its call by org alone, so a pinned caller could
rate, and thereby probe, another pin's call. It now applies the same
predicates as `/calls` and answers 404 for a foreign reference.
One test covers both: a pinned routed call is listed and readable by its
pin, invisible to another, reviewable only by its own.
Fragment: architecture/multi-tenancy.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(auth): refusals carry the pin; feedback reports read through it
Two omissions an independent review turned up.
The router's refusal fallback (`routers/call.py`, the one audit writer
outside `service.py`) wrote its row with no tags, so a pinned caller
could not see its own 404s and 403s — the pin-mismatch 403 above all,
the row a builder most needs while integrating. The handler now stashes
the pin next to the identity and the fallback records it.
`GET /feedback/{id}` was org-scoped with sequential ids, so a pinned
token could read another customer's free-text report. `Feedback.tags`
(revision 0042) snapshots the reporter's pin; the read applies the same
predicates as `/calls`, and a pinned reporter's `call_ids` verify only
against its own pin's rows. Unpinned behaviour is unchanged.
Fragment: architecture/multi-tenancy.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Jason Zhou <jason.zhou.design@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
SQLModel 0.0.45 rejects naive datetime bind parameters and may return
aware datetimes from the DB, breaking three crons:
- treg-asynctasks-settle: ValueError on WHERE next_check_at <= :now
- treg-catalog-stats: TypeError on created_at >= since comparison
- treg-arena-insights: ValueError on INSERT scan_until bind
Short-term fix: pin sqlmodel>=0.0.22,<0.0.45 → lockfile uses 0.0.44
Long-term prep: NaiveUTC annotations in place for eventual upgrade
Changes:
- Pin sqlmodel>=0.0.22,<0.0.45 in server extra
- Import NaiveDatetime from pydantic, define NaiveUTC type alias
- Update all datetime field annotations in models.py to use NaiveUTC
- Add regression tests for all three affected crons
Fixes production crashes in settle, catalog-stats, and arena-insights.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Jason Zhou <JayZeeDesign@users.noreply.github.com>
Through the real call path and the settle worker. Against the unfixed source the same
request is refused for the thirty-second 1080p ceiling.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
reAPI classifies a Seedance prompt that edits a @video reference as video
editing and then requires `duration: -1` (auto). The price table bounded
`duration` to 4-30, so -1 matched no row and priced at the fallback: thirty
seconds of 1080p. With `settle: table` the frozen request re-prices
identically at the terminal state, so that ceiling was the bill, not just the
hold, whatever resolution and length were actually rendered.
- Both reAPI Seedance rows settle on the terminal body's `usage.credits`.
`usage.unit: credit` is new: the credit's micro-USD worth comes from the
provider's fx.yaml `credit_rates_usd` entry and is frozen into the basis at
reserve. A credit basis with no frozen rate settles at the reserve, never
credits read as dollars.
- `duration: -1` is listed and priced by flat rows ahead of the per-second
rows: thirty seconds at the requested resolution, as a reserve only.
- A `times` multiplier is never non-positive, whatever minimum the field
declares, so admitting a sentinel cannot multiply a rate by it.
- The advertised per-second rate span skips a flat row that pins the duration.
- The validator accepts `credit` only when fx.yaml prices that provider.
Fragments updated: architecture/money.md, architecture/catalog.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- cost_view serves rate_usd_min / rate_usd / rate_unit for any table whose rows all
multiply by a duration field; the dashboard and CLI quote $0.119-$0.462/s instead of the
whole-call $0.47-$13.9/success range. usd / usd_min are unchanged for reserve and eligibility.
- Merged-row titles for the reAPI / PiAPI per-model keys are the plain model name.
- openrouter.yaml curates bytedance/seedance-2.5 (official rate card, settle: usage) on the
shared video-gen.seedance-2-5.generate key, so the default-filter row compares three routes;
the generated extended twin is removed, as the ingester's curated-model rule would.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SKjtrwx9eJ5Ewap8Fesp4J
Every org shares one provider account on the platform tier, and a provider
that honors Idempotency-Key (LeadsForge does) returns the FIRST org's job to
any other org that sends the same label. With this PR's
resource_ownership.produces rules that answer also made the second org the
recorded owner of the first org's job: reproduced live 2026-09-09, org B
read org A's enrichment results and was settled for them.
The relay now replaces the header with a digest of (org, label) before a
tier-4 request is built, so two orgs can never collide on the shared
account. Callers lose nothing: treg's own idempotency table already replays
their answer for the same label. A team's own key still relays the header
verbatim.
This is a fourth rewrite in the relay's faithfulness contract, so
non-negotiable 4 in AGENTS.md, the proxy-model fragment and the catalog
fragment move with it; the catalog fragment also lists LeadsForge among the
resource_ownership providers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
PROCESS_TIMEOUT_S wraps the claim's DB round trip as well as the poll. At
10 ms a loaded CI runner fired it before the claim, which the worker
correctly reports as backed_off with the row untouched - and the test,
which wants the post-claim poll_error path, failed on
`consecutive_failures == 1` (CI run 34036769106). The hung poll never
returns, so 0.5 s still cuts it; POLL_TIMEOUT_S keeps 0.01.
Two money regressions from the settlement-basis refactor, found by /code-review:
- An overflow child carried the PARENT's settlement basis, so an aggregator that reports no cost
settled at the parent's direct price (or table ceiling) instead of the aggregator reserve.
`_child` now carries its own observed-kind basis at the aggregator price. Regression test.
- The worker caught only httpx/JSON/RuntimeError; relay() raises GatewayFailed (an unset platform
key, an SSRF refusal), which escaped `_process` and, through a bare asyncio.gather, aborted the
whole tick every run while the reaper deliberately left those holds alone. `_process` now backs
off on any exception and `settle_due` gathers with return_exceptions. Regression test.
Also:
- A 2xx from an async endpoint that is not an accepted submission (not JSON, expect rule failed,
no task id / off-allow-list poll URL) settles at zero on the request path instead of becoming a
24-hour pending row (`_submission_rejected`); defer_submission no longer has a fail-closed branch.
- The CLI reads the submission through the domain's extract_submission (task id and the https +
allow-list rule the worker applies) and no longer calls .json() unguarded.
- reconcile.async_task_settlement bounds its query (pending rows or completed since the window).
- Activity artifacts are looked up by the archive's indexed key hash, not the unindexed req_url,
and a pruned carrier snapshot no longer raises.
- One dotted-path reader: settlement and settle.py use domain.asynctasks.json_path.
- Em dashes on lines this branch added are plain dashes (user house rule); plugin skills rebuilt.
- proxy-model.md documents the async branch of the call path; money/archive fragments updated.
sqlite 2546 passed; Postgres subset 240 passed (local postgres:16).
- Two workers racing one row after its lease lapses move money exactly once (Postgres FOR UPDATE
SKIP LOCKED; verified locally against postgres:16).
- A request cancelled at the pending-row commit boundary leaves a coherent outcome: the worker
now records a row whose hold was already closed as released instead of settling it at zero.
- Another org cannot read a task by its call_ref (views_for, /calls/{ref}, /calls).
- A descriptor cannot point at a retired or broken utility row (validator).
Findings from a read-only review of the whole branch against main, each with a regression test:
- A `times` multiplier is bounded by the input schema's min/max (finite, positive without a
declared minimum): `duration: 0` no longer reserves zero and `duration: 100` no longer bills
past the validated ceiling; an out-of-range value matches no row and prices at the fallback.
- The worker takes terminal evidence only from a 2xx poll. A 404/401 body that happens to carry
"status": "succeeded" backs off like any other error, the rule the CLI already applied.
- A static poll substitutes its parameter by the declared location (never a coincidental
placeholder) and percent-encodes a path value; the validator requires a pathParams id to have
exactly one `{name}` in the target path and a queryParams id none.
- One shell-safe formatter (`domain.asynctasks.fetch_command`, shlex-quoted) builds the retrieval
command for the CLI and the Activity API; provider strings shown on the terminal go through
`shown()`, which escapes control characters and terminal escape bytes.
- An idempotent replay of an async submission carries `X-Treg-Async` again: the router resolves
the endpoint from the path when the replay short-circuits marketplace resolution, so a retried
`--await` polls the task that is already running instead of printing the body and stopping.
- `settle: usage` requires an async descriptor (a synchronous response path has no consumer for
usage evidence); `async.interval` and numeric `when` values must be finite.
- The price floor reads nested input fields, so Replicate's `input.duration` minimum counts
(seedance-1-lite from $0.072, not $0.018); poll bodies are capped at 2 MiB.
- Dashboard: the timed-out tooltip says the hold was refunded and treg absorbed the charge; the
catalog fragment no longer describes the pre-change usage reserve.
2479 tests pass.
A `settle: usage` row reserved the matrix ceiling, so a $0.05 Wan 3.0 call demanded a $6.00
balance and a $1.00 team could not send any OpenRouter video. It now reserves what the table
says this request costs (the same figure the catalog shows) and settles the provider's reported
usage.cost, which may exceed the reserve; the ledger already takes the difference from the
balance and the next reserve is the gate. Live: Wan 2 s 480p reserved $0.10, settled $0.2125.
The two places that gap can land are now listed in reconcile.async_task_settlement and the
/admin/reconcile/asynctasks route: `overruns` (the team paid more than reserved, aggregated per
endpoint with the max ratio, which is where an unpublished minimum charge becomes data) and
`absorbed_shortfalls` (settle entries whose block_shortfall_micro > 0: the platform ate what
the team's blocks could not cover). A success whose terminal response carries no usage figure
settles at the reserve with reconcile_review and an ERROR alert, not at the ceiling.
Rate-card notes, the catalog get PRICE TABLE footer and the /access line say the matched row is
held and the reported cost is paid; money.md and the surface snapshot updated with the code.
The worker appended `?task_id=…` to the poll URL, but the relay composes the upstream query
from `request.query_items` alone, so MiniMax v1 polls arrived without the id and the provider
answered 2013 "invalid params" on every tick until the 24-hour deadline. Path-parameter polls
(H3, Replicate) and dynamic URLs carry nothing in the query, which is why it hid; a live Hailuo
2.3 task exposed it and settled at $0.28 on the first tick after the fix. Regression test pins
the composed request.
A random sample of six generation endpoints through the real CLI on treg's key found three
defects and three catalog lies; all are fixed here and the same six were re-run clean.
- Replicate polls a static catalog id. `--await` polled the absolute `urls.get` URL through
`/call/https://…`, which resolves only a team's own tool, so every team on treg's key got
"no registered tool for upstream 'api.replicate.com'" and never saw its result. Replicate also
publishes the stable `GET /v1/predictions/{id}`; it is listed as `replicate.predictions.get`
and the provider descriptor uses it. The dynamic-URL mode stays in the schema and worker for a
provider that offers nothing else (none listed today); catalog.md says it is BYOK-only until a
host-allow-listed relay exists.
- A 2xx that fails the endpoint's `expect` rule is a failure, not a task. MiniMax answers HTTP
200 with base_resp.status_code 2013 for a bad parameter; that was deferred as a task with no id
and charged the reserve. `_submission_accepted` gates deferral on the rule, and a
provider-reported zero (miss, failed envelope) now outranks a frozen price-table basis at
settle, so the caller pays nothing. Both MiniMax v1 rows carry the rule.
- Persistence failure of the pending row releases the hold (reason async_task_not_recorded) with
an ERROR alert instead of settling the basis: treg's own failure is treg's cost, the same
doctrine as the 24-hour deadline. The sample had charged a $4.80 ceiling this way.
- The CLI hands the submission body to stdout when it carries no task id, so the provider's
own "invalid params, …" is the diagnosis instead of a bare treg error.
- Catalog truth from live traffic: MiniMax-Hailuo-2.3-Fast is image-to-video only and 512P
needs first_frame_image, so the text-to-video row loses both (its own price table, example
fixed); Veo 3.1 Fast on Replicate works from text, so its capability is `.generate` and the
example no longer demands an image.
Fragments (catalog, money) updated with the code; 2469 tests pass.
- A price table prices out as a range everywhere. The store computes the cheapest row into
cost.table_min at load time and cost_view exposes usd_min beside usd (still the validated
ceiling that reserve and eligibility read). The wall, `treg catalog search`, the dashboard
and /access now show "$0.2-$1.96/success" for MiniMax H3 instead of the $1.96 ceiling alone,
which had made direct routes look several times dearer than they are.
- The 24-hour deadline releases the hold instead of settling it at the reserve: an outcome
nobody observed is the platform's cost, never the customer's. The row is flagged for
reconcile (`absorbed_timeouts`) and the worker logs an ERROR-level alert naming the endpoint,
since the usual cause is a provider that silently changed its status field.
- One reading of the async descriptor. The CLI's --await now imports json_path,
classify_terminal and artifact from treg.domain.asynctasks (stdlib-only) instead of carrying
its own copy; the module joins test_import_lightness and gets an import-linter contract that
forbids every server root, so the light install can never grow a database import that way.
artifact() returns the fetch target as {endpoint, name, value}; callers format the command.
- An endpoint's async block replaces the provider default whole (effective_async_descriptor);
the recursive merge with mode pruning had one user, which overrode everything anyway.
- The usage unit "token" is gone from the validator and settle() until a metered token-priced
listing exists with its fx rule and a live test; its unit derivation was wrong (it ignored
`per`) and no real traffic had exercised it.
- AsyncTaskRecord.request_data dropped: settlement_basis already freezes the request with the
rule. The migration is unreleased. The dead observed_override in the defer-failure branch
is gone; that branch settles the frozen basis, which is what it always did.
Fragments (catalog, money, import-boundaries, cli) updated with the code; 2468 tests pass.
The audit row froze a metered async submission's reserve as its charge, so Activity
kept showing "charged $x" for tasks that were still running or had been refunded.
`/calls` and `/calls/{ref}` now join the org's AsyncTaskRecords (application.asynctasks
.views_for), report `async_task` (state, reserve, settled amount, completion time) and
rewrite the charge to what actually hit the balance: null while pending, the settled
figure (0 after a refund) at a terminal state. The artifact is derived from the archived
terminal JSON through the descriptor (domain.asynctasks.artifact): the result URL for
path-mode rows, the exact `treg call` retrieval command for fetch-mode rows, plus the
descriptor's ttl_note. archive.load_terminal_responses reads the evidence back; treg
still never follows or stores media.
Dashboard Activity renders the state chip (generating… / done / failed · refunded /
timed out), "hold $x" while pending, and a `result ↗` link or `result via CLI` chip with
the expiry note. `treg audit` carries the same summary under `task`.
Docs: skill.md and llms.txt gain the "generate a video or an image" task now that the
mechanisms are built and verified (async descriptor, --await, shell-timeout warning for
CLI agents, lazy polling for MCP agents, reserve-then-settle money, expiring URLs);
USAGE.md documents `treg call --await`; plugin SKILL.md copies regenerated; fragments
for api, dashboard, cli, archive, money and the skill updated with the code.
Verified live on the demo server: a $0.003 FLUX Schnell submission showed
"generating… hold $0.003", and after one worker tick "done $0.003" with the result link.
Add the expand-only async task record and data-derived settlement basis, defer tier-4 holds until provider terminal state, and complete them through a multi-instance-safe worker. Archive terminal JSON, expose read-only reconciliation, add AIGC platform key slots and Render scheduling, and pin SQLite/Postgres behavior with E2E coverage.
The table migration ships with behavior because older code ignores it while the new request path cannot retain a hold safely without the durable record.
Live Postgres verification: Replicate FLUX Schnell settled a no-key team success at 3000 micro-USD. A real Seedance task accepted with an unreachable input image reached terminal failure and released its 72000-micro hold in full. Both terminal JSON responses were archived; catalog-priced provider spend was USD 0.003 total.