164 Commits
Author SHA1 Message Date
SToneX 0fd8c70c67 feat(search): serve the job-first answer to agents (search_experiment v2)
MCP catalog_search answered from the lexical ranker with the v1 judge's say
over it. The judge recovered a quarter of the searches the lexical gate
admitted nothing for and barely moved the rest: its page was the lexical
candidates filtered and reordered, with no vendor the words did not reach, and
an empty page told the agent to try other words, so half of all searches were
followed by another search. The engine behind /catalog/find (recall by job and
by meaning, one judge request, every vendor of a fitting job, a verdict that
can say the catalog lacks it) now answers agents in a new experiment mode.

- search_experiment gains mode v2: arm_for deals v2 as the majority arm, the
  same two holdouts keep the pure lexical and the pure v1 judged page, so v2
  is read against what it replaces. Arms are dealt by team and email where the
  search resolved them (an OAuth token rotates hourly), else by token.
- catalog_find: decide takes the verdict for its rule 8 (a person gets none,
  an agent keyword: its input always means something), an abstain keeps the
  judge's reason, expand returns its groups (expand_groups) and answer_v2
  splits into judge_and_decide and lay_out, so a holdout that only records
  lays out nothing. store gains routed_discovery_on and routed_parent, the
  one reader of each where four were.
- application/catalog_search lays the answer out for an agent (agent_page):
  the best limit // 2 jobs, rows dealt round-robin and laid out job by job,
  members the judge rated on their own leading their job, a routed parent
  leading a strong or closest job's group, a listed hub tool joining its job
  with no lexical gate, and per job the vendors the page left out.
- The verdicts an agent sees: strong, closest, name, and none only for a
  catalog gap (an empty page that says so, with catalog_request and no near
  misses). Not-a-task, an abstaining judge, a failure or a caller past the new
  per-caller cap (search_judge_max_per_caller_hour, the only bound on what an
  unmetered search can spend; it guards every dealt caller before any judge)
  serve the lexical page under keyword.
- The response adds verdict, reason, jobs, and per row job and more_providers;
  score is null on a judged answer (no probability reaches an agent); the hint
  says only what to do next. Both MCP surfaces answer alike.
- SearchLog: every served row records its job in every mode, so the report
  credits a call to any vendor of a job the page showed (a v2 hint sends the
  agent there); v2 rows carry find's readings and the verdict; the lexical
  holdout still has v2 judged and recorded, the counterfactual on a false
  none. The report gains conversion by job per arm and verdict, calls after a
  none, and the re-query rate per verdict.
- scripts/search_agent_bench.py scores the v2 answer offline against searches
  agents made and what they called next (job-hit, hit@limit, false-none by
  arm), beside the pages the log served, on find_bench's harness (the judge
  and the query vectors cached on disk, the card vectors warmed once).
- find_index: stored vectors decode as arrays, the index builds off the loop.

The HTTP route and the CLI still answer from the lexical ranker; they follow
once the route's hub read holds no session through a judge call and the CLI
sends its token.

Fragments: search-experiment.md, find.md, catalog.md, mcp-oauth.md; llms.txt
and skill.md (and the generated SKILL.md copies) say what the verdict means.
2026-10-01 20:19:52 +08:00
shehjad-dev ca1873bf79 feat(catalog): add Enrichlayer BYOK and platform tools 2026-10-01 02:24:03 +06:00
shehjad-dev 2293afb736 feat(catalog): add Search1API provider 2026-09-30 23:50:08 +06:00
Jason ZhouandClaude Opus 5.5 ca7770a3c7 fix(call): close abandoned idempotency claims with a stored 410; cut primary DB load (#748)
* fix(call): close abandoned idempotency claims with a stored 410 instead of 409 for 24h

A served call whose done-marking failed (a DB pool timeout) left its claim pending, and every
retry of that key answered 409 idempotency_in_progress until the 24h window expired.

- A claim is a lease owned by its call: it carries call_ref from the claim, the owner renews it
  every 5 minutes while running, and store, release and renewal are fenced on call_ref.
- A lease older than 15 minutes with no open hold and no pending async task under call_ref (and
  its call_ref: children) is closed by compare-and-swap on owner and lease timestamp with a
  stored terminal 410: idempotency_response_lost with the charge when the owner's ledger shows
  one (GET /calls/{id}/result may still have the answer), else idempotency_outcome_unknown.
  The key is never run again: a lapsed lease does not prove the owner stopped, and a second run
  under the same key is the double charge idempotency exists to prevent. Legacy rows without a
  call_ref keep answering 409 until they expire.
- Storing, releasing and renewing a claim retry once on a pool timeout, on a fresh session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(db): index the idempotency sweep and slow the arena refresh

Both run on the 1 vCPU primary whose CPU pinned during the pool-saturation windows that
stranded idempotency claims and settlements:

- _claim_idempotent sweeps expired labels on every keyed call, but idempotentcall had only
  membership_id, so the DELETE read every label the caller holds (tens of thousands for a
  batch caller). (membership_id, expires_at) makes it a range over the expired rows (0052).
- The arena collector re-aggregated the whole 30-day window every 120 s: three multi-second
  aggregates plus a retention DELETE over a multi-GB table, about a quarter of DB time.
  A 30-day leaderboard refreshes every 30 minutes now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 13:12:02 +10:00
shehjad-dev 1b512a8f8e Merge remote-tracking branch 'origin/main' into codex/octen-catalog
# Conflicts:
#	dsh/skills/treg/SKILL.md
#	plugin/skills/treg/SKILL.md
#	plugins/minimax/skills/treg/SKILL.md
#	plugins/treg/skills/treg/SKILL.md
#	skills/treg/SKILL.md
#	src/treg/oauth_providers.py
2026-09-30 03:37:59 +06:00
shehjad-dev 87ff7cc02f feat(catalog): add verified Octen search and extract routes 2026-09-30 03:36:14 +06:00
shehjad-dev 88c1849b9c fix(catalog): add Valyu shelf icons and refresh plugin skills 2026-09-30 02:51:39 +06:00
shehjad-dev a831ecbf76 docs(catalog): guide agents to verifier output facets 2026-09-29 23:43:06 +06:00
SToneX 94c6ff4f60 feat(catalog): add Gemini TTS and Lyria music on Google AI, with a music-gen platform
Four more google-ai rows, each spooled (a five-minute WAV or a long song exceeds 8 MiB as base64):

- google-ai.voice-gen.gemini-3-8-flash-tts and -flash-lite-tts: WAV speech in 30 prebuilt voices
  or a two-speaker dialogue, settled from usageMetadata (input $0.50/M, audio $9/M and $6/M).
- google-ai.music-gen.lyria-3-5 (full song, $0.08) and lyria-3-clip (30 s, $0.04): MP3 plus
  lyrics or section markers, a fixed price billed only when finishReason is STOP.

Music had no platform: `music-gen` joins the taxonomy, and like the other generation platforms
its answers are never cached. Capabilities are proposed in the provider file. Agent-facing files
list the new platform and the inline-media exception.
2026-09-30 00:59:35 +08:00
SToneX 732053599d feat(catalog): add Google AI (Gemini API) with Gemini 3 Pro Image
The first `google-ai` provider row: Gemini 3 Pro Image on Google's own generateContent, joining
reAPI and PiAPI on `image-gen.gemini-3-pro-image.generate` so catalog_get compares them. The
answer is synchronous with the image inline as base64, so the endpoint declares
`spooled_response: true` and settles from Google's usageMetadata token meters through
`usage.terms` at the documented paid-tier rates, holding the 1K/2K or 4K row. Reference images
pass as public URLs (fileData.fileUri), which Google fetches itself.

Registry entry (x-goog-api-key, free model-list probe), platform-key slot
`TREG_PLATFORM_KEY_GOOGLE_AI`, env-import probe, lettermark logo, NO_BALANCE_API reason
(postpaid Cloud billing) and the acknowledged capacity signature gap. Agent-facing files state
the inline-media exception to the 8 MiB limit.
2026-09-29 23:28:35 +08:00
shehjad-dev 6921fba4dd feat(catalog): add LiteScrape provider tools 2026-09-29 20:58:16 +06:00
shehjad-dev c60e372716 chore(plugin): sync provider count in skill copies 2026-09-29 20:06:39 +06:00
shehjad-dev eec7495787 feat(catalog): add Spider web endpoints and balance support 2026-09-29 18:19:20 +06:00
shehjad-dev 46cafed765 feat(catalog): add You.com BYOK and platform tools 2026-09-29 16:57:08 +06:00
shehjad-dev 7b9aeb942f feat(catalog): add Linkup web and research tools 2026-09-29 02:06:20 +06:00
shehjad-dev 65aa736dc1 feat(catalog): add Firecrawl web tools with bounded credit pricing 2026-09-28 22:19:05 +06:00
SToneX 3969eb9316 feat(catalog): show agent verdicts per endpoint on the platform page
Fold callreview into one vote per team per endpoint over 90 days, publish it on
GET /catalog/platforms/{slug} as `reviews` past five teams, and render a quiet
useful share on each endpoint line plus a Reviews tab with quotes drawn in
proportion to the vote. Provider lines show the access chip only for exceptions,
and the phone-width tab bar and provider line no longer collapse.

Fragments: architecture/feedback.md, interface/dashboard.md, interface/api.md.
2026-09-28 22:01:06 +08:00
Jason ZhouandClaude Opus 5.5 2de39612db fix(call): replays report zero cost; stored answers readable by call id (#703)
* fix(call): replays report zero cost; stored answers readable by call id

An idempotent replay echoed the first call's charge in X-Treg-Cost-Micro,
so a client summing that header counted one call twice (our own MCP tool
said "nothing was charged" beside a non-zero cost_usd). A replay now sends
X-Treg-Cost-Micro: 0 and the original charge in X-Treg-Original-Cost-Micro;
the CLI charge line reads the new header and falls back for older servers.

GET /calls/{id}/result took only the numeric audit row id, which a caller
holding the X-Treg-Call-Id it was given could not use without a second
lookup. It now accepts either. The endpoint was also absent from every
agent-facing doc, so a team that paid for answers its own code dropped had
no way to learn they could fetch them back; llms.txt, skill.md and
integrate.md now say so. Fragment updated: interface/api.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(context): replay cost header and call-id result lookup in money + archive

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 14:14:52 +10:00
b4ed9253d7 feat(catalog): serve herus13 Apify actors on the platform key with capped billing (#657)
* fix(catalog): keep Apify actor starts off treg's shared key

apify.web.scrape.job.start is priced free because the run bills later by the
actor's own pricing, and nothing on the shared-key path meters that run. On
treg's key it let any caller run any actor on treg's Apify account at no
charge. It now needs the team's own Apify key; job.status and job.results
stay open because per-team ownership already scopes them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): add Weibo, TikTok Shop, Lazada and Google Maps Apify actors

Combines #641-#644 into one catalog change: nine run-sync entries over four
herus13 actors, plus the lazada platform. The six Weibo entries serve on
treg's key at a price observed on it. Lazada, Google Maps and TikTok Shop are
own-key only: their actors bill run compute or a per-GB start fee that a
per-result price cannot meter.

Co-authored-by: Herus13 <bootforge.ai@gmail.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): require own keys for Weibo actor runs

* test(frontend): isolate landing interactions from continuous WebGL rendering

* feat(call): let platform_request pin query parameters

Some upstreams take their spend bounds as query options rather than body
fields (Apify's maxTotalChargeUsd, memory and timeout run options). A
queryParams pin must be sent exactly once and is compared as the pinned
value's type; own credentials keep the upstream contract.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill Apify platform calls by returned dataset rows

run-sync-get-dataset-items answers a bare array, so every Apify per_result
call settled at its estimate whatever it returned. Count the rows, add an
optional per-row-independent call_fee for the actor's start or compute
charge, and require maxItems (1-200) on the platform key so the hold is the
worst case. Apify's usageTotalUsd lags a finished run by minutes, so the
response body is the settlement evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): serve the herus13 Apify actors on treg's key

Weibo, Google Maps, TikTok Shop and Lazada now settle on the platform key by
counted dataset rows plus a flat call_fee (actor start, or Lazada's run
compute). memory and timeout are pinned so the fee is fixed and a run ends
before Apify's synchronous wait; TikTok Shop is held to keyword search.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bound Apify platform calls by maxTotalChargeUsd

maxItems does not bind pay-per-event actors whose own input sets the row
count, so a one-row hold could settle thousands of rows. Require the
maxTotalChargeUsd run option Apify enforces (at most $1), hold it plus
call_fee, accept only the run options each once in plain ASCII, and bill the
hold when a run reaches its cap. Pin meta-ads enrichment off and bill the
LinkedIn actor-start event on treg's key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): cap herus13 Apify runs by maxTotalChargeUsd on treg's key

Row notes name maxTotalChargeUsd as the enforced spend limit; Lazada's
compute fee is 0.015 under a 180-second timeout pin.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): name maxTotalChargeUsd as the Apify spend cap

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): treat an Apify run within two rows of its hold as capped

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(call): require a short timeout and a three-row cap on Apify platform runs

Past Apify's 300-second synchronous wait a run answers 408 and keeps billing,
so every Apify per_result platform call now names timeout <= 280. A cap under
call_fee plus three rows would bill an empty answer in full, so it is refused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(call): keep Apify platform runs inside every wait

treg's upstream read timeout (call_timeout_s, 180 s) and the MCP client's
120 s end the call before a 280 s run finishes, releasing the hold unbilled
while the run keeps billing. Bound timeout to 90 s and 30 s under
call_timeout_s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): pin herus13 Apify runs to a 90-second timeout on treg's key

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill a timed-out Apify platform run at its hold

A run that outlives its own timeout answers 400 run-failed with no rows, but
Apify billed its events up to the caller's cap and the run id in that body
reads the dataset. The caller chose the run's size and timeout, so the hold
settles instead of releasing. The minimum-cap check compares micro-USD.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): tell Lazada callers to keep platform runs small

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): release timed-out Apify runs; the account must stay Restricted

Billing the whole cap on a TIMED-OUT 400 overcharged callers: a run that
timed out after its start event cost Apify $0.00005 and would have billed the
full cap. The loss treg absorbs stays bounded by the $1 cap, and keeping the
Apify account's resource access Restricted stops anyone reading the unbilled
run's rows by id. Examples now show the 90-second platform timeout.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(money): name disconnects as a bounded Apify loss; call_fee wording

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): record Lazada's 0.05 minimum cap and 10-product floor

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(asynctasks): keep the Bright Data and CompanyEnrich platform keys the merge dropped

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): verify every Apify per-result price on the platform key

Each row's test_request ran on 2026-09-26 and Apify's chargedEventCounts
matched its rate card. The TikTok ad library actor returns rows again, so it
is verified with a captured example.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(catalog): route Google Maps and Weibo post detail through Apify

Adapters let treg.google.serp.maps and treg.weibo.post.detail choose the
Apify actors, with the run's spend cap, timeout and memory fixed so the
platform guard accepts the child call.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): drop Lazada's compute fee and keep account details out of notes

The actor's 2026-09-25 pricing no longer bills run compute to the caller (a
79-second run showed no platform usage), so its call_fee over-charged every
call. Notes describe observations by price tier, not by the account that
made them, and carry no run ids.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): state tiered Apify observations without the account's tier

Notes for plan-tiered actors record the events billed and that they matched
the rate card for the key's tier, not the dollar figure that would name it;
the repeated Weibo observation and the owner-meter notes go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settle): bill LinkedIn's actor-start once per query it runs

The LinkedIn jobs actor bills its actor-start event for every job title x
location searched, so a flat call_fee under-billed any multi-query call.
cost.call_fee_per names the body arrays whose lengths multiply the fee. The
apify.yaml header no longer names a plan or calls the TikTok actor broken.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(money): describe call_fee_per

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(catalog): accept only declared fields on LinkedIn job search

The actor bills an actor-start for every geo id it searches, and geoIds was
undeclared, so it passed through unbilled. The row now takes its declared
filters only (salary, easyApply, under10Applicants and industryIds added);
places go in locations, which call_fee_per counts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(catalog): leave headroom above the Apify spend cap

A run whose charges land exactly on maxTotalChargeUsd can be aborted by
Apify and answer 400 with no rows, which releases the hold. Lazada's
10-product floor costs exactly its 0.05 minimum cap, so its example now uses
0.06 and the notes say to set the cap above the expected spend.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Herus13 <bootforge.ai@gmail.com>
2026-09-26 15:04:26 +10:00
UncleCode ac5ed26e57 release: 0.22.0 2026-09-26 09:30:58 +08:00
UncleCode 2ff15db042 fix(hub): with TREG_HUB_TEAMS set, a team outside the list sees no trace of the hub
Before, the team list gated only the hub routes and /call/ of a hub id. The open surfaces kept the
plain flag, so with the hub on for one team every other user still read the hub sections of the
agent files, found approved hub tools in catalog search, could open catalog get and the share page,
and saw hub_create, hub_update and hub_mine in the MCP tool list (listed even with the flag off).

- `hub_app.visible_to(slug)`: a reader with a team is judged by `enabled_for`; a reader with no team
  sees the hub only when the list is empty.
- `routers/hub_gate.reader_team` resolves the key or session a request carries, only while a list
  is set. Catalog search, catalog get and its sibling rows, the share page, /skill.md and /llms.txt
  use it; MCP catalog_search resolves the bearer's team the same way.
- `mcp._HubToolsGate` drops the three hub tools from `tools/list` for anyone who cannot see the hub.
- `treg skill bootstrap` sends the key, so a listed team's agents get the hub sections.
- The static plugin files no longer carry the hub sections (owner, 2026-09-26; replaces the
  2026-09-16 decision to always ship them).

Updates docs/context/architecture/hub.md.
2026-09-26 07:52:48 +08:00
UncleCode 3ee1948c64 Merge origin/main into dev/hub: the tool hub, behind TREG_HUB_ENABLED
Brings main's 86 commits (dashboard boot and loading, legacy dashboard removal, test pruning,
overflow and routing fixes) together with the tool hub branch.

Conflicts, both sides kept unless noted:
- the legacy dashboard stays deleted, as on main;
- App.vue and the dashboard state: main's search page and loading states plus the hub pages;
- ci.yml: main's Postgres job, with the hub tests added to its list;
- dev-local.sh: main's server environment plus the hub flag passthrough;
- test_call_application_contract.py, test_marketplace_call.py: main's pruned files plus the
  hub branch's sync `settle: usage` test.

Not conflicts: main and the hub branch fixed the same CompanyEnrich empty-page billing; main's
rule runs first, so the hub branch's copy and its test are dropped. The Listing-tab test reads
the Vue source instead of the deleted legacy page.
2026-09-26 07:32:53 +08:00
Jason ZhouandClaude Opus 5.5 fd5a2c0033 fix: stop billing empty answers, make /admin/errors read-only, add treg --json call (#675)
* fix(billing): stop charging for empty and company-blind answers

Per-result searches settled at the requested page size even when the
vendor returned nothing: count the billed rows for CompanyEnrich (2-credit
minimum), Findymail employees, Icypeas bulk (FOUND only), TheCompaniesAPI
search (and simplified=true, which is free) and Serpstat (error envelopes
free, 1-credit minimum on an empty result). Icypeas profile-URL misses
settle at zero through an expect rule.

Routing: Prospeo 400 NO_RESULTS is a declared miss, not a caller fault;
the CompanyEnrich scroll adapter is removed so one empty question is not
routed twice; people.search declares scoping: [company_domain], so a
title-only provider is dropped from a company-scoped plan instead of
billing title-matched strangers once the others miss.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMPFwsMRowa1X7kXc5nZyw

* fix(admin): make GET /admin/errors read-only; purge evidence from a worker

Loading the errors page ran a platform-wide UPDATE blanking failed-call
evidence past the 14-day window. The purge moves to
application/evidence_retention.py behind 'treg-worker admin
purge-evidence', in bounded id batches. The view withholds evidence past
the window itself, so an unscheduled purge never widens what it shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMPFwsMRowa1X7kXc5nZyw

* feat(cli): treg --json call prints one parseable envelope

With --json, call prints {"result": <body>, "_treg": {http_status,
call_id, charged_micro|reserved_micro, ...}} on one line and nothing on
stderr, so a script that merges the streams still parses every answer.
Default output and --await are unchanged. Agent docs say to use it when
parsing, and to check a few results before looping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMPFwsMRowa1X7kXc5nZyw

* fix(billing): price counted rows at their credits and cap them at the hold

For a credit-priced per_result row the frozen unit is one provider credit,
so counting rows alone under-billed CompanyEnrich (2 credits a person, and
its 2-credit minimum) and Icypeas reverse-email bulk (10 a hit). Rows whose
catalog unit names an input entity reserve per thing asked about, so a row
count is capped at the hold: counting may lower a bill, never raise it.
simplified=true is free only on TheCompaniesAPI endpoints that declare it.

Also: a malformed contract scoping list fails the catalog load; --json
call keeps stderr silent on the WAF base64 retry; AGENTS.md records the
evidence-retention purge as the second callrecord writer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMPFwsMRowa1X7kXc5nZyw

* fix(billing): bill Icypeas company scrapes at the company rate; never exceed a zero hold

Also corrects the AGENTS.md callrecord-writer note and the scoping
comment (a bad list empties routing rather than failing the catalog).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMPFwsMRowa1X7kXc5nZyw

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 20:48:54 +10:00
UncleCode 8d5bd38125 fix(hub): nine fixes from the third hub simulation
The same cheating seller, after the round 5 fixes (docs/hub-listing-decisions.md round 6). Run
2's holes held; these are the new ones.

- A tool whose listing treg rejected serves only its maker's team: another team's call 404, the
  share page 410, catalog get 404. A rejected "Official Hunter.io" tool stayed callable by id.
- Team names are judged as they read: NFKC, lookalike letters and digits mapped, split into
  words. `trеg-hub` (Cyrillic е), `tregg`, `t-r-e-g`, `the treg team`, `hunter-io`, `Hunter.io
  data`, `apol1o`, `verified-partner` are refused; "official"/"verified" match as words, so
  `unverified-0` passes; platforms (people, web) match only as the whole name.
- A team tool may not point at treg itself (at /tools, and again in the runner): `treg-relay` ->
  <treg>/call let an approved tool call an unreviewed hub tool.
- The reviewer reads the code: the listing queue carries the script or steps and the own tools'
  base URLs, the update queue a diff of the code. A date switch in run.js was invisible.
- The public price counts only runs of the version priced; the maker's list prices the live one.
  An unapproved version's check run raised the approved tool's price.
- check.json may hold {"cases": [...]}, up to 5, all run at publish; the first failure names its
  case. One sample could not reach both paths of an email tool.
- The fee line names what runs paid: "provider fees (about $0.005 so far, at most $0.1 a run)".
- `treg hub ls` prints reasons under the newest row when no version is live.
- A maker cannot review its own hub tool; the review message says what serves now (nothing,
  when every approved version was retired).

Updates docs/context/architecture/hub.md.
2026-09-25 13:10:59 +08:00
UncleCode 2c21ea7c55 fix(hub): nine fixes from the second hub simulation (a cheating seller)
A simulated seller tried 14 ways to change what buyers get or pay (docs/hub-listing-decisions.md
round 5). Two were verified in the database; the rest were checked against the code.

- Once approved, always reviewed: HubListing.reviewed is set by the first approval and never
  cleared. Unlisting a reviewed tool moves it to state `unlisted` (out of search) instead of
  deleting the row, so its new versions and prices still wait and callers by id keep the approved
  version; listing it again needs no new review. A rejection that takes an approval back keeps the
  review too. Before, approve, unlist, change let v3 ($0.04 a run) and v4 ($0.05, no industry) go
  live at once under a summary that still promised $0.001.
- Reserved team names: a slug or name that is `treg`, starts with `treg-`, contains `official`, or
  is a catalog provider or platform slug cannot be created, renamed to, or publish a hub tool,
  unless a superadmin acts. A team `treg` had published `treg.companies-enrich`. A new user's
  email-derived team falls back to `team-<id>`.
- An answer with every declared output field empty settles no seller price.
- The price line names the provider-fee limit ("+ provider fees up to $1 a run"), and "you pay
  only what completes" becomes a plain note that a failed run still pays the steps that ran.
- The public price rests on runs by other teams once there are any; the maker's own free runs
  and checks are marked "(from the maker's own tests)".
- The contract names the hosts of the maker's own tools a version calls (`sends_inputs_to`).
- A rejection's reason stays while the maker asks again; `treg hub ls` prints it.
- `treg hub price --help` says what it does on a script; a publish that changes the price says
  so (`price_note`).
- `treg org create` says the active team changed; `treg balance` sums earned credit in one line.

HubListing gains `reviewed` (in the unreleased 0050).
Updates docs/context/architecture/hub.md.
2026-09-25 12:23:07 +08:00
UncleCode 1668edb7ea fix(hub): ten fixes from the first hub simulation
Two agents used the hub as a maker and a buyer on a local server; each finding below was checked
against the code or the database before this fix.

- verified-email and lead-pipeline accept only a work email: never a free mail provider, and at the
  input company's domain when one is given. A finder returned a Gmail address for a company's CMO,
  the checker called it valid, and the tool charged its fee.
- A script's price headline shows its cap after the observed range ("$0.0103/run so far · seller
  up to $0.02 a run"): recent runs may all have charged little.
- `treg catalog get` on a hub tool no longer labels the whole run cost "per successful run to the
  maker"; it prints the seller price and, separately, what a run cost a caller.
- After a publish, the CLI's example call uses the tool's own inputs, not a fixed figma.com body.
- `treg hub ls` shows the maker's price and the search state (no / requested / yes / rejected,
  update waits) on the version callers get, marked ◀.
- `treg hub price` on a script says it moved only the cap; the fee is the ctx.charge lines.
- The recipe rules (summary 1-200 characters; the input keys; required, example, int max) are in
  skill.md, llms.txt and the `treg hub init` output.
- A version in `checking` or `review` is never shown on the public views (catalog get, share
  page); before, `catalog get <id>@N` showed an unreviewed summary and price to anyone.
- New `treg whoami`; `treg balance` names the team. `treg whoami` used to run the system whoami
  through `treg with` and print the machine's user name.
- `_kv` keeps a space after a 7-character label ("listingrequested" in `treg hub list`).

Updates docs/context/architecture/hub.md.
2026-09-25 11:04:31 +08:00
shehjad-dev 4eb30f3037 fix(catalog): route Fetchin LinkedIn tools 2026-09-25 01:12:44 +06:00
shehjad-dev 035c3959e1 release: 0.21.3 2026-09-24 23:31:04 +06:00
shehjad-dev e6d628d63a feat(catalog): add Serper tools 2026-09-24 22:56:37 +06:00
shehjad-dev dc9d4f9e80 feat(catalog): add ScrapeGraphAI provider
Add ScrapeGraphAI key auth, shared-key capacity collection, catalog pricing and rate policy, fifteen live-verified tools, web adapters, examples, and the supplied logo. Keep monitor mutations BYOK-only while exposing bounded scrape, extract, search, and crawl surfaces on both BYOK and platform credentials.
2026-09-24 22:02:58 +06:00
shehjad-dev 75bd4a1a0e feat(enrichment): add async Wiza email and phone 2026-09-24 20:09:28 +06:00
UncleCode b36832a403 feat(hub): every update to a listed tool waits for treg's review
Jason: a maker could get a tool approved, then publish a version that charges more or returns
worse data, and it would serve every search caller at once (docs/hub-listing-decisions.md round 4,
which replaces round 2's "an approval stays across new versions").

- A version that passes its check on an approved tool becomes `review`, not `live`, so every
  surface that reads `live` (calls by id, search, the place beside providers, the share page) keeps
  the approved version. A newer waiting version makes the older one `superseded`. The maker's team
  can call a waiting version by `<id>@N`; nobody else can.
- `treg hub price` on an approved tool is checked at once and stored as pending; the approved
  price keeps serving.
- GET /admin/hub/updates (what serves now beside what would replace it) and POST
  /admin/hub/updates/{id} {decision, reason}: approve makes the version live and applies the
  price; reject marks the version `rejected`, drops the price and gives the maker the reason.
- Unlisting, or a rejected listing, releases what waits: an unlisted tool is self-serve.
- HubListing gains pending_version, pending_pricing, update_reason (in the unreleased 0050).
- CLI, the dashboard's Listing tab and a new "Hub updates waiting" queue on the Admin page say
  it; skill.md, llms.txt, USAGE.md; plugins and surface snapshots regenerated.

Tests: a new version waits while v1 serves and only the maker can pin v2; approve serves v2;
reject keeps v1 with a reason; superseded; a price change waits and serves after approval;
unlisting releases both.
Updates docs/context/architecture/hub.md.
2026-09-24 17:45:07 +08:00
UncleCode 744200d44f fix(hub): delete a team's listing rows with the team; record the new surface
- cascade_delete_org now deletes HubListing: a team with a listing request could not be deleted
  (the foreign key failed with a 500). Found by test_org_delete_clears_EVERY_org_scoped_table.
- tests/snapshots: the two admin listing routes and their schema (scripts/dump_surface.py).
- The five plugin copies of the skill, regenerated after the pricing, listing and capability text
  (scripts/build_plugin.py).
2026-09-24 15:39:51 +08:00
UncleCode 986d5ffa57 merge: main into dev/hub, hub migrations renumbered 0044-0049
Brings main's 24 commits since the 2026-09-21 merge: the new Vue dashboard in frontend/ with
the account-based rollout (legacy frozen as src/treg/web/dashboard-legacy/), Olostep and
Keenable, routed web search, TrestleIQ, call_media and resources_list on MCP, provider
resources, the search log.

Conflicts, every one "both sides added": imports in call/service.py and routers/web.py; the
error-owner table in call/types.py; dev-local.sh keeps the TREG_HUB_ENABLED passthrough in
front of main's new SERVER_ENV; the legacy dashboard keeps both the Resources and the Hub
nav entries and view names; test_mcp.py lists main's two new tools and the hub's three;
test_marketplace_call.py keeps both new test blocks; .gitignore keeps both. MAP.md and
docs/context/README.md regenerated. tests/test_dashboard_markup.py removed, as on main.

The hub's six migrations move from 0041-0046 to 0044-0049 above main's 0043; a fresh database
upgrades to one head, 0049.
2026-09-24 12:38:40 +08:00
UncleCode 9757ff544c feat(hub): ctx.charge, an own-key step's cost billed to the caller
An own-key step costs the caller nothing and the maker real money at a vendor treg cannot
see; a waterfall over four such vendors costs $0.001 one run and $0.50 the next, and no price
table can say which branch ran. So the script says it: `ctx.charge(usd, label)` after the
step, once it has seen the answer. `pricing.max_charge_usd` in the manifest is the most all
charges may total in one run, allowed with every mode, script tools only; without it a
charge is refused and the run stops. Held with the fee on the same hold, refused past the
cap or past 20 lines, one `charged` trace line per charge with the maker's label,
`usage.charged_micro` their sum, settled to the maker as fee + charges, released with the
fee when the run fails. The price label reads "+ own-key steps up to $X per run".

Proven end to end: a $0.01 fee plus charges of $0.001 and $0.03 (a zero charge dropped)
bills the caller $0.041 and the maker earns $0.041; a charge past the cap fails the run and
the caller pays nothing. (owner + Jason, 2026-09-24)

Also: the callmatrix refusal test reaches the sandbox bridge by its parallel-era name, and
search-console-health's summary fits the 200-character rule.

Fragment updated: architecture/hub; llms.txt, skill.md and the plugin mirrors say the same.
2026-09-24 12:11:19 +08:00
Jason ZhouandClaude Opus 5.5 4a2e03a1c8 feat(hub): price in the maker's words, lead with what a run costs
Pricing modes are renamed to what a maker would say: per_call (was flat),
per_result (was per_unit; the count is `results`, `units` still read) and
percent (was cost_plus). The old names still validate and are stored
canonically, so no published tool changes.

max_price_usd is no longer required. The hold is derived from what bounds
the run: the caller's ceiling for percent (fees + part <= ceiling), the
results_from input for per_result. The settled price never passes the hold.

Every price surface (dashboard, treg hub ls, public page, catalog_get,
search rows) now leads with what recent successful runs actually cost,
provider fees and the seller part together, as one number or a low-high
range. price_ranges reads a run's output only for the maker's own runs of a
per_result tool, since catalog search calls it for every listed tool.

Dashboard: the overview cards no longer overflow (the observed cost and the
maker's price are separate cards, long values wrap), and the Price tab is
display-only with a copyable prompt for the maker's agent. A form could set
only a per-call price and silently flattened a variable one.

Adds the engineering-team-size recipe (a script ladder over three catalog
tools) with its matrix test.

Docs: hub.md, hub-pricing-decisions.md (round 3), USAGE, CHANGELOG,
llms.txt, skill.md and the plugin skill copies.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 13:30:01 +10:00
shehjad-dev ff7cabb9c8 chore(plugin): regenerate provider mirrors 2026-09-23 22:54:51 +06:00
shehjad-dev d396f644ce chore(plugin): regenerate provider mirrors 2026-09-23 21:29:48 +06:00
shehjad-dev f42487625f release: 0.21.2 2026-09-23 20:08:29 +06:00
shehjad-dev 5ed9ab31ee feat(catalog): add TinyFish tools 2026-09-23 18:39:57 +06:00
shehjad-dev 49da738b03 feat(async): settle provider-native usage meters 2026-09-23 18:25:30 +06:00
shehjad-dev c296245d5b feat(catalog): add Adyntel provider 2026-09-23 14:32:52 +06:00
shehjad-dev adc4266d98 release: 0.21.1 2026-09-23 13:08:59 +06:00
UncleCode e727466155 docs(agent): point agents at the six first-party hub tools before the routed chains
skill.md and llms.txt gain one block, inside the hub markers so it vanishes when the hub is
off: six jobs are already built as tools, call the tool and not the chain. One table maps
"you want" to the tool id and its inputs and to the hand-built chain it replaces. The block
sits BEFORE the routed guidance so an agent reads the finished tool first. The "never re-send
the same find" rule now names treg-hub.verified-email as the one-call form. Five plugin
mirrors regenerated.

Why (review 2026-09-22): 360 teams chain find → verify by hand, 53 chain the whole lead
pipeline, 73 call up to ten providers for one question; almost every caller is an agent
following this file, so this file is where the behaviour changes.
2026-09-23 13:59:42 +08:00
shehjad-dev d838e851fe release: 0.21.0 2026-09-22 23:49:25 +06:00
shehjad-dev 5abda88df8 fix(skill): regenerate treg mirrors 2026-09-22 22:23:07 +06:00
shehjad-dev 26773fa8e1 feat(fishaudio): complete managed voice workflows 2026-09-22 20:46:06 +06:00
shehjad-dev dd14ed7b98 feat(fishaudio): add organization-scoped voice tools 2026-09-22 19:13:02 +06:00
6e667a4c6f feat(auth): isolate pinned agent history and shared-provider async reads (#616)
* feat(auth): scope pinned agent history and async reads

* fix(auth): a routed parent row carries the pin; reviews use the pinned read scope

`_audit_parent` in `application/call/route.py` wrote the successful routed
parent audit row without `_tag_telemetry`, so once history is filtered by
pin a customer's own routed calls vanished from `/calls`, its result 404'd
and `/calls/{ref}` returned `call: null`. The refusal path already carried
the tags; this aligns the success path.

`POST /reviews` resolved its call by org alone, so a pinned caller could
rate, and thereby probe, another pin's call. It now applies the same
predicates as `/calls` and answers 404 for a foreign reference.

One test covers both: a pinned routed call is listed and readable by its
pin, invisible to another, reviewable only by its own.

Fragment: architecture/multi-tenancy.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(auth): refusals carry the pin; feedback reports read through it

Two omissions an independent review turned up.

The router's refusal fallback (`routers/call.py`, the one audit writer
outside `service.py`) wrote its row with no tags, so a pinned caller
could not see its own 404s and 403s — the pin-mismatch 403 above all,
the row a builder most needs while integrating. The handler now stashes
the pin next to the identity and the fallback records it.

`GET /feedback/{id}` was org-scoped with sequential ids, so a pinned
token could read another customer's free-text report. `Feedback.tags`
(revision 0042) snapshots the reporter's pin; the read applies the same
predicates as `/calls`, and a pinned reporter's `call_ids` verify only
against its own pin's rows. Unpinned behaviour is unchanged.

Fragment: architecture/multi-tenancy.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Jason Zhou <jason.zhou.design@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 20:49:08 +10:00
shehjad-dev dde71d82f3 chore(plugin): refresh catalog provider count 2026-09-21 23:59:16 +06:00