mirror of
https://github.com/superdesigndev/treg.git
synced 2026-10-02 03:24:35 +08:00
feat(find): record the verdict on SearchLog; the report runs on Postgres
A v2 answer's verdict was only implied by its SearchLog row (an empty shown for none, owner judged for both strong and closest), and the experiment report stratifies by it. Migration 0056 adds searchlog.verdict, the reason after a colon (none:gap), written by v2 finds. scripts/search_experiment_report.sql: the interleaving-credit block never ran on Postgres (round(double precision, int) does not exist; cast to numeric), the mode its arms are read from is a psql variable, and latency percentiles leave out answers served from the judge's in-process cache (0 ms). Fragments: find.md, data-model.md, search-experiment.md.
This commit is contained in:
@@ -47,6 +47,7 @@ sources:
|
||||
- src/treg/alembic/versions/0033_signup_promo_eligibility.py
|
||||
- src/treg/alembic/versions/0041_searchlog.py
|
||||
- src/treg/alembic/versions/0055_find_v2_log.py
|
||||
- src/treg/alembic/versions/0056_searchlog_verdict.py
|
||||
- src/treg/timeutil.py
|
||||
- src/treg/infra/db.py
|
||||
- src/treg/domain/referrals.py
|
||||
@@ -366,7 +367,8 @@ uses this metadata, never the encrypted token's shape.
|
||||
Nothing is written while `search_experiment` is `off`.
|
||||
`/catalog/find` writes the same row with `mode=find`, `source=web-find` and no identity; 0055 adds
|
||||
`engine` (v1 | v2) and v2's readings: `platform_choice`, `platform_conf`, `name_p`, `recall_ms`,
|
||||
`embed_ms`, `embed_error`, and `units` as `[kind, id, p]` rows ([find](find.md)).
|
||||
`embed_ms`, `embed_error`, and `units` as `[kind, id, p]` rows ([find](find.md)); 0056 adds
|
||||
`verdict`, the verdict a v2 answer ended on with its reason after a colon (`none:gap`).
|
||||
- **`RunRecord`** - the **server-side run** audit row (a `treg run --server` CLI execution - the "kind"
|
||||
`server_run` in usage rollups): `org_id`, `user_email`, `bundle_name` (holds the **tool** name since the
|
||||
tool-side run unification; column name is historical), `argv` (JSON - never carries a secret value;
|
||||
|
||||
@@ -10,6 +10,7 @@ sources:
|
||||
- tests/test_find_index.py
|
||||
- tests/test_embed.py
|
||||
- src/treg/alembic/versions/0055_find_v2_log.py
|
||||
- src/treg/alembic/versions/0056_searchlog_verdict.py
|
||||
- scripts/find_bench.py
|
||||
- tests/fixtures/find_bench.yaml
|
||||
- tests/test_find_bench.py
|
||||
@@ -192,13 +193,14 @@ the right platforms and vendors); `judged` gains `reason` (on
|
||||
## What is recorded
|
||||
|
||||
`SearchLog` (mode `find`, source `web-find`, no identity) with `engine`; v2 also writes
|
||||
`platform_choice`, `platform_conf`, `name_p`, `recall_ms`, and `units` as `[kind, id, p]`;
|
||||
`verdict` (the reason after a colon: `none:gap`), `platform_choice`, `platform_conf`, `name_p`,
|
||||
`recall_ms`, and `units` as `[kind, id, p]`;
|
||||
`baseline_ids` is every endpoint the units reach, `judged` the kept units. `embed_ms` and
|
||||
`embed_error` record the query's vector (`off` without a key, `not_ready` while the card vectors
|
||||
build, else the client's reason). The `judged` event carries the same as `embed: {ms, error}`. `SearchMiss` gains `reason` (`gap`,
|
||||
`not_task`, `judge_off` for an empty keyword fallback, `scope` for a shelf's `none`) and `engine`.
|
||||
Only the served engine files a miss: in `shadow` v1 does, and v2's empty answers show in its
|
||||
SearchLog row only, so a find never counts twice in the misses. Migration 0055. Fire-and-forget through
|
||||
SearchLog row only, so a find never counts twice in the misses. Migrations 0055 and 0056. Fire-and-forget through
|
||||
`audit`, like every row there.
|
||||
|
||||
## The pages
|
||||
|
||||
@@ -102,8 +102,9 @@ the database.
|
||||
|
||||
## Reading it
|
||||
|
||||
`scripts/search_experiment_report.sql` (Postgres, read-only) joins `searchlog` to `callrecord` by
|
||||
team + email within ten minutes of the search, on endpoints that were on the served page:
|
||||
`scripts/search_experiment_report.sql` (Postgres, read-only; `-v mode=` picks the mode whose arms
|
||||
it reads) joins `searchlog` to `callrecord` by team + email within ten minutes of the search, on
|
||||
endpoints that were on the served page:
|
||||
|
||||
1. volume and health per arm — differs share, empty-baseline share, judge error rate, p50/p95 judge
|
||||
latency, tokens;
|
||||
|
||||
@@ -5,10 +5,11 @@
|
||||
-- each row's owner; `callrecord` (audit) carries the call. No labels anywhere — the join IS the label.
|
||||
--
|
||||
-- Run against the read replica: psql "$TREG_READ_DATABASE_URL" -f scripts/search_experiment_report.sql
|
||||
-- Postgres only (jsonb functions). Every block is read-only.
|
||||
-- (`-v mode=v2` reads another mode's arms). Postgres only (jsonb functions). Every block is read-only.
|
||||
|
||||
\set window '30 days'
|
||||
\set followup '10 minutes'
|
||||
\set mode 'interleave'
|
||||
|
||||
-- 1. Volume and health: how many searches, how often the pages differ, what the judge cost.
|
||||
-- `differs` is the population the experiment can say anything about; an identical page is a
|
||||
@@ -19,8 +20,9 @@ SELECT mode, arm,
|
||||
round(100.0 * avg(differs::int), 1) AS differs_pct,
|
||||
round(100.0 * avg((baseline_total = 0)::int), 1) AS baseline_empty_pct,
|
||||
round(100.0 * avg((judge_error IS NOT NULL)::int), 1) AS judge_error_pct,
|
||||
percentile_cont(0.5) WITHIN GROUP (ORDER BY judge_ms) AS judge_ms_p50,
|
||||
percentile_cont(0.95) WITHIN GROUP (ORDER BY judge_ms) AS judge_ms_p95,
|
||||
-- a judge answer served from the in-process cache costs 0 ms; the latency is the live one's
|
||||
percentile_cont(0.5) WITHIN GROUP (ORDER BY judge_ms) FILTER (WHERE judge_ms > 0) AS judge_ms_p50,
|
||||
percentile_cont(0.95) WITHIN GROUP (ORDER BY judge_ms) FILTER (WHERE judge_ms > 0) AS judge_ms_p95,
|
||||
sum(judge_tokens_in) AS judge_tokens_in
|
||||
FROM searchlog
|
||||
WHERE created_at > now() - :'window'::interval
|
||||
@@ -34,7 +36,7 @@ WITH pages AS (
|
||||
SELECT s.id, s.arm, s.org_id, s.user_email, s.created_at, (s.baseline_total = 0) AS baseline_empty,
|
||||
ARRAY(SELECT jsonb_array_elements(s.shown::jsonb)->>0) AS shown_ids
|
||||
FROM searchlog s
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = 'interleave' AND s.org_id IS NOT NULL
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = :'mode' AND s.org_id IS NOT NULL
|
||||
),
|
||||
converted AS (
|
||||
SELECT p.id, bool_or(c.id IS NOT NULL) AS converted
|
||||
@@ -63,7 +65,7 @@ WITH il AS (
|
||||
ARRAY(SELECT jsonb_array_elements(s.judged::jsonb)->>0) AS judged_ids,
|
||||
ARRAY(SELECT jsonb_array_elements(s.shown::jsonb)->>0) AS shown_ids
|
||||
FROM searchlog s
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = 'interleave' AND s.arm = 'interleave'
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = :'mode' AND s.arm = 'interleave'
|
||||
AND s.judged IS NOT NULL AND s.differs AND s.org_id IS NOT NULL
|
||||
),
|
||||
first_call AS (
|
||||
@@ -95,8 +97,8 @@ SELECT sum((credit = 'judged')::int) AS judged_wins,
|
||||
sum((credit = 'tie')::int) AS ties,
|
||||
(SELECT count(*) FROM il) AS interleaved_searches_with_disagreement,
|
||||
round(
|
||||
(sum((credit = 'judged')::int) - sum((credit = 'baseline')::int))
|
||||
/ sqrt(nullif(sum((credit IN ('judged', 'baseline'))::int), 0)), 2) AS z
|
||||
((sum((credit = 'judged')::int) - sum((credit = 'baseline')::int))
|
||||
/ sqrt(nullif(sum((credit IN ('judged', 'baseline'))::int), 0)))::numeric, 2) AS z
|
||||
FROM points;
|
||||
|
||||
-- 4. Re-query rate: a second search by the same caller within two minutes with no call in between
|
||||
@@ -104,7 +106,7 @@ FROM points;
|
||||
WITH s1 AS (
|
||||
SELECT s.*, lead(s.created_at) OVER (PARTITION BY s.org_id, s.user_email ORDER BY s.created_at) AS next_search
|
||||
FROM searchlog s
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = 'interleave' AND s.org_id IS NOT NULL
|
||||
WHERE s.created_at > now() - :'window'::interval AND s.mode = :'mode' AND s.org_id IS NOT NULL
|
||||
)
|
||||
SELECT arm, count(*) AS searches,
|
||||
round(100.0 * avg((next_search IS NOT NULL AND next_search < created_at + interval '2 minutes'
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
"""searchlog: the verdict a judged search ended on
|
||||
|
||||
Revision ID: 0056
|
||||
Revises: 0055
|
||||
Create Date: 2026-10-01
|
||||
|
||||
A v2 answer's verdict (strong | closest | name | none | keyword) was only implied by the row: a
|
||||
`none` by an empty `shown`, strong and closest both by owner `judged`. The report stratifies
|
||||
conversion by it, so the verdict is its own column, with the reason behind a none or a keyword
|
||||
fallback after a colon (`none:gap`, `keyword:not_task`). Nullable, no default: v1 rows and rows
|
||||
written before read as unknown.
|
||||
"""
|
||||
from collections.abc import Sequence
|
||||
|
||||
import sqlalchemy as sa
|
||||
from alembic import op
|
||||
|
||||
revision: str = "0056"
|
||||
down_revision: str | Sequence[str] | None = "0055"
|
||||
branch_labels: str | Sequence[str] | None = None
|
||||
depends_on: str | Sequence[str] | None = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
op.add_column("searchlog", sa.Column("verdict", sa.String(), nullable=True))
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
with op.batch_alter_table("searchlog") as batch:
|
||||
batch.drop_column("verdict")
|
||||
@@ -602,6 +602,7 @@ def _log_v2(query: str, cat: catalog_store.Catalog, found: Found, r: Recalled, r
|
||||
platform_choice=plat.get("choice"), platform_conf=plat.get("confidence"),
|
||||
name_p=None if found.name_p is None else round(found.name_p, 3), recall_ms=round(r.recall_ms),
|
||||
embed_ms=r.embed.ms, embed_error=r.embed.error,
|
||||
verdict=f"{found.verdict}:{found.reason}" if found.reason else found.verdict,
|
||||
units=[[c.unit.kind, c.unit.id, None if p is None else round(p, 3)] for c, p in zip(found.cands, probs)])
|
||||
if served and (found.verdict == NONE or (found.verdict == KEYWORD and not found.rows)):
|
||||
audit.record_search_miss(query=query, source="web-find", engine="v2",
|
||||
|
||||
@@ -1564,6 +1564,9 @@ class SearchLog(SQLModel, table=True):
|
||||
embed_ms: int | None = Field(default=None)
|
||||
embed_error: str | None = Field(default=None)
|
||||
units: list | None = Field(default=None, sa_column=Column("units", JSON, nullable=True))
|
||||
# the verdict a v2 answer ended on, the reason after a colon where there is one
|
||||
# (strong | closest | name | none:gap | keyword | keyword:not_task ...); None for v1
|
||||
verdict: str | None = Field(default=None)
|
||||
|
||||
|
||||
class CapacityPolicy(SQLModel, table=True):
|
||||
|
||||
@@ -333,7 +333,7 @@ async def test_v2_streams_units_and_the_answer_and_logs_its_readings(clients, mo
|
||||
async with session_maker() as s:
|
||||
(row,) = (await s.execute(select(SearchLog))).scalars().all()
|
||||
assert row.engine == "v2" and row.platform_choice == "people" and row.platform_conf == 0.9
|
||||
assert row.name_p == 0.0 and row.recall_ms is not None
|
||||
assert row.name_p == 0.0 and row.recall_ms is not None and row.verdict == "strong"
|
||||
assert ["job", "people.email.find", 0.92] in row.units and dict(row.judged) == {"people.email.find": 0.92}
|
||||
|
||||
|
||||
@@ -449,7 +449,7 @@ async def test_shadow_files_one_miss_from_the_engine_it_serves(clients, monkeypa
|
||||
misses = (await s.execute(select(SearchMiss))).scalars().all()
|
||||
logs = (await s.execute(select(SearchLog))).scalars().all()
|
||||
assert [(m.engine, m.source) for m in misses] == [("v1", "web-find")]
|
||||
assert sorted(r.engine for r in logs) == ["v1", "v2"]
|
||||
assert sorted((r.engine, r.verdict) for r in logs) == [("v1", None), ("v2", "none:gap")]
|
||||
|
||||
|
||||
async def test_v2_first_event_does_not_wait_for_the_query_vector(clients, monkeypatch):
|
||||
|
||||
Reference in New Issue
Block a user