100 Commits
Author SHA1 Message Date
Ray 6d23caf416 Unify the document tree across local and cloud (#541)
get_tree returns one node shape in local and cloud mode: {title, node_id,
start_index, end_index, summary, text, nodes}. page_index and
prefix_summary no longer appear. The SDK only renames fields on the way
out, so a document indexed before keeps its own ranges and summaries.

New local indexes, standard and flash:
- A parent whose first child starts on a later page gets a first child
  "<parent title> (intro)" that holds those pages.
- A parent's range covers its whole subtree, and its summary is written
  from its children's summaries. Standard mode now summarizes with
  summarize_tree, as flash does.
- A node the model leaves unsummarized falls back to its subsection titles
  or its own text.
- The standard large-node split acts on leaves only.

A node's text is its own pages. A parent's runs onto the page its first
child starts on, and is empty when its intro holds those pages.
2026-10-01 23:20:09 +08:00
Ray f0d67c1224 test: summary concurrency test waits for the expected overlap (#547)
The mock held each call for a fixed 20 ms and expected all three to
overlap inside it. On a slow CI runner the three summaries start further
apart than that, and the control run read (3, 2) instead of (3, 3).

Each call still holds 20 ms, so a lane with no cap still lets the others
in. It then holds until its lane reaches the expected peak, with a 5 s
deadline. An ignored cap and too little overlap both still fail.
2026-10-01 21:50:16 +08:00
Ray 60c8ff8a6f feat(sdk): node navigation helpers get_node, get_node_parent, get_node_path, get_node_map (#542)
They take the tree, not a doc_id: fetch it once (get_document_structure)
and navigate locally, instead of one request per step.

get_node / get_node_parent / get_node_path share one depth-first walk,
O(n) per call. They return the tree's own nodes rather than copies
(unlike get_nodes / get_leaf_nodes), so node["nodes"] keeps working.
get_node_map is create_node_mapping plus the input check;
create_node_mapping stays as the 0.2.8 surface.

Wrong input raises TypeError naming get_document_structure(doc_id) /
get_tree(doc_id)['result'] instead of reading as "not found": the whole
get_tree() response, a None tree, or a non-string id.

create_node_mapping (so get_node_map too) walks past a node whose nodes
is None instead of raising TypeError, as get_node already did.
2026-10-01 21:33:16 +08:00
Ray f279431eb4 Restructure repo layout (#539)
- Remove docs/naming-rules.md
- Move examples/documents/results/ to examples/results/
2026-09-30 16:22:34 +08:00
Ray d2693d8079 ci(publish): gate releases on the live cloud tests again (#538)
Reverts #537.
2026-09-29 23:50:13 +08:00
Ray 619cbd89f6 ci(publish): run the release gate without the live cloud tests (#537)
The CI PageIndex account is over the 1,000 active-page free tier, so every
test_live_* call returns HTTP 403 and blocks the release. Without
PAGEINDEX_API_KEY the live tests skip; the offline suite still gates the publish.
2026-09-28 22:01:02 +08:00
Ray 037a7dbacf perf: summaries run deepest-first and start while expand is still deciding (#432)
Flash indexing spends most of its wall time in summaries, and until now that stage waited for expand to finish and then ran its calls in whatever order the tree recursion produced. This branch makes the summary stage run deepest node first and start while expand is still deciding, so the LLM channels never sit idle waiting on the expand chain.

**What changes**

- `_PriorityGate`: the summary semaphore admits the queued call with the most work still above it (depth = calls left on the node's path to the root, its own included), FIFO within a depth. Cancellation-safe like `asyncio.Semaphore`.
- Tasks are created deepest node first, so the first admissions are the deep leaves rather than whichever shallow leaves the recursion reached first.
- `summarize_tree` becomes a thin wrapper over `SummaryScheduler`: `mark_final(nodes)` says those nodes will not gain, lose or swap children and starts their subtrees; `finish()` awaits the roots. Same task order, gate and error semantics as before.
- `optimize(on_final=...)` reports which nodes are final as it goes: after each round's merges, at each expand candidate's decision (together with what it grew), and for the whole tree at the end. A node is final when it is collapsed under the trigger, collapsed and already judged by expand, or has children — the cost merge cannot fire on a surviving node after the first round (see the commit message for the argument).
- Same-page fusion moves to where duplicates arise (right after a collapsing merge, right after expand attaches children) instead of the next round's start, so no node waits a round for it. The nine corpus PDFs produce byte-identical merge-only trees; SpaceX just stops after two rounds instead of a third that did nothing.
- `page_index_flash` runs expand and summaries on one event loop when both are on; every other combination keeps the old path.

**Measured** (same hour, end to end via `submit_document`)

| | before | after |
|---|---|---|
| fed-2023 (222 p) | 97.9 s | 72.6 s |
| PRML (758 p) | 174.3 s | 136.8 s |

Summary-stage only (fed, 182 calls, 64 wide): FIFO 58–62 s → gate 50–57 s → gate + deepest-first 45 s.

Same calls, same prompts; outputs are order-independent. Peak in flight is now the expand cap plus the summary cap (32 + 64).

**Tests** cover the ordering, cancellation, scheduler, final-node reporting, immediate-fusion and one-loop overlap cases, and every knob's path from the client and the CLI to the model calls.

**Summary prompt and indexing knobs**

The summary prompts no longer ask for the `points` list that `parse_summary` discarded, and cap the summary at `summary_max_words` (default 150). Measured on gpt-5.6-luna, mirror A/B, summary stage only: per-call latency 9.7 → 5.3 s (−45%), fed-2023 47.5 → 30.7 s (−35%), PRML 71.1 → 38.1 s (−46%), output tokens −65%. Summaries come out ~1160 chars instead of ~670 and carry the specifics that used to sit in the discarded list; a blinded pairwise judge (claude-sonnet-5, source in view) prefers them 21-1-0 over the old ones. Deleting the list without a cap is not enough: the model then pours it into the summary (3× longer) and parents slow down more than the leaves gain.

Four indexing knobs are settable from the SDK (flat arguments or the `index=` slot) and the CLI: `summary_max_words`, `summary_concurrency`, `use_embedded_toc`, `optimize` (`"full"` / `"merge"` / `"off"`). `summary_concurrency` bounds both lanes: expand's gate becomes min(32, the cap), so one knob lowers the whole indexing lane on a tight quota (the lanes overlap, so up to cap + min(32, cap) calls run at once). Defaults are unchanged.

The two summary knobs are flash-only: `submit_document(mode="standard")` refuses them rather than index without the cap, as the CLI already does. Both must be positive integers, checked before the PDF is opened; a direct `page_index_flash` call that passed `0` (read as the default until now) or a whole-number float such as `8.0` now raises `ValueError`.
2026-09-24 19:42:46 +08:00
Ray 8eff773368 Anthropic lanes: send the model as written; the prefix picks the route (#443)
- chat(protocol="messages"), anthropic_runner_config() and claude_agent_config() honor chat_model: a model you set carries over; the SDK's stock default is never sent.
- Model names are sent as written; only the routing prefix is read. bedrock/, vertex_ai/ and azure_ai/ select AnthropicBedrock, AnthropicVertex and AnthropicFoundry (and the matching CLAUDE_CODE_USE_* switch for Claude Code); anything else goes to Anthropic directly.
- Bedrock: InvokeModel rejects top-level cache_control for Opus 4.6 and earlier, so the Messages lane moves an explicit breakpoint onto each turn's newest tool result; anthropic_runner_config() omits the key on that route.
- The max_tokens default is 8192, or budget_tokens + 8192 with thinking; no LiteLLM ceiling lookup, no claude-3 list.
- The anthropic extra requires >=0.122.0 (the Bedrock/Vertex tool runner).
2026-09-24 19:20:36 +08:00
Ray 9c4c3ff2dd test(sdk): cover list_documents regression boundaries 2026-09-21 17:27:50 +08:00
Ray ed06ec69ed fix(sdk): list_documents review followups
- Forward recursive value instead of hardcoding True (#3)
- Raise limit cap from 100 to 10000 to match server (#4)
- Accept folder_id='root'/'' in local mode (#6)
- Fix wire test to assert presence not identity (#6, #9)
- Delete restating comments (#15)

Addresses PR #522 review findings 2-7.
2026-09-21 17:27:50 +08:00
Ray 91238b33b6 README banner: drop fixed height so it scales on PyPI (#528) 2026-09-21 12:19:31 +08:00
Ray 71714e86d9 get_document_id accepts a path: strip the folder prefix since names are unique 2026-09-20 18:05:03 +08:00
Ray 6674d6c4b4 resolve_citations: paired cite tags leave no closing tag behind (#525)
<cite doc= page=>quoted</cite> and <cite doc= page=></cite> are tag
shapes the parser accepts and the chat renderer matches, but the
rewrite replaced only the opening tag, so the display text kept a raw
</cite>: "[[1]](#pageindex-citation-01)quoted</cite>". The chat UI hides
that because its HTML pipeline drops a stray end tag; this text goes to
any host. _CITE_TAG_RE now takes an optional "text</cite>" tail and the
link keeps the text. A closing tag is consumed only together with the
citation it closes, so a tag left unresolved, or a </cite> in quoted
HTML, stays as written. The parser reads group 1 as before: 0 diffs on
30,000 random tag inputs, and 1 MB bodies scan in about 15 ms.

The Returns section said every block-level citation carries bbox,
block_type and text. get_citations() adds them only when the block
could be read, so a local document or a 403/404 block has none, and
c['bbox'] raised KeyError for a reader of this docstring alone.

The highlight test drew on a 1000 px square at scale 1000 and probed
one pixel on the corner of the transposed rectangle, so an x/y swap, an
ignored scale and a whole-page fill all passed. It now draws on a
500x2000 page and checks one pixel inside the region and one outside;
all three mutants fail it.
2026-09-20 17:47:00 +08:00
Ray 0445ee80bd feat(sdk): align upload and folder names with shared naming rules (#524)
* feat(sdk): align upload and folder names with shared naming rules

* fix(sdk): normalize filenames before multipart upload
2026-09-20 17:33:34 +08:00
Ray 8e386ff626 Citation display helpers: numbered links, region highlight, image URLs, folder paths
BREAKING: resolve_citations(), released in 0.2.17, is renamed
get_citations(); same arguments, same list. resolve_citations() now returns
{'answer', 'citations'}: the answer with each citation tag replaced by
[[i]](#pageindex-citation-0i), and the get_citations() entries led by
'anchor' and 'index'. Numbers follow the tags as written, so a repeated
citation reuses its number and a block cited under the wrong page still
gets its link while its entry carries the block's real page. One shared
_citation_key() parses a matched tag for both the parser and the rewrite.

highlight_region(image, bbox, scale=1000) draws the cited region on a page
image (PIL.Image or bytes in, PIL.Image out). Pillow becomes a dependency.

get_page_image(doc_id, page) and get_document_image(doc_id, img_id) return
short-lived URLs. Cloud-only. They need the compute routes that return
presigned URLs; against an older server they raise PageIndexAPIError.

get_document_path, get_folder_path and get_folder_id translate between ids
and readable paths. Paths are built from list_folders() parent links,
because the server's `path` field exists only on /docs listing rows. A path
shared by two folders raises instead of picking one.
2026-09-19 21:53:39 +08:00
Ray 2a7363692f list_documents: pass through the API's recursive flag
Listing a folder returned only its direct contents, with no way to
reach the documents in its sub-folders; the endpoint has taken
recursive all along, the SDK just never sent it. Listed documents also
carry a folder path, served by the cloud.

Local mode accepts the flag and ignores it — there are no folders to
descend into — and reports the same keys a cloud listing has, path
among them, always None.
2026-09-19 21:52:43 +08:00
Ray bd72efe09d Tighten instruction parity alarm (#516)
- Markers: 'folder' → 'folder_id', 'get_folder_structure' → 'get_folder_structure(',
  'recursive' → 'recursive='; the browse-folder routing line moves to exact-match
- Section order: shared headers must appear in the same relative order
- None guard: _assert_instructions_local_parity(None) → readable failure
- Format enumeration: live test reads prompts/list and asserts the server's
  cited_answer format set matches LOCAL_CITATION_PROMPTS
2026-09-18 14:25:40 +08:00
Ray a64bc4f7fd ci: pass PAGEINDEX_API_KEY to publish.yml test gate (#514)
#512 added the secret to tests.yml but missed publish.yml — the
release gate still skips every live parity test.
2026-09-18 14:14:13 +08:00
Ray c840b69b25 Sync browse_documents query: "Required when sort=relevance" (#515)
Sync browse_documents query description with server (#486)

pageindex-chat #486 fixed 'Only used' → 'Required when sort=relevance'.
Sync the local contract and snapshot to match.
2026-09-18 14:11:00 +08:00
Ray 18eb5c9b3c Review fixes: unify format validation, pin key set, self-check exception list
- Hoist format guard above the api_key branch so both lanes reject
  unknown formats (cloud was silently returning the markdown prompt)
- Docstring: 'any other value raises' (the server does not reject)
- assert set(LOCAL_CITATION_PROMPTS) == {markdown, cite} pins keys
- assert set(_LOCAL_ONLY_LINES) <= frozen catches stale entries
2026-09-18 01:16:40 +08:00
Ray fd3c12ea04 Detect deleted shared agent instructions in parity checks 2026-09-18 01:16:40 +08:00
Ray d059a8b378 Sync browse_documents query description with server
The server (PR #484) changed the query parameter description from
"must be omitted when sort='time'" to "Only used when sort='relevance'".
Update the local contract and snapshot to match.
2026-09-18 01:16:40 +08:00
Ray bf0645cc28 Drop the footnote citation format; alarm on instruction drift
pageindex-chat removed `footnote` from the cited_answer prompt
(2fc3d298), leaving the SDK advertising a format the cloud now rejects:
`citation_prompt("footnote")` raised a raw ProtocolError against the
cloud while local mode happily returned the frozen copy — the same call,
two behaviours.

Also adds the drift alarm the instructions never had. The tool contract
and the citation prompts each have a live parity test; AGENT_INSTRUCTIONS
only had a non-empty check, so a cloud edit to shared guidance could
leave local agents on stale instructions unnoticed. The new test lets a
cloud line through only when it names a tool or parameter local mode does
not have.

Verified against the live server: 8 live tests pass, 167 offline.
2026-09-18 01:16:40 +08:00
Ray d44799f33f ci: enable live parity tests in CI (#512)
ci: pass PAGEINDEX_API_KEY to pytest so live parity tests run

The 6 live tests (drift alarms that verify the local contract against
api.pageindex.ai) are gated on PAGEINDEX_API_KEY. Without the secret
they silently skip, making the CI green but the guards ineffective.

The secret must still be added in repo Settings → Secrets → Actions.
Fork PRs never receive it (GitHub default for `pull_request` trigger).
2026-09-17 18:07:15 +08:00
Ray 8fede179c6 Parity test: Sources line no longer served by the cloud (#508)
Chat PR #480 removed the Sources trailer from the MCP server's
cited_answer prompt, so the only dropped line is get_document_image().
2026-09-17 03:02:22 +08:00
Ray ac0eed7444 v0.2.18: flash improvements, prompt alignment, get_document_id (#507)
* Flash: drop the size and landscape bail, fall back to page nodes

extract_toc refused any PDF under 300 text weight, or 200 on its densest
page, or with mostly-landscape pages, and returned an empty outline. Both
rules threw away documents the detector handles: three to five pages of
heading plus a line yield every heading, and an ordinary slide deck yields
one node per slide. Only the no-alphabetic-script condition stays, and its
result now says so with toc_source="unreadable" instead of borrowing
"detected"; toc_source is set on every result (detected, bookmarks, pages,
unreadable).

When detection yields nothing for a document that has text, the pages are
the tree: page_index_flash emits one node per text-bearing page, titled by
its first block and labelled toc_source="pages". A flat tree larger than
FLAT_TREE_MAX_NODES is returned without the optimize and summary passes,
since the managed pipelines refuse it: flash_rejection_reason() gives the
local client and the CLI one policy, pointing at standard mode for an
oversized flat tree, and for an unreadable PDF naming the missing text
(standard mode stays the hint there too, since a PDF of nothing but
numbers lands in the same bucket and the model can still read it).

The nine example PDFs extract byte-identically before and after: the
removed rules never fired on ordinary documents.

* Flash: layout decides, never script; the page fallback covers every page

extract_toc refused a PDF whose dominant script fell in Scholar's "other"
family, which is not a script but a bucket: Arabic, Hebrew, Persian, Urdu,
Devanagari, Bengali, Tamil, Thai, Khmer, Georgian, Armenian, Amharic, and
any document with no letters at all, each reported as "no alphabetic text".
Three more rules keyed on script inside the estimator: an unnumbered heading
in a script other than the body's was dropped, so a Chinese report lost its
English section titles; a kana-majority Japanese document had every detected
heading discarded; a mostly-landscape document picked its title from page one
without the body-paragraph check the portrait path applies, so a slide deck's
title became slide one's body text. These are Scholar's precision-over-recall
scope limits for an index of Latin and CJK papers; on PageIndex's default
local mode they were silent refusals and silent losses. They are deleted
here, in this repo's copy of the port, together with the Cyrillic-only
density threshold; the private scholar/ tree stays a faithful port and the
new tests guard the fork. With the rules gone the same layout yields the same
three headings in Japanese, Hindi, Arabic, Hebrew, Thai and mixed
Chinese/English, and English is unchanged. Language now only decides which
cues are available: case, keyword tables, numbering styles.

The page fallback emitted a node only for pages carrying text, each spanning
one page, so a partly-OCR'd 20-page PDF indexed as six islands with fourteen
pages in no node's range and unreachable by retrieval, accepted without an
error. Titles came from the page's first raw line, on a real document the
running header or the page number. Every page is now a node titled "Page N",
FLAT_TREE_MAX_NODES bounds document pages, and toc_source="unreadable" means
exactly that no page carries text. The refusal says what was measured, no
text layer, run OCR, instead of "no alphabetic text ... try standard mode",
which asserted a cause never checked and pointed at a mode fed the same bytes.

The non-Latin fixtures are PyMuPDF-generated with open-licensed font subsets
embedded (FiraGO, Droid Sans Fallback); tests/data/flash/make_fixtures.py
regenerates them byte-identically.

* Flash parser: a multi-code-point glyph no longer crashes the RTL sign

A ToUnicode value can be several code points: a Devanagari conjunct, a Thai
cluster, an Arabic ligature. _rtl_sign passed the whole value to
unicodedata.bidirectional, which takes exactly one character, so every real
Hindi and Thai PDF raised TypeError in the char merge, before any rule ran.
The first code point now decides the direction, as _reverse_if_rtl already
does for the same values; the empty string is LTR.

* Flash: title scoring stops looking at script; the docs describe page nodes as emitted

The title scorer halved any candidate whose dominant script differed from
the document's, the last place flash changed a decision on script. Without
the factor the nine example PDFs and the four fixture PDFs extract
byte-identically.

get_leaf_nodes reads `nodes` with .get like its siblings, so a flat page
tree, whose nodes carry no `nodes` key, walks instead of raising KeyError.

The README output block and the page_index_flash docstring now say what a
node actually carries: node_id always, `nodes` only with children,
`summary` only when summaries ran, and a flat tree past
FLAT_TREE_MAX_NODES pages coming back unsummarized and unoptimized. The
local client's page-fallback test feeds that real node shape.

* Flash: a hierarchy that starts after page 1 gets a Preface node, as in standard mode

Retrieval reaches a page only through a node range. When the first heading
or the first bookmark sits on a later page, everything before it was in no
node: a memo's first section once its heading became the document title, a
title slide's body, a report's cover, contents and letter to shareholders,
the cover pages of a bookmark tree. Standard mode has always inserted a
Preface node for exactly this (utils.add_preface_if_needed); flash now does
the same, next to the page fallback and before the optimize and summary
passes, and renumbers the node ids.

Of the nine example PDFs, the two Federal Reserve reports and Four Lectures
gain the node (pages 1-4, 1-2 and 1); the other six and every extract_toc
result are unchanged. The four fixtures gain a Preface over their title page.

* Align local prompt copies with the chat MCP server

The frozen citation prompt carried an invented "Sources" footer and
the agent instructions had a pagination tutorial chat never shipped;
the persistence protocol was missing the rephrase-and-retry step chat
has had since April.

* get_document_id(name): look up a document ID by its display name

One API call via the new server-side name filter on /docs, so no
client-side pagination. Cloud and local both covered.

* Fix citation parity test for the Sources line removal

The test asserted the frozen copy matched the cloud exactly minus
get_document_image(); the Sources line is now a second deliberate
omission in the cite format.
2026-09-16 21:17:37 +08:00
Ray bfbd4b305c Flash: layout decides, never script; the page fallback covers every page (#502)
Flash returned an empty structure, and `submit_document(mode="flash")` and the CLI a hard error, for any PDF under 300 text weight, under 200 on its densest page, or with mostly-landscape pages. Both rules threw away documents the detector handles. Four more rules keyed on the document's script: the "other" script family (Arabic, Hebrew, Persian, Urdu, Devanagari, Bengali, Tamil, Thai, Khmer, Georgian, Armenian, Amharic, and numbers-only text) was refused as "no alphabetic text"; an unnumbered heading in a script other than the body's was dropped, so a Chinese report lost its English section titles; a kana-majority Japanese document had every detected heading discarded; a mostly-landscape document picked its title from page one without the body-paragraph check, so a slide deck's title became slide one's body text. These are Scholar's scope limits for an index of Latin and CJK papers; on PageIndex's default local mode they were silent refusals and silent losses.

**What changes**

- Layout decides, never script. The size and landscape bails, the script gate, the cross-script heading drop, the Japanese outline nullifier, the landscape title branch, the Cyrillic-only density threshold and the title scorer's cross-script penalty are deleted from this repo's copy of the port; the private `scholar/` tree stays a faithful port and the new tests guard the fork. Language now only decides which cues are available: case, keyword tables, numbering styles.
- When detection finds no hierarchy, `page_index_flash` returns one node per page titled `Page N`, covering every page, labelled `toc_source="pages"`. A flat tree over `FLAT_TREE_MAX_NODES` (10) pages comes back without the optimize and summary passes and is refused by the local client and the CLI through one shared `flash_rejection_reason()`, pointing at standard mode.
- Every page is in some node. A hierarchy that starts after page 1 (a memo whose first heading became the document title, a title slide, a report's cover and contents, a bookmark outline that begins on page 3) is preceded by a `Preface` node covering the pages before it, the node standard mode has always inserted for the same case; until now those pages were reachable from no node.
- `toc_source="unreadable"` means exactly that no page carries text; the refusal says so and points at OCR, not at standard mode, which would receive the same bytes.
- The character-level parser no longer raises on a glyph whose ToUnicode value is several code points (a Devanagari conjunct, a Thai cluster, an Arabic ligature); real Hindi and Thai PDFs used to fail with a `TypeError` before any rule ran.
- `toc_source` is present on every result: `detected`, `bookmarks`, `hybrid`, `pages`, `unreadable`. The README and the `page_index_flash` docstring list them, and describe a node as emitted: `node_id` on every node, `nodes` only on entries with children, `summary` only when summaries ran.
- `get_leaf_nodes` walks a flat page tree instead of raising `KeyError` on a node without a `nodes` key; it was the one tree helper reading the key unguarded.

**Behaviour change**

Small documents, slide decks, and Japanese, Arabic, Hebrew, Indic, Thai and mixed-script documents that used to fail flash indexing or lose headings now index; with the rules gone the same layout yields the same headings in every one of those scripts, and English is unchanged. A garbage text layer that still has layout structure now indexes as a garbage-titled tree instead of being refused. A Chinese-body report whose cover sets an English title over a Chinese subtitle now picks its title by layout; the deleted penalty could hand `doc_title` to a body paragraph. `extract_toc` yields the same nine example trees, node for node, before and after; `page_index_flash` adds the `Preface` node to the three whose hierarchy starts late (the two Federal Reserve reports, pages 1-4 and 1-2, and Four Lectures, page 1), the node standard mode already gives them, and leaves the other six identical.

**Tests**

Fixtures for Japanese, Chinese with English headings, Hindi and Arabic under `tests/data/flash/`, PyMuPDF-generated with open-licensed font subsets embedded; `make_fixtures.py` regenerates them byte-identically. Green on all three CI legs locally (with and without agent frameworks, pypdfium2 4 and 5).
2026-09-13 18:05:42 +08:00
Ray a30a805151 Block citations resolve to bounding boxes: get_block() and resolve_citations() (#503)
* Block citations resolve to bounding boxes: get_block() and resolve_citations()

resolve_citations(answer, doc_ids=None) reads the <cite doc= page= block=/>
and <doc=…;page=…;block=…> tags of a cited answer and returns one entry per
distinct citation with the document id and, for block-level citations on
cloud documents, the block's page, bbox, type and text from the new
GET /doc/{doc_id}/block/{block_id}/ route, wrapped as get_block(). Local
mode stays page-level; a name shared by two documents raises instead of
guessing; a block the document lacks keeps its entry without a bbox.

* resolve_citations: paired quotes, tag-edged names, tolerant block lookups

_CITE_ATTR_RE pairs its quotes, so an apostrophe inside a quoted name
(Moody's Outlook.pdf) stays in the name instead of ending it.

_OLD_CITATION_RE stops a name at < or >, so a <doc=name> tag without its
own ';' no longer swallows the text up to the next tag's ';' and that
citation with it.

The block lookup keeps its bbox-less entry on a status-less raise (local
mode, where pages have no blocks) and on 403, as doc_targeting_block
does; a transport failure still propagates with its status.

The library listing is _all_documents(), which tolerates an absent
'total' and ends on the empty page; ids are deduplicated before the
collision check, so a document re-served after an upload shifted the
window is one document, not a collision with itself.

doc_ids=[] fails loud like chat(doc_id=[]) instead of resolving nothing.

Docstrings: the bbox coordinate space on both get_block(); the
shared-document promise dropped, since the metadata route is owner-only.

* Citation tag scan is linear

_CITE_TAG_RE's second alternative (the paired form) was unreachable: the
first alternative matches the opening tag whenever a '>' follows, and
without one neither alternative can match. Dropping it, stopping the tag
body at '<' and letting the group absorb the whitespace makes an
unterminated '<cite ' scan linear: '<cite ' plus 4000 spaces went from
2 minutes to 0.1 ms, 99 KB of unterminated tags from 27 minutes to
about a millisecond. Output is identical on every shape tried, the
paired form included.

* Finish the tag-edge and listing-entry sweep

Two holes of the kinds already closed, left open one line and one loop
away from their siblings.

The old format's block group still crossed tag boundaries the way its
name group used to: "<doc=a.pdf;page=3;block=p3_text_5 <cite doc=..." put
the whole following citation inside block_id and, since a block lookup
that 404s is now kept without a bbox, lost it silently. It stops at the
tag's edge like the name beside it; a block id has never contained < or
>, the server writes p{page}_{type}_{seq}.

The library listing reads entries with .get() like the rest of the
codebase does: _all_documents() tolerates a listing whose shape varies,
so reading its result with raw subscripts put the KeyError back one
level down. An entry without a name or an id cannot answer a citation
either way, so it is skipped.

* resolve_citations: block ids trimmed, linear attribute scan, doc_id= like chat()

add() trims both values itself: the name, which each caller used to
trim, and the block id, which neither did. "block=p3_text_5 >" and
block=" p3_text_5 " now look up p3_text_5. A padded id 404s, and a 404
is kept without a bbox, so the padding lost the bbox silently and handed
the padded id back to the caller.

_CITE_ATTR_RE anchors the attribute name at a word boundary. Without
\b the greedy (\w+) restarted at every character of a long word inside
a tag body: '<cite ' + 64 KB of x + '>' cost 19 s on the caller's
thread. Anchored, 64 KB is 1.6 ms and 1 MB 27 ms. Output is identical:
the leftmost match of (\w+)= always begins at a word boundary (0 diffs
on 30,000 random inputs). The linear-time test now reaches this
expression; its inputs had no '>' and never did.

The scope parameter is doc_id, the name chat(), chat_completions() and
document_context() give the same str-or-list argument, and the one the
docstring already used to describe it. The integration builders keep
their doc_ids.
2026-09-13 17:03:33 +08:00
Ray c75800506c Three docstring truths the 0.2.16 merges outran (#501)
- chat(instructions=) called the managed system prompt the carrier of
  "the tool guidance and the document context". Targeting moved to the
  first user message, so it carries the tool guidance alone.
- chat(citations=) said the managed endpoint's resolved citations come
  back "only from chat_completions". They ride the response envelope,
  which protocol="chat_completions" returns whole as well; what cannot
  reach them is the answer lane, which returns the answer string.
- Rejoins a line as_claude_mcp()'s paragraph lost to a merge.

Each clause was written on a branch cut before the change that voided it,
and survived the merge because neither side conflicted.
2026-09-11 17:36:02 +08:00
Ray fc2dc41c16 Client instructions: a standing persona for every answer surface (#497)
* Client instructions: a standing persona for every answer surface

PageIndexClient(instructions=...) sets standing guidance for the answering
agent — persona, language, format — appended after the managed system
prompt wherever an answer is produced: chat() and chat_completions() on
both engines, the Responses and Messages protocol lanes, agent_instructions()
and the three *_agent_config() bundles. It is a client-level argument like
mode=, not a chat-side spelling: it combines with any chat= and never
selects own-model chat on its own. Blank configures nothing, as
chat(instructions="") does. chat(instructions=) adds to it per call; the
prompt order is managed base, client, call, history system rows.

One insertion point serves every own-model surface (_base_instructions).
The managed cloud chat takes exactly one system message, first: the
client's instructions, the call's, and the history's system rows now fold
into it, in that order — so chat(instructions=) works on a managed client
(refused since #460, although the endpoint has accepted custom
instructions since Sep 1), and a system row anywhere in the history no
longer 400s. The answer lane's messages contract is the same on both
engines: text history only — the endpoint refuses tool rows and
structured content itself; the SDK says so first, with the protocol-lane
pointer.

Live-verified against the production managed chat and the live MCP
instructions.

Claude-Session: https://claude.ai/code/session_01J8fbpdM5pz2JLjiNY11usy

* Managed fold: always send the canonical history

The managed payload is now the same shape every time — one leading
system row when there is any system text, then the role/content history —
instead of forwarding the caller's list untouched when nothing folded. That
branch let a blank system row sit mid-history and reach the endpoint's
"only one system message, first" refusal; blank text now configures
nothing, like a blank instructions= does.

Claude-Session: https://claude.ai/code/session_01J8fbpdM5pz2JLjiNY11usy

* Managed fold: lift system rows only, forward the rest verbatim

The managed chat_completions lane ran the whole payload through the
own-model validator: tool rows, structured content and tuples were
refused before the wire, and every field beyond role/content was
stripped (tool_calls, name, the endpoint's own citations). The SDK is
a thin skin over the cloud. It folds the client's instructions and the
history's system/developer rows into the one leading system row the
endpoint takes, and sends everything else as given; the endpoint
decides what it accepts.

Also:
- _managed_instructions drops blank system texts, as the managed fold
  and _anthropic_system already do, so both engines build the same
  prompt for the same input (and share one prompt-cache key).
- An end-to-end test drives chat() on the own-model lane and asserts
  the persona and the call's system text reach the model; the
  white-box builder tests alone stayed green with the persona removed
  from the answer lane.
- as_claude_mcp(): the MCP-instructions channel carries the tool
  guidance only; the client's instructions ride system_prompt.
- Comments that restated code or an assertion removed; the chat=
  combination test asserts chat_model too. Blank instructions
  configure nothing at the constructor, as chat(instructions="") does.

Claude-Session: https://claude.ai/code/session_01TgdXZx63aMwrpaKFTQFFKx

* Chat history: normalize the container, validate only what the SDK reads

Both lanes now take any iterable of message dicts. chat()'s instructions
prepend expanded only lists, so a tuple or generator on the managed lane
silently dropped the call's instructions once _require_own_chat no longer
refused it; own-model rejected the same shapes outright. list() once at
each reader instead. None and other non-iterables fail with Python's own
TypeError, as in the OpenAI SDK.

_system_text refuses a system row carrying non-text parts instead of
keeping the text parts and dropping the rest: the SDK folds that row into
its own system text, so it is the reader and must say what it could not
read. The endpoint stays the authority on every row the fold leaves in
place.

Restore the messages docstring sentence 108875b rewrote: tool-role turns
are rejected on both engines (the endpoint 400s them), so the managed
lane does not take them "verbatim"; rename the test that carried that
claim. Cover the blank-text filter, which no test guarded, and drop a
truncated comment.

Claude-Session: https://claude.ai/code/session_01UBpLu7TJUvrvLvFacs8WYT

* fix: ctor type error names the answer, not a lane managed clients cannot use

A plain managed client is the caller most likely to pass Messages-style
blocks as instructions=, and the old message sent it to
chat(protocol="messages"), which that same client refuses. Say "pass
text", and make the block pointer the exact working call: model= is
required on that lane, and it needs a chat_model= client.

Claude-Session: https://claude.ai/code/session_018od4o3YBqrny2vzGhrX4q5

* Rebase onto the merged lanes: two pins the lift and the targeting move void

The chat_completions protocol lane's guard test still pinned the managed
refusal of instructions=, which this branch lifts on purpose; its own
managed-fold tests cover what the lane now does. And _anthropic_system
lost its doc_id argument when targeting moved to the first user message,
so the prompt-order assertion calls it with what it takes.
2026-09-10 20:05:27 +08:00
Ray a3364c8a1b Document targeting moves to the first user message; folders join chat() (#495)
* Document targeting is conversation content: document_context() replaces doc_id on the agent surfaces

The doc-targeting block ("The user has specified document: ...") was
appended to the system prompt by agent_instructions() and the three
*_agent_config bundles, and placed as its own system block by the Messages
chat lane. A per-request target in the system prompt breaks the cached
prefix, sticks across turns, and on cloud splices SDK text onto the
server-served instructions. The cloud's managed chat puts the same block at
the head of the conversation, so every surface now does the same.

- chat(): all three lanes prepend the block as the first user message
  (the Messages lane moved off its system block).
- New client.document_context(doc_id): the block text for callers who own
  the conversation (the framework routes), to lead their first message.
- BREAKING: doc_id removed from agent_instructions(), openai_agent_config(),
  anthropic_runner_config(), claude_agent_config(), agent_tools(),
  as_openai_tools(), as_anthropic_tools() and as_claude_mcp(). The
  tool-layer allowlist was local-only and raised on cloud; it stays
  internal to local chat(doc_id=).
- The block is one get_document per id: names are unique per library
  (uploads suffix a taken name), so the shadow check and its listing sweep
  went. The cloud endpoint carries metadata, so local get_document() gained
  the key for parity. Rendering matches the cloud's: one document is an
  object, several a list.

Live-verified on a local store and the cloud library across chat() (answer
lane, chat_completions, responses, messages) and the OpenAI Agents,
Anthropic tool_runner and Claude Agent SDK routes.

Claude-Session: https://claude.ai/code/session_01VxguPoTv2BmzS9d3erzTrS

* Agent surfaces: keyword-only tails so a stale positional doc_id raises

doc_id sat first in the positional list on agent_instructions,
openai_agent_config and claude_agent_config, and second on
anthropic_runner_config and as_claude_mcp. With it removed, a v0.2.15 call
such as openai_agent_config("pi-a") no longer failed: the id landed on
include_management (truthy, so the management tools came along and
targeting silently vanished) or on server_name. A bare * after the
surviving leading positional turns those calls into an immediate
TypeError; every in-repo caller already passes keywords.

Claude-Session: https://claude.ai/code/session_017FumozBm2xbT2SG6WBxjMe

* Fix the delegate build_claude_mcp missed by the keyword-only tails

5451e4e put a `*` on the client methods that dropped doc_id, so a stale
positional call raises instead of landing on the next parameter.
build_claude_mcp, which as_claude_mcp delegates to, dropped doc_ids the
same way and let server_name move into its slot: a stale
build_claude_mcp(client, False, ['pi-a']) built an in-process server
named ['pi-a'], or on cloud returned the http config with the argument
dropped. Now keyword-only, covered by the same test.

Doc truths: document_context no longer claims a later turn can move on
to another document (chat's docstrings say to keep doc_id constant, and
the block is prepended at the head of the conversation, so a retarget
lands behind the history); submit_document's metadata inventory and
CloudAPI.get_document's key list both name the metadata and folderId
keys get_document returns.

Docstring trims: the "names are unique per library" rationale sentence
and the multi-line test docstrings, one of which still described the
listing backfill this PR removed.

Claude-Session: https://claude.ai/code/session_018FoZJ97cuM9hAjrk3goZ6u

* Folders on chat: chat(folder_id=) and folder_context()

The managed /chat/completions already takes folder_id and renders a
folder targeting block ahead of the document block (pageindex-compute
api.py:491, :3182); the SDK had no way to send it, and own-model chat
had no folder analogue of document_context(). Both land here.

folder_context(folder_id) renders that block as the managed chat
renders its own: the name, an id/name/description metadata row (the
server's FolderContext), and the directive to discover the folder's
documents with browse_documents / search_documents(folder_id=...,
recursive=true). One list_folders() call picks the folder by id: the
endpoint has no LIMIT and lists the same userId + api_surface
partition the managed chat validates against, so what the helper finds
is what the server accepts, and it carries the same plan gate, so a
non-Max key fails here rather than mid-run. A missing folder raises.
Cloud-only: local libraries have no folders.

chat(folder_id=) places it. The managed lane sends the field and lets
the server render and gate it; the three own-model lanes lead the
conversation with one user message, folder block then document block,
joined as the server joins them. folder_id joins the prompt-cache key
for the reason doc_id does: the block is byte-identical for every
conversation about a folder.

"root" (and "") targets nothing on every lane and in every mode, as
the server treats folder_id="root": folder_context returns "" and chat
places nothing. folder_id is keyword-only on chat() and trailing on
every other signature, so no positional call shifts.

Tests: the wire field on the managed lane; block order and root/"" on
the shared renderer; folder_id threaded through the chat, responses
and messages engines; the local cloud-only refusal.

Claude-Session: https://claude.ai/code/session_01NsQjN5HMmmX4n5NVsZWYaw

* Keep folder-less cache keys stable; folder_id keyword-only everywhere

_conversation_cache_key splices folder_id into the seed only when one
is set, so a conversation without a folder keeps the prompt_cache_key it
had before folders existed instead of cold-starting on upgrade; the
released key is pinned.

folder_id gets the bare * on chat_completions, _responses and _messages,
keyword-only on every chat surface as on chat(); every caller already
passes it by keyword.

The three *_agent_config docstrings name the context helpers in the
order the lanes place them: folder, then documents.

test_doc_targeting_is_one_lookup_per_document now counts get_document
calls; it accepted a second lookup per id before.

Claude-Session: https://claude.ai/code/session_01T2i8viYf3Cb6wZcjToLEDD

* Rebase onto the chat_completions protocol lane: folder_id joins its exact payload

The lane's payload assertion is exact, and folder_id is a managed chat
request field, so it appears there too — unset on a folder-less call.

* folder_id rides the chat_completions protocol lane too

chat(protocol="chat_completions") reached main while this branch was
open, and it is a managed chat_completions call site of its own: a
folder_id passed there was dropped on the way to the endpoint, scoping
nothing with no error. It carries it now, asserted on the wire beside
the answer lane's.
2026-09-10 20:02:25 +08:00
Ray d1a4477c29 chat(protocol="chat_completions") and the compat note on chat_completions() (#493)
* chat(protocol="chat_completions") and the compat note on chat_completions()

chat() gains the third protocol value: the answer lane's own engine
with its Chat Completions envelope kept (chunk dicts when streaming,
instructions as a leading system row). It is the one protocol the
managed cloud chat serves, so that lane opens without a chat model;
the own-model knobs still refuse there.

chat_completions() stays, unchanged in signature, with a docstring
that marks it as kept for existing code and points new code at
chat(). Error strings that steered callers to it now name the protocol
lane; the cookbook's two cells use chat().

The managed endpoint takes extra_body as its own request fields
(temperature, enable_citations), merged last under the skeleton
refusal, so new code reaches them without the old door.

Claude-Session: https://claude.ai/code/session_015b7YGZ8LvYp2Q3oGNN8Lfc

* Review fixes: type and document the chat_completions protocol value

The protocol overloads name "chat_completions" so the literal narrows
to the envelope dict and the chunk-dict iterator; the protocol arg doc
lists it and notes the managed chat serves it; the show_process texts
no longer promise a transcript the Chat Completions lane has none of.

Claude-Session: https://claude.ai/code/session_015b7YGZ8LvYp2Q3oGNN8Lfc

* Review fixes: name the lanes in error text, finish the chat() steer sweep

- _split_chat_messages rejected tool rows and structured content with
  "use chat(protocol=...)", which is the lane the caller just used now
  that chat_completions is a protocol value; name responses / messages.
- submit_query / get_retrieval still steered to chat_completions(), the
  door this branch demotes to compat; steer to chat() like the rest.
- Three docstring claims narrowed to where they hold: the compat note's
  "everything is chat(protocol=...)" excepts the text-only stream (that
  is chat(stream=True, show_process=False)); system rows join the
  managed prompt only with your own chat model, the managed endpoint
  forwards them verbatim; extra_body is verbatim on Responses, Messages
  and the managed endpoint, while own-model chat_completions splits it
  like the answer lane.
- Tests: drop the warnings guard around a deprecation that does not
  exist and the comments restating assertions; the managed-lane
  skeleton test now calls the lane it is named for.

Claude-Session: https://claude.ai/code/session_014bdfnYbWsejWhuphPt1aSv

* Fix managed chat stream overrides in extra_body

* Clarify chat completions protocol parameter documentation

* chat(): keyword-only after messages, type the protocol value

chat(messages, doc_id, stream, ...) and chat_completions(messages, stream,
doc_id, ...) disagree on positions 2 and 3, and the compat note now tells
existing callers the two are interchangeable. A positional rewrite bound
doc_id=True and stream="pi-1": an unscoped answer, billed, no error. From
messages on, chat() takes keyword arguments only, so that rewrite is a
TypeError. Nothing in the repo passed chat() a positional after messages.

BREAKING: chat(messages, doc_id) and chat(messages, doc_id, stream) by
position no longer bind; spell doc_id= and stream=.

Also: the implementation and catch-all overload now carry the same
Literal as the protocol overloads, so a misspelt protocol fails
type-checking instead of only at runtime; _refuse_skeleton names the
managed lane's door (a system row), since instructions= is refused there;
_require_own_chat's comment no longer claims to be the gate for every
protocol.

Claude-Session: https://claude.ai/code/session_01DCBitejpuoNpD8cmQ9HR7j

* fix: prevent managed chat extras from overriding document scope

* extra_body: one gate for every lane, before any I/O

The managed endpoint's refusals of stream and doc_id sat in
cloud_api.chat_completions, one call site of four. The other lanes
took the same keys and broke worse: an own-model OpenAI backend sent
stream:false under an SSE parser and returned an empty answer, and
LiteLLM raised a bare KeyError. The merge had no Mapping check
either, so a list splatted into the payload as fabricated fields.

_refuse_skeleton now owns all of it: a non-dict is refused, and
stream / doc_id join the refused keys beside the skeleton. chat()
and chat_completions() call it first, so the own-model half no
longer runs doc targeting and the MCP initialize fetch before
refusing. cloud_api is a plain transport again.

The skeleton remedy no longer splits by client type. A system row
works on every chat-shaped lane and instructions= becomes one, so
the two labels were inverted for chat_completions() callers.
Managed chat(instructions=) opens when #497 lands.

Tests: non-dicts and the argument keys at the gate, both public
doors refusing before any lane is entered, a foreign managed key
pinned as forwarded. The keyword-only test builds its client outside
pytest.raises and matches the positional message.

Claude-Session: https://claude.ai/code/session_01RTqXsk6Y3iGn9iXZHrzpv6
2026-09-10 19:37:29 +08:00
Ray 5f7a39e175 Tool-path rate limits: retry at the bridge, then fail the run fast (#492)
* Tool-path rate limits: retry at the bridge, then fail the run fast

A PageIndex cloud 429 (or 5xx) on a tool call used to reach the model as
an INTERNAL_ERROR envelope saying "try again": the model re-called once
with no wait, then wrote the failure into its answer, and chat() returned
normally with no status anywhere. The same 429 before the loop (the
doc_id targeting lookup) already propagated raw.

- McpBridge mounts a urllib3 Retry: 429/502/503 and connection failures,
  three attempts, 0/2/4 s apart or as Retry-After says; read timeouts
  are never replayed (240 s each, and the server may have acted); a
  Retry-After past a minute is a quota, not a blip, so the backoff runs
  instead of sleeping it out. Exhausted, the last response falls through
  to the existing >= 400 branch, so the status_code survives.
- _bridge_invoker re-raises 429/5xx alongside 401/403. The frameworks
  turn a raised tool exception back into model-visible text, so each
  chat() door gets its own escape: the in-process MCPServer's
  failure_error_function lets a PageIndex-caused failure propagate and
  _translate_run_error unwraps it from the framework's wrapper (which
  also un-flattens the 401 case); the Messages lane runs each turn's
  tools through the runner's public generate_tool_call_response() and
  raises before the next model call.
- _model_backend_error keeps the provider's status_code.

Claude Agent SDK tools cannot fail fast: the SDK MCP server converts
handler exceptions into JSON-RPC errors for Claude Code by design.

Claude-Session: https://claude.ai/code/session_014S88dcSz7jykegAWyWZk8E

* Tool-path fail-fast: cover unreachable servers and all 5xx

The bridge retry is now a plain urllib3 Retry: 429 and every 5xx retried
three times at the fixed 0/2/4 s backoff, Retry-After ignored. That drops
the _Retry subclass, whose get_retry_after raised InvalidHeader on a
non-integer header (turning a 429 into "could not reach the server"),
honoured a 60 s Retry-After three times over, and let a 413 carrying
Retry-After replay. 500 and 504 join the forcelist so the invoker's "what
survived the bridge's retries" holds for every status it re-raises. Retry
is imported from requests.adapters, the declared dependency.

The invoker re-raises transport failures too: once the bridge's own
connection retries fail, the model cannot reach the server either, and
the envelope only sent it round the retry loop.

The handshake error blames the API key only on 401/403: a rate-limited
handshake is now a run-terminating error and was telling users to rotate
a working key.

Docstrings on agent_tools()/build_agent_tools and the Anthropic adapter
state the real raise set: 401/403, post-retry 429/5xx, unreachable server.

Claude-Session: https://claude.ai/code/session_013xk3xt9KgHNTjsYmFLKxbu

* fix: surface Messages tool failures before advancing runner

* fix: require urllib3 1.26 for MCP retries

* fix: keep the Messages fail-fast quiet and single-path

Raise ToolError from the tools chat(protocol="messages") runs instead of
the raw PageIndexAPIError: the Anthropic runner log.exception()s any
other exception, so every fail-fast printed a 20-line traceback from
anthropic's internals before the SDK raised its own error. The lane
still records the failure and raises it right after the runner's tool
batch, so the ToolError content never reaches the model.

Drop the two post-loop _messages_fail_fast calls: the runner executes
tools only through the public generate_tool_call_response (0.108.0
through 1.4.0), which checked_tool_response wraps, so they could never
fire. Annotate _pageindex_cause for the py.typed package.

Claude-Session: https://claude.ai/code/session_01CfYSeq8kM7HjfF79TGbsiT

* fix: fail fast on the cloud's account-limit tool errors; one 5xx list

RATE_LIMITED / USAGE_LIMIT_REACHED (pageindex-chat #472) arrive as a
normal tool error inside HTTP 200, already retried server-side; the
invoker re-raises them as 429 / 402, the way a post-retry status
escapes, so every lane fails fast without a per-lane change.

The bridge retries the whole 5xx range, the same range the invoker
re-raises; a non-JSON 200 body (a JSONDecodeError is a RequestException
too) stays a model-visible envelope, since the server was reached. The
MCP stub tests keep the machine's proxy out of 127.0.0.1.

Claude-Session: https://claude.ai/code/session_01F7vqmZdnWeKC9SUBytDrdf

* fix: raise the account-limit escape outside the invoker's own try

_raise_account_limit fired inside _invoke's try, so its escape depended
on 429 and 402 also appearing in the except's re-raise tuple: two lists
in one function that had to agree. The check now runs after the try,
where the except cannot swallow it, and 402 leaves the tuple (the MCP
route never answers HTTP 402; it was there only to let the raise through).

The mcp_stub fixture also sets the lowercase no_proxy: requests reads
that spelling first, so a machine with no_proxy set still routed the
stub requests through its proxy despite NO_PROXY.
2026-09-10 19:34:05 +08:00
Ray ebf19b21db Process display elides an image's base64 payload (#489)
The woven [tool_result] line clipped the framework's image item to its
first 200 characters, which is the head of a data URL. The line now shows
the item as it is with the payload elided: {"type": "image", "image_url":
"data:image/png;base64,..."}. The item itself is unchanged.

Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
2026-09-06 20:42:00 +08:00
Ray d0d6aa39e8 Tool results reach the frameworks as MCP content; the frameworks render it (#487)
* Tool results reach the frameworks as MCP content; the frameworks render it

get_document_image returns an MCP image block, and the SDK flattened every
tool result to text on its way to a framework, so an own-model chat received
"[image/png content omitted: ~256 KB]" where the hosted lanes (HostedMCPTool,
the http MCP config, the managed chat) received the page.

The rule now: the SDK carries MCP types and renders nothing. Each framework
gets the tool set the way it consumes MCP, and converts the content itself.

- McpBridge.call_tool returns the content blocks untouched; the invoke
  contract behind _tool_specs is (content blocks, is_error) on the cloud and
  local paths alike (local tools send one text block).
- openai-agents: an in-process MCPServer over the tool specs, and the
  FunctionTools are MCPUtil.to_function_tool over it. The framework's own
  MCP conversion builds the tools (schema verbatim, strict off) and renders
  results (text items, image data URLs); its failure pipeline answers
  malformed arguments. The hand-built FunctionTools are gone.
- Anthropic tool runner: results validated as CallToolResult and rendered by
  the Anthropic SDK's mcp_content (text blocks, base64 image blocks), on the
  error channel too.
- Claude Agent SDK: the content passes through, it is MCP already.
- Plain functions (agent_tools): render_text, the former flattening, keeps
  the size stub for binary content; a string cannot carry an image.

Wire consequences: tool results on the openai-agents lanes are the
framework's structured items (a single text item for a text result) rather
than a bare string, and the Anthropic tool_result content is a block list.
ChatStream tool_result events carry that structured output; the woven
display shows its text.

mcp becomes a declared dependency: it already arrives with openai-agents,
and both adapters use its types.

Verified live on a 15-page paper, page 1, across chat() on the LiteLLM lane
(gpt-5.6-luna), chat(protocol="responses"), and chat(protocol="messages",
claude-sonnet-4-6): each answer described the red attribution text and the
arXiv stamp rotated along the left margin.

Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW

* Relay test asserts the wire aliases: mcp 1.x and 2.x differ in attribute names

CI installs mcp 2.x, where CallToolResult exposes is_error and content
blocks exposes mime_type as attributes, with the wire names (isError,
mimeType) as aliases; mcp 1.x uses the wire names as attributes. The
adapters only ever validate from and dump to the wire shape, so they run on
both; the test now does the same.

Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
2026-09-06 19:13:20 +08:00
Ray 2858384a75 docs: the chart images load by absolute URL so PyPI renders them (#475)
pyproject publishes README.md as the PyPI description, and PyPI does not
resolve a relative path against the GitHub repo — the five charts render
as broken images on the project page (verified on 0.2.14). Every chart
now loads from raw.githubusercontent.com; the ten other images in the
file already did.

PyPI's sanitizer drops <source>, so only the <img> fallbacks matter
there; the srcset URLs change too, to keep GitHub's dark variants
spelled the same way.
2026-09-04 15:30:44 +08:00
Ray 0400366d35 litellm terminal noise: quiet as a default, and the no-fetch stamp on every lane (#473)
* fix: flip litellm's banner switch at the provider lookups, not after them

_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.

Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* refactor: the preload thread only imports — every litellm entry quiets itself now

Since 30652ca each get_llm_provider caller and every completion helper
flips the banner switch and clamps litellm's logger before touching
litellm, so the background import's own _quiet_litellm() had no path
left that depended on it; it was also the one place that mutated
process-global logging at a moment the caller could not predict. The
env stamps stay: import-time records still need them. The config-time
repair test no longer needs the preload guard pinned.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* fix: litellm quiet is a default, not a clamp — and the no-fetch stamp covers every lane

_quiet_litellm() set litellm's loggers to ERROR on every entry, overriding whatever the host, litellm._turn_on_debug(), or LITELLM_LOG=DEBUG had set. It now sets a level only while a logger is still NOTSET: the default stays ERROR, anything set explicitly wins, and there is no per-call setLevel churn.

LITELLM_LOCAL_MODEL_COST_MAP moves from the client's preload to utils' import. The CLI pipeline and openai_agent_config on a cloud client import litellm without the preload and fetched the remote model map on first use (offline: a WARNING). Every litellm lane imports utils first.

The LITELLM_LOG=ERROR stamp is gone: it pinned litellm's stderr handler at ERROR for the whole process (silencing _turn_on_debug for good) while the logger level was what actually gated the chatter — output verified identical with and without it across plain, logging-configured, and debug hosts.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* style: essential comments only

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* test: the utils-import stamp check lives with its package-surface sibling

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* refactor: drop the CLI's duplicate no-fetch stamp — utils' import sets it

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* fix: the thinking-budget ceiling imported litellm ahead of the stamp

_default_max_tokens is reached from anthropic_runner_config/messages(). On a managed cloud client the constructor never imports pageindex.utils, and local_chat deliberately keeps utils off its import path, so that lane imported litellm with LITELLM_LOCAL_MODEL_COST_MAP unset: the model cost map was fetched over the network (3517 entries vs the bundled 2982) and offline printed the very warning this branch removes. Verified end to end through the public API, before and after.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
2026-09-03 18:37:11 +08:00
Ray 85180eeacd fix: extra_body refuses the skeleton keys; sampling fields ride ModelSettings on LiteLLM-routed models (#472)
* fix: extra_body refuses the skeleton keys

extra_body merges last on every lane, so a caller's system / instructions / input / messages / tools replaced the managed prompt, the conversation or the doc tools wholesale. The run succeeded and answered without tool guidance or document scoping, silently. Those are the SDK's on every lane (the three-layer rule: skeleton, named knobs, caller extras); the door for the prompt is instructions=. Refused at the two seams where extra_body meets the wire, before the Anthropic transport exists on the Messages lane.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit 4d18854378)

* fix: extra_body sampling fields ride their ModelSettings field on LiteLLM-routed models

On the answer lane, extra_body for a non-OpenAI destination becomes ModelSettings.extra_args, the LiteLLM kwargs channel (LiteLLM would plant extra_body as a literal request field, which Anthropic rejects). openai-agents passes ModelSettings' own fields to litellm.acompletion by name beside **extra_args, so a caller's temperature / top_p / max_tokens / penalties / tool_choice in extra_args collided with the explicit keyword and died as a raw TypeError inside the framework. Split by the framework's own contract: keys that are ModelSettings fields ride their field (the caller's value winning over ours), the rest stay LiteLLM kwargs. The field list is the public dataclass, not a hand-kept name list; bad values now fail ModelSettings validation and surface as PageIndexAPIError. response_format is not a ModelSettings field and still has no door on this lane; documented.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit c1ad980505)

* docs: chat_completions extra_body names the skeleton carve-out

chat_completions shares the seam that now refuses the skeleton keys; its docstring still promised an unconditional merged-last win. Same sentence as chat().

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit f9072f9fff)
2026-09-03 18:27:16 +08:00
Ray c8c8293aa7 ChatStream moves to a light module so chat()'s type hints resolve at runtime (#471)
* fix: ChatStream lives in a light module, so chat()'s type hints resolve at runtime

ChatStream was importable in client.py only under TYPE_CHECKING (a real
import would have dragged local_chat's asyncio stack into `import
pageindex`), which left chat()'s return annotation a dangling string:
typing.get_type_hints(PageIndexClient.chat) raised NameError, and so
did anything that introspects signatures — agents' function_tool
(client.chat) died on it before looking at a single parameter. The
class touches neither asyncio nor the agent frameworks, so it moves to
pageindex/chat_stream.py, client.py imports it for real, and the
package exports it directly instead of lazily. `import pageindex` still
leaves local_chat unloaded.

Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY

* fix: the ChatStream move keeps its own invariants — future annotations, a guarded eager path, the old import path pinned

Review of #471 found the move's guards thinner than they look:

- chat_stream.py had no `from __future__ import annotations`, unlike
  every sibling module, which made its `-> "ChatStream"` quotes
  load-bearing: unquoting them — the very edit this move made in
  client.py, and what `ruff --select UP037 --fix` does — broke
  `import pageindex` outright.
- The import in client.py is the whole fix and reads like a
  typing-only one; a comment says why it must stay real.
- The lazy-import test's denylist named no framework, so agents,
  litellm, openai or anthropic could join the eager path with a green
  suite — the cost the module was split out to avoid.
- The type-hints walk had no floor, so it could silently stop covering
  anything, and nothing pinned `pageindex.local_chat.ChatStream`, the
  path the class shipped under in 0.2.11-0.2.14.
- local_chat's module docstring still claimed the class.

All four guards mutation-checked red.

* style: the walk-collapsed assertion message fits the 79-col convention

* style: no rationale comments — the guard is the test, the why is the commit
2026-09-03 18:19:15 +08:00
Ray 564bc97c0f Chat process follow-ups: a mid-stream error raises, .events survives partial reads, bad show_process chokes first (#470)
* fix: .events survives partial reads; hidden call lines still label results; bad show_process chokes first

- ChatStream.events delegated with `yield from`, so a dropped handle
  (next(stream.events), for ... break) closed the shared run on GC and
  the rest of the run silently vanished. A plain loop leaves it alone.
- _weave filled call_args only past the tool_call visibility guard, so
  with call lines hidden the standalone result lines never carried the
  arguments they promise.
- show_process is validated before the stream check: an invalid value
  is refused as such instead of being told to add stream=True and then
  refused again; the managed lane's duplicate choke goes with it.
- Docstring: show_process is not own-model-only.

Claude-Session: https://claude.ai/code/session_016M3qaQedSK7L4DwysFRmk2

* fix: mid-stream error chunk raises instead of ending as a short answer; stream docstring says show_process is on by default

- The managed endpoint reports a server-side failure as a final
  {"error": ...} chunk after the partial answer (api.py refunds the
  credits, then yields it). Neither chunk decoder looked at it, so
  chat(stream=True) and chat_completions(stream=True) in both modes
  ended as an apparently complete short answer with no exception. One
  guard in each decoder raises PageIndexAPIError; the partial answer
  is still delivered first.
- The `stream:` arg and the Returns block still described the pre-PR
  contract (bare text chunks); only the show_process paragraph said it
  is on by default.

Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY
2026-09-03 15:52:20 +08:00
Ray 1ccb31bf9d fix: Responses effort rides extra_body; envelope reports metadata; "" is unset on every lane (#469)
* fix: Responses effort rides extra_body; chat() docs say what the wire does

The rule chat() follows, written down: chat() names PageIndex's own
parameters plus a subset of LiteLLM's unified vocabulary; anything
vendor-specific rides extra_body under the lane's wire names. Three
layers meet on the wire — the skeleton (managed prompt, conversation,
tools) is the SDK's, the named knobs are translated per lane, and
extra_body is the caller's, merged last so it wins. A key the SDK writes
into a nested object must go in through extra_body too, so the caller's
other keys in that object survive — the SDK merge is shallow.

- Responses lane: reasoning_effort joins the caller's extra_body
  "reasoning" object (their keys win) instead of riding a separate
  reasoning= that extra_body's object replaced whole on the wire.
  Mirrors the Messages lane's output_config. The envelope's
  given.get("reasoning") now reports the merged object for free.
- reasoning_effort="" is unset on both protocol lanes, like model=""
  and instructions="".
- extra_body docstring: drop the invitation to override the SDK's
  system/input — that is the skeleton, and a fixed input breaks the
  tool loop; name the wire fields per lane instead of one mixed list.
- messages docstring: every system row joins the managed prompt on the
  answer lane, not only a leading one (_split_chat_messages hoists all).
- max_turns docstring: the OpenAI lanes raise at the cap, the Messages
  lane returns the truncated run — the divergence was documented only on
  the private door.
- Migration message: parameters go by keyword; the doors' sampling and
  thinking fields ride extra_body.
- chat_model docstring: responses is no longer a chat surface.
- _split_chat_messages refusals no longer name chat_completions, a
  method the chat() caller never typed.
- Tests: the Responses door equivalence takes the extra_body shape;
  effort/extra_body collision, key survival and "" pinned on both
  protocol lanes; the ModelSettings spy asserts the values reached the
  wire, not only the envelope.

Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE

* fix: Responses envelope reports the metadata sent

metadata is a Responses request field the caller sets through
extra_body; the envelope hard-coded None. Same source as the other
caller-set fields: given.get().

Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE

* fix: "" is unset on the answer lane; managed-cloud gate reads falsy as unset

reasoning_effort="" reached LiteLLM as a literal empty effort on the
answer lane while the protocol lanes already treated it as unset. And
chat_completions' managed-cloud own-model gate, which chat() routes
through, refused model="" / reasoning_effort="" / {} as knobs the caller
never set. Non-numeric knobs now read falsy as unset, as the local lane
always has (model or chat_model, if extra_body, backend or {}); numeric
ones keep `is not None`.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit b15858b156)
2026-09-03 15:52:16 +08:00
Ray a7995a4d22 Ignore the default results/ output directory (#468)
`run_pageindex.py` writes its trees to `./results` (`output_dir = './results'`,
lines 142 and 198), so running the CLI leaves untracked JSON at the repo root.
That output has already been committed by accident ten times — 51 files, 3.2 MB,
across commits from "first commit" through "improve tree optimization".

`examples/documents/results/` is unaffected: this only ignores the top-level
directory the CLI writes to.

Claude-Session: https://claude.ai/code/session_01KXmNuHrkQEkefNS12V7apV
2026-09-03 14:29:25 +08:00
Ray 5189cf2892 chat(protocol=): the protocol doors move behind the front door (#460)
* feat: chat(protocol=) — the protocol doors move behind the front door

responses() and messages() become _responses()/_messages(): the same
engines, reachable as chat(protocol="responses"|"messages") with the
protocol's own input and output shapes (transcript items / content
blocks in, the envelope or native stream out). The old names are the
vendor SDKs' own, and an agent-written client.messages(...) now fails
fast with the way in — a runtime-only __getattr__, invisible to the
type checker so attribute typos on the client still get flagged.

chat() gains the knobs that lost their public home: instructions
(appended after the managed prompt — a string on every lane, Messages
system blocks with protocol="messages"), max_turns, backend,
extra_headers, extra_body. reasoning_effort lands natively on each
lane: LiteLLM's kwarg, Responses reasoning.effort, Anthropic
output_config.effort. show_process stays the answer lane's view.

Six overloads keep the return types narrow for py.typed consumers; the
answer lane's contract is unchanged. The managed cloud chat rejects the
own-model knobs as before.

Claude-Session: https://claude.ai/code/session_017uvMD9eatgdaupLGFUjuKM

* fix: chat(protocol=) honors extra_body thinking; honest envelope and remedies

- Messages lane: through the public door the thinking budget rides
  extra_body, so the default max_tokens lift reads it there (extra_body
  wins on the wire, so it wins in the lift); the documented
  extra_body={"thinking": ...} no longer 400s on max_tokens < budget.
- Responses lane: temperature/top_p/reasoning/max_output_tokens reach the
  wire via extra_body only; the envelope now reports what was sent instead
  of the never-set locals.
- One _require_own_chat refusal for chat(protocol=...), the doors behind
  it, and instructions. The doors' own copies had already drifted from
  chat()'s text, and every copy pointed a local client with chat_model
  blank at the managed chat it does not have; a door reached directly on
  such a client fell through to a raw AttributeError. The check is the
  one chat_completions already makes.
- Messages lane treats model="" as unset, like every other model check.
- The protocol refusal for show_process runs before the stream=True hint,
  so the first remedy offered is the right one.
- instructions="" configures nothing (no empty system row, cache key
  unchanged), matching the protocol lanes.
- Responses validation names messages, the public parameter, not input.
- Stale docstring pointer to messages() fixed.
- A protocol=None, stream: bool overload restores str | ChatStream for a
  runtime-variable stream; the catch-all had widened it to a 4-way union.
- Tests: door X refuses like chat(protocol=X) on managed and blank-local
  clients; the protocol gate and the non-stream show_process order are
  asserted by distinct messages; chat(stream=True) joins the max_turns
  matrix; each fix carries a red-verified assertion.

Claude-Session: https://claude.ai/code/session_01M9GdjnuDHHwDKqMzvPWCcj
2026-09-02 06:59:55 +08:00
Ray 91b6c365e0 Chat process display: labels are the event type names (#459)
* style: label woven lines by their event type names

[tool_call] and [tool_result] replace the "[tool]" label and the "->"
arrow, so the text view's labels are exactly the .events type names
(config keys stay plural — they switch a class of lines; each line is
one instance).

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* style: tool_result lines flush left — no nesting indent

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* style: one vocabulary — show_process keys are the event type names

tool_call / tool_result (singular) everywhere: event types, text labels,
and now the config keys, which select event types by name.

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
2026-09-02 04:28:38 +08:00
Ray ad7956d3c3 Keep litellm's terminal noise out of answers: stdout banners, logger chatter, retry print (#455)
* fix: mute litellm's stdout Provider List banner in the litellm lanes

litellm's OpenRouter adapter probes supports_reasoning() with the
provider-stripped model name on every completion, so any model missing
from its static map (e.g. openrouter/z-ai/glm-5.3-flash) makes
get_llm_provider print a red "Provider List:" banner straight into
stdout — interleaved with the streamed answer, once per agent turn.
suppress_debug_info is litellm's own embedder switch (its Router sets
it too) and gates only this banner and the "Give Feedback / Get Help"
one; errors still raise with their full text. Applied at the same lazy
hook points as the existing litellm repairs, plus the background
preload, so merely importing pageindex still leaves the host's litellm
untouched.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL

* fix: retry notice rides logging, not the caller's stdout

llm_completion/llm_acompletion printed '* Retrying *' straight into
stdout on every retried request — the channel that belongs to answers
and CLI output. The notice moves to logging.warning beside the error
line that already accompanies it.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL

* fix: gate litellm's stderr WARNING chatter alongside the stdout banners

The banner mute grows into _quiet_litellm: litellm's own logger sprays
WARNING records (remote-map fetch fallbacks, cost hiccups) onto stderr
from inside requests — not actionable for SDK callers, whose real
failures raise as exceptions. The preload stamps LITELLM_LOG=ERROR
before litellm's import initializes its logger (setdefault, so an
explicit caller choice wins, and litellm honors a chosen level itself);
the hook's setLevel covers litellm imported before us, plus the dotted
litellm namespace its adapters log under.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL
2026-09-01 22:16:23 +08:00
Ray 555c370f66 Show the chat run: show_process weaving and the ChatStream events view (#454)
* feat: show the chat run — show_process weaving and the ChatStream events view

chat(stream=True) now returns a ChatStream: iterating it yields the
answer text with the run woven in by default — "[thinking] " sections,
one "[tool] name arguments" line per call with its clipped result —
and .events yields the run as typed dicts (thinking/answer deltas,
tool_call with parsed arguments, tool_result with the full output).
One run serves one view; close() kills it like a closed generator.

- show_process: on by default ("on where available"); False for the
  bare answer stream; a dict (ChatProcessOptions: thinking, tool_calls,
  tool_results, max_chars) selects the parts. Explicit True without
  stream=True raises.
- Managed clients weave what the endpoint serves: tool-call lines
  parsed from its block_metadata chunk tags (that wire carries no
  thinking and no tool results); old-wire chunks stay plain answer
  text, and the bare answer view no longer leaks tool-argument JSON.
- Engine: one typed-event primitive (_chat_events_agen /
  _cloud_chunk_events) with _weave as a pure renderer over it; the
  chat lane's prologue is shared via _chat_agent, behavior unchanged
  on chat_completions/responses/messages.

453 tests green (16 new, red-verified), no-openai-agents leg simulated,
pyright flat vs main.

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* fix: pair woven tool results with their calls; drop empty deltas; validate show_process before the managed request

Three review findings on the show_process weave, all red-verified:

- _weave nested every tool result under the most recent [tool] line,
  which misattributes results when a turn makes parallel calls (the SDK
  streams all calls, then all results). A result now nests only when the
  line above is its own call (by call_id); otherwise it stands alone
  with its call's clipped arguments echoed, so same-name parallel calls
  stay tellable apart. Hidden-call mode falls out unchanged (no stored
  arguments, no echo).

- Empty deltas now stop at the event source. chat_completions and the
  managed chunk lane both filter them; the new local lane did not, so a
  mid-stream "" (litellm forwards annotated/provider-field empties)
  leaked into the bare view and flipped _weave sections, splitting one
  thinking burst into repeated labels.

- The managed streaming lane sent the billed request before
  _process_options ran, so a config typo cost a real chat call. chat()
  now chokes on bad show_process before dispatching, matching the local
  lane's validate-first order.

Two docstring truths: close()'s "a run never consumed never starts"
holds only for own-model chat (the managed request is already on the
wire), and the module docstring now covers the managed chunk weave.

456 tests green; the three new ones red-verified; managed-lane paths
re-run with the agents package blocked; pyright adds nothing on touched
lines.

Claude-Session: https://claude.ai/code/session_01XeeD2214z6Vd6qKcAi9ZvJ

* fix: seal open tool blocks from the answer; typed chokes; chat() stream overloads

Four review fixes plus two coverage gaps, each red- or
mutation-verified:

- _cloud_chunk_events treated any non-tool_use tag inside an open tool
  block as answer text, so argument JSON leaked into the
  show_process=False answer — the one meant to be appended back as
  conversation history. Inside an open block nothing is answer:
  argument chunks now accumulate under any tag. And non-string
  argument pieces stringify at the join instead of killing the whole
  stream with a raw TypeError.

- _process_options sorted unknown keys before repr-ing them, so
  mixed-type keys ({1: True, "foo": 1}) raised a bare TypeError past
  the caller's `except PageIndexAPIError`; sorting the reprs keeps the
  single error type.

- chat() gains @overload on stream, so the docstring's own `.events`
  usage type-checks for py.typed consumers (previously pyright ruled
  `Cannot access attribute "events" for class "str"` on the exact
  documented snippet). pageindex/ error count unchanged (234).

- Coverage: the streaming lane's whole `finally` could be deleted with
  the suite still green — the new abandonment test pins the teardown
  (pump exits, turn 2 emits nothing, _aclose_backend closes the
  per-call client). And FakeModel emitted only the reasoning event
  production never sends (litellm folds reasoning into
  reasoning_content, which arrives as summary deltas); it now
  alternates variants, so dropping either from the isinstance tuple
  goes red.

459 tests green; managed-path tests re-run with the agents package
blocked; flake8 parity on every touched file.

Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP

* fix: reading .events is inert — consuming claims the view; honest falsy show_process message

ChatStream.events was a property whose getter latched the stream's one
view on mere attribute access: a debugger variable pane, hasattr, or
getattr(stream, "events", None) — which PageIndexAPIError escapes, as
getattr only swallows AttributeError — was enough to make a later
`for chunk in stream:` refuse, with nothing consumed. The getter now
returns a lazy generator: the managed refusal, the view claim and the
run start all happen on first consumption, so introspection is
side-effect free and the text view stays usable after a probe.

And the stream=False guard's message told falsy-but-not-False values
("show_process=0", "") that they passed show_process=True; the check
itself is the ruled falsy-{} trap and stands, but the message now
names the off values and echoes what was got.

461 tests green (2 red-verified new: inert read on both lanes, plus
the falsy-message case); changed managed-path tests re-run with the
agents package blocked; pyright pageindex/ 234 -> 234.

Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
2026-09-01 22:16:14 +08:00
Ray 7d18cc033c docs: drop stray blank lines in the cloud quickstart snippet
Claude-Session: https://claude.ai/code/session_01Wh75QagUhr5DPVATjmGrCe
2026-08-31 19:00:00 +08:00
Ray 9fee239b17 docs: correct what the index model does (#441)
* docs: correct what the index model does

The index model does not build the tree structure — Flash extracts it
from the document layout without an LLM. The model only summarizes and
refines the tree.

Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD

* docs: name PageIndex Flash in the submit_document note

Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD
2026-08-28 01:57:30 +08:00
Ray 21f2b1018e docs: FinanceBench chart as local light/dark assets (#440)
* docs: swap the FinanceBench chart for local light/dark assets

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: replace the FinanceBench chart with light/dark assets

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: narrow the FinanceBench chart to 60%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: set the FinanceBench chart to 65%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: set the FinanceBench chart to 70%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF
2026-08-28 00:56:37 +08:00
Ray 688ce200fa docs: hide the FinanceBench accuracy chart
Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF
2026-08-28 00:39:20 +08:00
Ray d947ab6c93 docs: README structure and layout pass (#438)
* docs: the benchmark charts render at 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render centered at 80% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 85% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 90% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the collapsed sections drop the spacer <br>

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the header link row drops Discord

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Updates entry comments out the MCP/API pointer

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the one-line summary stands without the Why it works heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: TEMP four punchline variants side by side

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the one-line summary reads as a centered pull quote

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: TEMP alert-style variants next to the pull quote

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the pull quote sits under an In one sentence heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the pull quote sits under a TL;DR heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote bolds the claims and drops the vector DB and chunking

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote reads no vector DBs or chunking

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote leaves retrieval unbolded

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Retrieve step leads with agentically

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote reads left-aligned

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote wraps on its own

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example uses a plain system-message string

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example keeps the system prompt inside the code block

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example inlines the system prompt

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Cloud example drops the doubled blank line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench case study returns to Benchmarks

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench paragraph speaks of PageIndex directly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence names the corpus and the margin plainly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence drops the corpus aside

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence keeps the benchmark description

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Benchmarks splits into the local open-source run and FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: top-level sections use single-hash headings again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed Local mode

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed PageIndex Local

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 80% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench chart stays at 90% and the benchmark gloss goes in parentheses

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench chart returns to its original 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark charts render at 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed Running PageIndex locally, charts at 75%

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example indents the prompt continuation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations prompt breaks before the cite tag

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations prompt breaks before using

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the collapsed usage guides follow the Quickstart

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a Usage section holds the two collapsed guides

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h3 with their steps one level down

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h2 so their steps keep their levels

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h3, inner levels untouched

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: each usage step collapses on its own under Usage

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Usage keeps its two guides as headings with collapsed items inside

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the *_config note folds into Other agent frameworks

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Usage and the Detailed Usage Guide open with a line of orientation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in drops the happy-path idiom

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in names the two guides directly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the guide and integration lead-ins say what each block holds

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in contrasts direct use with integration

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in lists its two guides; their intros stay general

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in is one line with (i) and (ii)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in says SDK client and your own agent

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in uses parallel verbs

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in says access rather than call

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in drops the verb in (i)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in reads as one plain sentence pair

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the client guide is headed Use PageIndex through the SDK client

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in lists its two ways as bullets

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in returns to one (i)/(ii) line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Ready to Try It links cover only the noun

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage section is headed Usage Guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the summary heading reads tl;dr

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items use bold summaries so they sit close together

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: expanded items get a line of air under their summary

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items try h4 summaries again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items settle on bold summaries with a spacer

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a line of air between neighbouring collapsed items

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Step 2 anchor lives inside its summary so item gaps match

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the two guide headings rely on their generated anchors

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide opens with a plainer line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in drops workflow

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in names its three steps

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in, polished

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in points below

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a line of air before the integration heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: no spacer before the integration heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the two usage guides are labelled (a) and (b)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench closing line names the evaluation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench closing line links only the results

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench claim links straight to its benchmark subsection

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the indexing-cost line names the model in prose

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench subsection is headed Leading on FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench heading reads Leading accuracy on FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Updates lists the PageIndex File System again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the File System update entry links in plain weight

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the header link row gains Blog

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the summary heading reads TL;DR

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Model Recommendations points at the Usage Guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Model Recommendations names the SDK client usage guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the usage-guide pointer promises more than models

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the usage-guide pointer, one word shorter

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in says each example covers one framework

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in says a different framework

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: each framework fold shows the one-call and explicit forms end to end

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the streaming example uses a literal and the quickstart drops trailing spaces

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry says where the structure comes from

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry, tighter

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry says built by an LLM

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry, Ray's wording

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry keeps extracted heuristically

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Updates dates read [Aug '26]

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k
2026-08-28 00:14:31 +08:00
Ray d4ce6ee65f docs: the run_messages comment keeps the principle, drops the version numbers
The rationale to preserve is that execution policy belongs to the
vendor and only the runner's history is ground truth; the exact 1.0/1.1
switch point lives in the previous commit's message.
2026-08-27 03:37:47 +08:00
Ray 070293bdca test: the messages max_tokens test follows anthropic 1.1's stop policy
anthropic 1.1.0 replaced the runner's refusal special case with an
explicit stop-reason table: every non-tool_use stop is terminal and its
tool_use blocks are never executed (1.0 executed a max_tokens turn's
complete blocks). The envelope already keys on the runner's history, so
the product adapts by design; only the test had the 1.0 behavior baked
in. It now asserts the version-appropriate shape on both sides, and the
run_messages comment describing the old behavior is reworded to name
the policy split.

Verified: 434 green on anthropic 1.1.0 (CI's failing config) and on
0.120.2 (the pre-1.1 branch); no-frameworks collection stays clean.
2026-08-27 03:37:47 +08:00
Ray f47f27b7b3 docs: README restructure with quickstart, benchmarks, and a collapsed usage guide (#431)
Quickstart on the v0.2.11 index=/chat= spellings, a citations example, indexing-time benchmarks with new charts, the own-agent integration in its own collapsed block, a rewritten PageIndex Cloud section, the reasoning positioning restored, and prose em dashes replaced with plainer punctuation. Earlier snapshots of this branch landed via #418/#419; this squash carries the tail.
2026-08-27 03:28:17 +08:00
Ray 31f01910cd docs: API-key URLs follow the dashboard move to developer.pageindex.ai
dash.pageindex.ai/api-keys now 307-redirects to
developer.pageindex.ai/api-keys, and the README (PR #431) already
standardized on the developer domain — the client docstring and the
CloudClient keyless error message were the last SDK references to the
old host.
2026-08-27 02:54:46 +08:00
Ray 57252e2686 fix: the LiteLLM lanes hide litellm's bridge usage warning (#430)
Streaming chat() on an OpenAI gpt-5.4+ model with function tools makes
litellm re-route chat.completions through the Responses API. When the
stream ends, litellm's logging stores a chat-shaped usage dict inside a
ResponseAPIUsage field and model_dump()s it, so pydantic prints
"Expected `ResponseAPIUsage` - serialized value may not be as expected"
once per streamed turn. litellm does this on purpose
(litellm_logging._get_assembled_streaming_response, 1.97 and 1.98) and
the answer is unaffected.

Both seams that hand a LiteLLM-routed model to openai-agents — the chat
lanes' model builder and openai_agent_config()'s litellm/ lane —
register one warnings filter matching exactly that message; every other
warning still surfaces. A process-wide filter is the only placement
that works: litellm emits the warning from its async success handler
on the worker thread, out of catch_warnings' reach.

Claude-Session: https://claude.ai/code/session_01APJbbp3jpRxJhvchcU2Aed
2026-08-26 20:03:34 +08:00
Ray 174f95f35b fix: the slots accept any Mapping at runtime; eight docstrings stop calling cloud tool scoping server-side (#429)
fix: the slots accept any Mapping at runtime, as their annotation admits; eight docstrings stop calling cloud tool scoping server-side

4e9c56c widened index=/chat= to Mapping[str, Any] so the exported
TypedDicts pass a checker, but _resolve_index_slot/_resolve_chat_slot
still dispatched on isinstance(..., dict): a MappingProxyType or ChainMap
was pyright-clean and raised "must be a string or a dict" at construction.
The resolvers now narrow on Mapping — the comprehension already copies,
so a read-only proxy proves the caller's mapping is never mutated.

4e9c56c corrected three of eleven "scoping is server-side" sites; the
remaining eight said the same untrue thing about doc_id on cloud (its
tools carry no allowlist — targeting is prompt-level, as the runtime
error already explains). Deleted rather than reworded.

local_chat.py's module docstring predates own-model chat over the cloud
bridge; storage_path's prose now names the PathLike 5e2dc9b typed.

Claude-Session: https://claude.ai/code/session_01VQ6mruXZBgw9Hjii8KPbQP
2026-08-26 17:00:41 +08:00
Ray b9a9a3b5aa fix: the .env search ends at the cwd tree; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike (#428)
* fix: .env stays unset when the cwd tree has none; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike

find_dotenv(usecwd=True) returns '' when nothing is reachable from the
cwd, and `or None` turned that into load_dotenv's own upward walk from
utils.py — the install-dir leak the cwd search was added to replace. A
pip-installed SDK could load another project's .env from above
site-packages, silently.

_local_chat treats a blank chat_model as "managed chat", which a client
without an api_key does not have: chat_completions() then reached for
LocalAPI.chat_completions and raised a bare AttributeError. The managed
branch now refuses as a PageIndexAPIError naming chat_model.

py.typed made the annotations authoritative while storage_path was typed
str; _ARG_TYPES accepts os.PathLike, so Path(...) ran fine and failed the
user's type check. Both signatures and LocalIndexConfig now say so.

Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb

* fix: the exported config shapes pass into index=/chat=; a comment and two docstrings stop overclaiming

The slots were annotated dict[str, Any]. A TypedDict is consistent with
Mapping[str, object], never with dict (PEP 589: a dict-typed receiver
could write arbitrary keys through it), so the four shapes types.py
exports — and py.typed advertises to installed callers' checkers —
could not be passed to the one place they describe. pyright on a probe
that does exactly that: 9 errors before, 0 after. The constructor only
reads the slot (items(), then a fresh conf dict), so Mapping is the
honest bound; a plain dict is a Mapping, and TypedDict instances are
plain dicts at runtime, so nothing moves at runtime.

The _ARG_TYPES comment said "every value" is shape-checked; api_key is
not in the table (its empty check is separate, its type check stays
unchecked by ruling), so the comment now speaks for the table only.

_local_doc_scope and _require_local_scope still explained the cloud
drop as "scoping is server-side" — true of the managed chat, which
never reaches either function. What reaches them on a cloud client is
own-model chat and the config helpers, whose cloud tools take no
allowlist: targeting there is prompt-level only, as the error message
between them already said.

434 passed; pyright on pageindex/ unchanged at 235 (0 in the touched
files, before and after).

Claude-Session: https://claude.ai/code/session_01TxG8u8x29XRnK4yscZVCch

* test: the install-dir .env test is named for what it asserts

Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb
2026-08-26 15:08:34 +08:00
Ray 920db2b1b1 feat: the client grows two sides — documents and chat each pick their home (#424)
* feat: the client grows two sides — documents and chat each pick their home

One client, two independent switches: api_key decides where documents
live (the PageIndex cloud, or the local store); a configured chat model
decides who answers (your own model in your process, or the managed
cloud chat). Their free combination opens the bridge — cloud documents,
your model — and the fourth cell stays unspellable.

- index=/chat= slots: string shorthand or grouped dict, 1:1 with the
  flat arguments; one spelling per side, sides mix freely
- optional "type" everywhere (top-level and in either dict): always
  omittable, checked against the content, meaningful alone —
  type="cloud" is a keyless cloud spelling
- PAGEINDEX_API_KEY is read only when the code explicitly says cloud
  (PageIndexCloudClient(), type="cloud", "pageindex-cloud",
  {"type": "cloud"}); a bare PageIndexClient() stays local
- bare mode words ("cloud", "local", …) are reserved: they error
  with the real spellings instead of silently parsing as model names
- bridge chat runs the in-process agent over the live cloud MCP tools
  and instructions; doc_id targets at the prompt level; citations stay
  managed-only; an auth-shaped backend failure explains whose
  credentials run the model
- typed shapes (IndexConfig, ChatConfig) ship as optional annotations

Every previously working program is byte-for-byte unchanged: the only
behavioral deltas are error paths — reworded guidance, and the
api_key+chat_model combination graduating from an error into the
bridge.

* fix: the constructor refuses empty and mistyped values on every spelling

- .env keys reach all four keyless-cloud spellings: utils' import-time
  load_dotenv now runs before every PAGEINDEX_API_KEY read
- an empty chat-side value ("", {}) errors instead of silently selecting
  own-model chat on the default model; None-valued slot keys mean absent,
  exactly like the flat arguments
- _local_chat derives from chat_model, so a post-construction assignment
  switches the whole client, never half of it
- model= beside a slot gets the split guidance (index_model=/chat_model=)
  instead of "two spellings of the same thing"
- the messages door wraps provider failures through _model_backend_error,
  and 401s count as auth-shaped even without "api key" in the text
- keyless-cloud hints name the spelling that actually combines; slot
  strings are stripped; wrong-typed values raise PageIndexAPIError
- retrieve_model/chat_backend docs drop the stale "Local mode only";
  the local-scope refusal no longer claims bridge tools are server-scoped

* fix: type= cross-checks the index slot; the cloud pinned class frees its chat side

- type= beside index= now does what the docstring promises: agreement
  passes, disagreement errors, and a mistyped value reports the
  vocabulary error instead of a spelling collision
- PageIndexCloudClient grows the chat-side arguments (chat=, chat_model,
  retrieve_model, chat_backend), so "pin the index side" is literally
  true and the chat surfaces' construct-with-chat_model guidance is
  followable on it
- the four chat doors' doc_id entries carry the enforcement split the
  config helpers already state (local: tool-layer allowlist; cloud:
  prompt-level / server-side)
- types.py stops claiming slot keys share the flat names — the side
  prefix is factored out, index={"model"} is index_model=

* docs: bridge-reachable wording — dependency errors say own-model chat, hints name a chat= model

- the three framework-missing errors said "in local mode", which is
  wrong on a bridge client (cloud documents + own model) — they now
  explain the dependency the way the surfaces do: your own chat model
- the construct-with guidance reads "(or a chat= model)": a bare
  chat="pageindex-cloud" is also chat= but selects the managed side
- the mechanical Local-only → Own-model-chat-only substitution left
  orphan fragments and two overlong lines; those paragraphs re-flowed

* fix: managed chat reads None; the bridge stops paying per-turn tool lists

- a managed-chat cloud client stores chat_model/chat_backend as None, so
  the documented attribute reads instead of raising AttributeError;
  _local_chat derives from "is a chat model configured"
- McpBridge caches tools/list per session — every chat turn rebuilds the
  tool set, and the round trip was pure latency; the 404 session-expiry
  reset drops the cache with the session
- run_messages builds tools before the transport: on a bridge client
  that build is network I/O, and a failure there stranded a per-call
  anthropic client ahead of the try/finally

* refactor: the side declaration is spelled mode=, not type=

"type" is Python's own word — a builtin, and "data type" beside the
TypedDict shapes; "mode" is what the SDK already calls the two sides
("local mode", "cloud mode"). Same grammar everywhere the declaration
appears: the top-level argument, the index dict, the chat dict, the
typed shapes. The rename also frees the builtin inside the constructor,
so the shape-check error names the offending class through type() again.
"type" in a slot dict is now an ordinary unknown key.

* fix: the reserved-word errors stop calling "cloud" not a mode word

With the declaration key spelled mode=, 'index="cloud" is not a mode
word' contradicted its own remedy, index={"mode": "cloud"} — "cloud" is
exactly a mode value. The four bare strings are reserved words; the
message now says so.

* fix: a blank tools/list is not cached; the auth note's managed exit is chat-lane only

- McpBridge.list_tools caches only a non-empty list — a transient blank
  (a deploy blip, a gate misconfiguration) would otherwise run every later
  turn with zero tools while the instructions still name them, and only a
  404 session reset could clear it
- the 401 architecture note appends "drop the chat model configuration"
  only on the chat lane: responses() and messages() refuse a client
  without an own model, so on those lanes the exit sent the caller in a
  circle
- CloudIndexConfig says api_key is omittable only while mode: "cloud"
  stays — index={} refuses as an empty dict rather than reading the env

* fix: the bridge fetches tools/list per call again; .env resolves from the cwd

- McpBridge.list_tools no longer caches: the tool set is built once per
  SDK call (Agent(tools=...) ahead of Runner.run; build_anthropic_tools
  ahead of tool_runner), not per model turn, so the cache saved one round
  trip per later call while a mid-pagination 404 replayed a dead cursor
  into a duplicated (and cached) list, and the list went out by reference
  across a lock dropped between miss and store
- utils.load_dotenv searches upward from the cwd: a bare load_dotenv()
  walked up from utils.py, which is site-packages for an installed SDK,
  so the four keyless-cloud spellings never saw a project-root .env; the
  package-relative walk stays as the fallback
- the emptiness guard strips strings: chat_model=" " selected own-model
  chat, the silent flip the guard's own comment rules out
- the _local_chat comment stops advertising post-construction assignment
  as a full mode switch

* fix: the pinned classes take index=/chat=; "cloud"/"local" are mode words; a blank chat_model stays managed

- PageIndexLocalClient takes index= and chat=, PageIndexCloudClient takes
  index= — the grouped spelling of the flat vocabulary each already took;
  their refusals name the class and an exit that class can take, and the
  mode cross-check runs before any environment read
- "cloud" and "local" are accepted wherever "pageindex-cloud" was (index=,
  chat=, mode=, {"mode": ...}), case- and whitespace-insensitive; "hosted"
  and "managed" still refuse, pointing at the real word
- every spelling strips its strings, and the slot spellings' type/empty
  errors name the slot key (index["model"]), not the flat argument
- _local_chat treats a blank chat_model as managed: the constructor
  refuses "", so assignment agrees instead of opening the bridge on a
  nameless model; openai_agent_config carries no model then either
- an empty MCP tools/list raises like empty instructions does — a
  zero-tool agent would answer from the model's own knowledge silently
- enable_citations names the real gate (managed vs own chat), not
  "cloud-only", on a cloud own-model client
- pageindex/py.typed: the exported config TypedDicts reach installed
  type-checked callers

* test: the two framework-door tests skip without openai-agents

as_openai_tools() and openai_agent_config() need the agents package, which
the "without frameworks" CI legs do not install — the same importorskip
every other test on those doors already carries.
2026-08-26 04:13:33 +08:00
Ray 416e304f51 perf: expand schedules dependency-exact at thirty-two concurrent proposals (#422)
* perf: expand proposes a wave of nodes concurrently

The expand loop awaited one propose_children at a time — 20-30 nodes at
~3s each put 1-3 minutes of pure round-trip latency on every default
local submit. Nodes waiting in a wave are all frontier leaves whose
decisions cannot affect each other, so the model half now runs
concurrently (EXPAND_CONCURRENCY = 8) while the apply half stays serial
in wave order: decisions, log entries, and child ids land exactly as
before, and children attach into the next wave. A fatal classification
still aborts the run right after the wave's gather.

Benchmarked on real PDFs with a fixed-latency fake model: 408 pages
21.1s -> 3.0s, 758 pages 28.2s -> 3.5s (7-8x); final trees byte-identical
to the serial pass on both. The cap stays low on purpose: expand treats
an exhausted retry ladder as fatal, and a wide burst on a rate-limited
account would trip exactly that — 8 already collapses minutes to seconds.

* perf: expand schedules dependency-exact instead of in waves

A child's only prerequisite is its own parent's apply, so each kept
node gathers its children directly rather than waiting for its whole
generation to finish. Same recursive shape as summarize_tree; the
semaphore still caps in-flight proposals at 8; trees are unchanged.

* perf: expand admits thirty-two concurrent proposals

Cap sweeps on six real documents put the speed plateau at 32: the
ready frontier tops out at 21-28 nodes on few-hundred-page PDFs, so
64 buys nothing while doubling the burst. Live runs at 32 cut the
expand phase 24-30% on the two documents wide enough to feel it,
with zero ladder retries anywhere - and summaries already burst
twice as wide through the same ladder.
2026-08-22 14:49:01 +08:00
Ray 8289729aff fix: indexing failures surface loud, chat envelopes append verbatim (#421)
Indexing: dead credentials or a missing model fail the run instead of
storing a document with blank summaries; a 400 (context_length_exceeded)
skips the retry ladder — the prompt will not shrink — and stays a
per-prompt failure the run absorbs; all-empty model replies can no longer
store a retrieval-ready document; the one-sentence doc description
absorbs its own context overflow instead of discarding a fully indexed
document; the heading-less flash refusal points at mode='standard'.

Chat: messages() output is append-verbatim clean — unset response-only
defaults are dropped (no "caller": null the request schema rejects);
Claude cache marks follow the wire routing; model_settings and name are
openai_agent_config parameters; one Anthropic client per backend; lifted
thinking defaults are clamped to the model's output ceiling from
LiteLLM's capability map.

Store and inputs: lone surrogates are scrubbed from page text and the
stored basename, so the returned name is byte-for-byte the stored name
and the rename warning fires; NaN/Infinity metadata is rejected at the
gate; every cloud error now carries its HTTP status.

CLI: the flash lane resolves the summary model through ConfigLoader like
the standard and markdown lanes; an empty flash structure errors like the
SDK instead of writing "structure": [] with exit 0; --summary-model
reaches the markdown lane; the SDK page-spec surface keeps 0.2.10's
whitespace tolerance while the tool layer stays strict.

pypdfium2 stays on the 5.x line for every install; the 4.x code paths are
tested compatibility insurance with their own CI leg; process-pool
construction failure falls back to the sequential parse; a py3.10 GC
flake in text extraction is fixed.


Port of feat/local-chat 0667e3b..1993740 (28 commits); README and assets untouched.
2026-08-21 21:35:45 +08:00
Ray 418d1549bb docs: weave the original README's highlights back into the SDK rewrite 2026-08-19 23:03:07 +08:00
Ray b628190529 ci: dev tags publish to PyPI only — skip the GitHub Release (#414)
Dev builds need an explicit ==pin to install, so their GitHub Releases
carry no install value and double the feed next to the same-day stable.
2026-08-19 22:16:39 +08:00
Ray ba0ef02d78 fix: Aug 18-19 review hardening — five max-effort rounds across tools, chat, and flash (#413)
Squash-port of feat/local-chat's post-dev5 wave (03ffab3..49a24e1): five
review rounds of reproduced-then-fixed findings. Highlights: standalone
agent_instructions back on the strict shadow check (the *_agent_config
bundles keep the relaxed in-set check they can prove); caller-owned
http_client survives the per-call backend closes; get_tree keeps
key_items under the flash merge default; OpenAI-protocol classification
follows litellm's own routing (azure/openrouter/deepseek/groq/xai);
thinking-aware max_tokens default shared by messages() and
anthropic_runner_config(thinking=); 401/403 re-raise instead of
retry-coaching envelopes, bridge errors carry status_code; cloud
discovery + instructions ride the ?tools=read endpoint matching the
tool gate (live-verified); the chat lane's litellm model resolution
deduplicated into utils._litellm_model; explicit optimize= wins over
the deprecated optimize_expand (DeprecationWarning added); mcp 2.0
compatibility; spawn-worker and doc-scope guard fixes; the suite is
.env-independent and pins the LitellmModel._fetch_response seam.

Per-finding rationale in the ported commit messages on feat/local-chat.
2026-08-19 21:58:37 +08:00
Ray ae2a5b49b5 feat: backend connection overrides; extra_headers on every local door (#411)
* feat: backend connection overrides; extra_headers on every local door

Two clients, two configs — the gap this closes. index_backend /
chat_backend on the constructor (and per-call backend on the chat
doors, mirroring per-call model) carry connection params in each
lane's own vocabulary, verbatim: the indexing gateways take the dict
as LiteLLM call kwargs (a contextvar scopes it per operation, and the
env-var key pre-check yields to it), the chat lane lifts api_key /
base_url into LitellmModel's two pinned constructor slots and rides
the rest as call kwargs, responses() and messages() hand the dict to
their SDK client constructors. Per-call keys win over the client's;
the openai-SDK fast path normalizes LiteLLM's api_base spelling.
Config bundles deliberately don't carry it — you run those in your
own environment (docstring says so).

extra_headers lands on all three protocol doors, each engine merging
caller headers verbatim (anthropic-beta wire-proven on messages).
Wire-probed exception, documented on the chat door: LiteLLM's
anthropic adapter owns the anthropic-beta header and drops the
caller's value — Anthropic beta flags belong on messages(). Cloud
rejects the new knobs like the other local-only params, and the
docstrings now say credentials belong in backend, never extra_body
(on the bare lane they would leak into the JSON body without
touching auth — wire-probed).

* docs: constructor Args list gains index_backend / chat_backend

The class docstring documents every constructor argument; the backend
wave added two without entries. Same phrasing as the per-call docs:
index lane is LiteLLM vocabulary verbatim, chat_backend reaches
whichever door runs (api_key/base_url portable across all three).

* chore: three audit leftovers

The classic pipeline's ThreadPoolExecutor import (unused since the
0.2.9 merge) goes. The bare-model missing-key error now also names
the backend={'api_key': ...} route, which satisfies the same check.
messages() joins the uniform: a bad backend dict wraps as
PageIndexAPIError like the responses door, instead of leaking the
SDK's raw constructor error.

* chore: narrow the anthropic wrap to TypeError; key advice names chat_backend

Anthropic's constructor raises only TypeError in our supported range
(unknown kwargs, conflicting credentials) — a missing key defers to
request time, so catching AnthropicError there guarded an impossible
case. And chat() has no backend parameter, so the missing-key advice
now names chat_backend alongside the per-call route.
2026-08-17 17:34:19 +08:00
Ray 08ea1d975c fix: py3.10 litellm type repair; chat() gains reasoning_effort (#410)
* fix: rebuild litellm's Message/Delta types on Python 3.10

litellm 1.97.0 ships Message and Delta annotations whose nested forward
refs (ChatCompletionReasoningSummaryTextBlock et al) do not resolve on
3.10, so every completion() dies constructing its response object —
non-stream and stream alike (upstream BerriAI/litellm#36384, open, no
patch release; 1.96.2 is clean, so the floor raise surfaced it, and
pydantic 2.12/2.13 both reproduce). The repair rebuilds the two models
once with their defining modules' namespaces at our three completion
gateways; version-gated to <3.11 and best-effort, so it is a no-op on
healthy interpreters and future fixed litellm releases. Verified on a
3.10 venv: the previously failing anthropic wire test and the full
suite pass (250 green, matching CI's matrix leg).

* feat: chat() takes reasoning_effort — the front door's one thinking knob

Ruled in as a business-level control alongside model: who answers, and
how hard it thinks. Same name, values, and verbatim semantics as
chat_completions underneath (LiteLLM's cross-provider tier string);
unset sends nothing so each backend's own default behavior applies.
Sampling and wire-level knobs deliberately stay off the front door.
2026-08-17 01:34:22 +08:00
Ray bc1c1740ff feat: model knobs, LiteLLM-verbatim chat lane, per-door passthrough params (#409)
* docs: drop the demo's install step — openai-agents ships with the SDK now

* feat: the chat lane routes every model through LiteLLM — bare names included

The direct-OpenAI special case existed to dodge LiteLLM's import cost,
and it made OpenAI's own Responses-first models fail on the front door:
gpt-5.6-sol 400s on chatcmpl+tools while reasoning is on (server-side
policy — wire-captured with no reasoning_effort in our request).
LiteLLM 1.97 translates such calls onto /v1/responses; 1.84 does not,
so the sol-class 400 now carries its two exits (upgrade litellm /
responses()).

Routing after the flip: chat protocol — bare names are OpenAI-compatible
shorthand (wire form openai/<name>; OPENAI_API_KEY / OPENAI_BASE_URL
still select the backend, and the missing key stays a build-time
failure), litellm/ strips, openai/ opts out to the OpenAI SDK directly;
responses protocol unchanged (OpenAI-SDK native, LiteLLM refused).

The import cost is handled instead of dodged: local clients preload
litellm on a background thread (first call then perceives 0.0s), and
pageindex sets LITELLM_LOCAL_MODEL_COST_MAP=True via setdefault —
LiteLLM's import otherwise blocks on a network fetch of its price map
(fresh venv: 5.6s -> 1.3s; offline it hangs to the timeout).

Also restores prompt_cache_key delivery, found dead during the flip's
gating verification: openai-agents 0.20 no longer derives it from
RunConfig.group_id, so both lanes sent nothing. ModelSettings.extra_body
is the one channel all three model classes put on the wire (the bare
kwarg is dropped by LiteLLM; extra_args[extra_body] collides with the
responses model's own parameter — both wire-verified), and it is scoped
to OpenAI destinations: LiteLLM plants extra_body as a literal field in
other providers' bodies, and Anthropic rejects unknown fields — the
anthropic wire test now pins the absence.

Verified before landing: mock-server matrix (OPENAI_BASE_URL + bare
name works through LiteLLM; gpt-named self-hosted models are NOT
bridged off a custom base_url; prompt_cache_key on the wire in every
OpenAI lane with distinct per-conversation keys; anthropic body clean)
and live (sol answers through chat(), gpt-5.4 unchanged, responses()
bare unchanged with the key on its wire).

* refactor: no prefix-triggered direct lane — chat model names are LiteLLM's, verbatim

Ray's ruling on the flip's remaining carve-out: a routing decision must
never hide in a model-name prefix. openai/ now means what LiteLLM says
it means (its openai provider), like every other name on the chat lane —
the grammar is LiteLLM's with zero exceptions.

The two defenses for keeping a direct carve-out had no concrete victim:
debugging isolation (litellm is unavoidable in indexing anyway, and
responses() IS the OpenAI-SDK-native door), and endpoint determinism
(litellm sends chatcmpl for openai-provider models except the gpt-5
bridge, which never fires against a custom base_url — wire-verified).
If a direct escape is ever needed, it will be a declared parameter,
never name grammar.

openai/-prefixed names keep the build-time OPENAI_API_KEY check for
parity with bare names; responses() is untouched (bare and openai/
still drive the OpenAI SDK — LiteLLM cannot speak that protocol).

* fix: third-audit findings — extra_body naming drift, litellm floor hint

The prompt_cache_key delivery channel went through three iterations and
settled on ModelSettings.extra_body; the _conversation_cache_key
docstring still named extra_args from the middle iteration.

The litellm install hint said >=1.30, below both our own pyproject floor
(>=1.84.0) and the floor openai-agents' litellm extra declares (>=1.83).
A user in a broken environment following it would land on a version the
package itself rules out. The hint now matches the declared floor.

* test: pin the bundle door's Agents-SDK model grammar; rename the cache-key test to extra_body

The config bundle hands its model string to the Agents SDK's own
MultiProvider grammar, which refuses unknown prefixes (probe on 0.20:
'anthropic/x' -> UserError: Unknown prefix). _normalize_retrieve_model's
litellm/ spelling is what keeps that door working — a link the existing
self-referential assert (config["model"] == client.retrieve_model)
could not catch. Pinned with a provider-slashed name.

Also renames the cache-key delivery test to its real channel,
extra_body — the extra_args name survived from the superseded delivery
attempt.

* refactor: name the retrieve_model helper for its reason — the Agents SDK's grammar

_normalize_retrieve_model said what it does, not why. The litellm/
spelling exists because the Agents SDK resolves raw model strings with
its own prefix grammar and refuses unknown prefixes — the name now
points at that constraint.

* test: the bundle-grammar test skips without openai-agents, like its file's siblings

Every agents-dependent test in this file importorskips; without the
guard this one errors where the others skip.

* feat: index_model + chat_model — two-knob model surface with full legacy fallback

The documented surface becomes two role knobs: index_model builds the
index, chat_model answers on the chat surfaces. model turns into the
set-both umbrella (its 0.2.8 indexing semantics are a strict subset, so
old configs run unchanged); summary_model and retrieve_model stay
accepted as legacy role names.

Resolution lives in ConfigLoader.load(), the one seam every consumer
already passes through (client, CLI standard/md paths, flash's
summary fallback, tree_optimize's default_model): new names win over
old, specific over general, model sets every role, and code constants
close each chain. The packaged yaml no longer ships model keys — key
presence is what separates a user's explicit choice from a built-in
default, and _validate_keys accepts the five model names explicitly.

Consequences: with no config at all, classic-mode structure extraction
now uses DEFAULT_INDEX_MODEL (gpt-5.6-luna) instead of the yaml's old
gpt-4o-2024-11-20 line (ratified; flash-default users see no change).
client.retrieve_model becomes a read-only alias for client.chat_model.

The resolution matrix test pins one row per released generation:
0.2.8 (model), 0.3.0.dev (model+retrieve_model), 0.2.10.dev (all three
legacy names), the new pair, umbrella-only, and mixed.

* feat: --index-model on the CLI; README flag docs follow

The CLI leads with --index-model; --model stays as its legacy synonym
(the CLI only indexes, so the umbrella and the index role coincide).
The flash branch's summary fallback gains the index position, and the
standard branch now forwards --summary-model, which it had silently
ignored — the flag's help always claimed it worked there. The md
branch's unfiltered model=None no longer clobbers the default: the
resolver treats None as unset.

* feat: default chat model becomes gpt-5.6-sol

Ray's pick for the out-of-box QA default; indexing stays on luna. sol
runs tools on the Responses lane — current litellm bridges chat()
there automatically; older litellm gets the guided 400 naming both
exits.

* chore: trim the model-keys comment to the constraint

The per-generation history lives in 7244ee4's message.

* feat: per-door reasoning passthrough — reasoning_effort / reasoning / thinking

Each chat door gains its own protocol's native thinking control,
forwarded verbatim with no invented vocabulary and no default of ours:
chat_completions(reasoning_effort=...), responses(reasoning={...}),
messages(thinking={...}). Unset sends nothing, so backend defaults
(sol: medium, adaptive) are untouched. chat() stays answer-only.

Delivery channels, each verified: the chat door rides
extra_args["reasoning_effort"] — LiteLLM's own top-level kwarg on
every supported openai-agents version, admitting non-enum values
("none"); newer openai-agents promotes it to the top-level argument
and pops the duplicate. Wire-captured on a mock backend
(/v1/chat/completions body carries it) and coexists with the Claude
cache marker in one dict. The responses door rides
ModelSettings.reasoning — coerced to the typed openai Reasoning object
and forwarded verbatim by the Responses model; the envelope echoes the
caller's dict. The messages door joins the existing anthropic
passthrough dict, asserted through the real tool runner.

LiteLLM semantics observed and accepted as-is: unknown models refuse
the param loudly with LiteLLM's own remedies, and gpt-5.4+ names with
an explicit effort route to /v1/responses even against a custom
api_base (its documented pre-existing arm). The sol-class 400 guidance
now names the third exit — an explicit effort routes on older litellm
releases too.

Cloud chat_completions rejects the new parameter like model/max_turns;
responses()/messages() are local-only already.

* feat: extra_body escape hatch on the three protocol doors

The industry-standard per-request extension channel (openai/anthropic
SDK trio): a dict merged verbatim into the backend request, last, so
caller keys win over SDK-set ones. Routing per door: OpenAI-compatible
destinations get a true body merge (ModelSettings.extra_body / the
anthropic SDK's native extra_body); LiteLLM-routed providers take the
keys as LiteLLM's own top-level kwargs instead, since LiteLLM plants
extra_body as literal fields other providers reject. Cloud mode rejects
it like the other local-only knobs. chat() stays answer-only.

* chore: litellm floor 1.84 -> 1.97.0

1.97.0 is where the unset-effort chatcmpl->responses bridge landed
(responses_api_bridge_check's on_constraint_enforcing_endpoint arm,
A/B-verified against 1.96.2), so sol-class models work through the
chat lane out of the box instead of 400ing until a manual upgrade.
Three spots move together: the pyproject floor, the requirements.txt
CI pin, and the install hint.

* feat: named top_p/max_tokens on the chat door, max_output_tokens on responses

extra_body could not carry these: openai-agents' LitellmModel passes
every ModelSettings sampling field as an explicit keyword and unpacks
extra_args into the same call, so the common knobs collided with a bare
TypeError on the LiteLLM lane (reproduced against a stub — Python call
semantics, callee-independent). Named params ride ModelSettings fields,
the one channel clean on every lane; responses() uses the protocol's
own name (openai_responses maps ModelSettings.max_tokens to
max_output_tokens on the wire) and the envelope now echoes the real
value instead of a constant None. Both caps bound each backend call in
the agent loop, not the whole run — documented. Cloud rejects them like
the other local-only knobs; the long tail (frequency_penalty etc.)
stays extra_body-blocked-loudly on that lane by choice.

* chore: the missing-key error names chat_model as the other exit

The default chat_model is what put keyless users on the OpenAI lane,
so the error now points at the knob that picks a different backend.
2026-08-17 00:42:17 +08:00
Ray 08a912c69c feat: openai-agents becomes a base dependency (#406)
feat: openai-agents becomes a base dependency — the chat engine ships with the SDK

chat() is the SDK's front door, and its engine lived behind a
vendor-named extra: pip install pageindex could index a document but
failed on the first chat call, and chatting with Claude required
installing '[openai]'. Measured before moving: the base tree already
carries litellm (75 MB) + openai (13 MB), openai-agents adds ~15 MB
(agents 8.1 + mcp 1.7 + griffe 1.4 + small pure-python deps), and
current litellm's openai range (>=2.20,<3) intersects cleanly with
openai-agents' (>=2.45,<3).

The [openai] extra stays declared but empty, so existing
pip install 'pageindex[openai]' commands keep resolving. Error
messages and docstrings drop the extra; requirements.txt gains the
dependency, so CI now runs the openai-agents test lane instead of
skipping it. Extras now mean exactly one thing: a vendor's own SDK
surface ([anthropic] for messages()/tool runner, [claude] for the
Claude Agent SDK).
2026-08-15 02:23:32 +08:00
Ray ec4851059f feat: chat() front door and Claude cache marking on LiteLLM routes (#405)
* feat: chat() — the answer-out front door over chat_completions

* feat: cache-mark the managed prefix on anthropic-routed LiteLLM models

* refactor: cache predicate asks litellm's own provider resolution

* feat: extend cache marking to Claude on Bedrock and Vertex — both live-verified

* fix: point anthropic-extra users at messages(); guard the two silent vendor chains

* docs: the max-tokens table is a closed set — litellm's map prunes EOL'd entries, so it cannot replace it

* chore: trim the max-tokens docstring to the contract

* docs: state the local text-only history contract — cloud forwards tool turns, local rejects; extra fields drop
2026-08-15 01:39:55 +08:00
Ray 7c6e3c1c8f feat: Flash with full optimization becomes the default local indexing mode (#404)
* fix: break the phantom exception chain in _run_sync

Move asyncio.run(coro) out of the except RuntimeError block so real
errors no longer carry a bogus "no running event loop" context in
their traceback.

* feat: Flash with full optimization becomes the default local indexing mode

Every entrance now defaults to Flash with the full optimize pass
(deterministic merge, then LLM expand), replacing the standard LLM-built
tree as the default:

- submit_document(): mode=None now means "flash"; pass mode="standard"
  for the LLM-built tree. _index_flash runs optimize="full" with the
  expand model = summary_model, and fails fast with the missing key
  name(s) via litellm.validate_environment before any work.
- page_index_flash(): optimize takes "full" (default) / "merge" / False;
  True is accepted as "full" for compatibility, unknown values raise
  instead of silently degrading to merge-only. optimize_expand stays
  honored for legacy callers.
- CLI: --mode {flash,standard} replaces --flash (kept as a hidden
  compatibility alias that forces flash). --optimize defaults to full in
  flash mode with an `off` choice; explicitly passing it outside flash
  still errors. Standard-only tuning flags (--toc-check-pages,
  --max-*-per-node, --if-add-*) now error in flash mode instead of being
  silently ignored, mirroring the existing flash-only flag errors. The
  key pre-check runs only when an LLM will actually be called, so
  --no-summary --optimize off|merge works keyless. Output drops the
  _structure_flash suffix — always <name>_structure.json.

On the Disney earnings PDF the optimized default is also faster than
unoptimized flash (fewer nodes to summarize) and fixes hierarchy
mistakes; both modes emit identical schemas end to end.

Docs updated to match (mode flag, defaults, LLM usage honesty); tests
pin the new defaults: stored mode == "flash", optimize passthrough, and
the unknown-optimize rejection.
2026-08-13 18:28:11 +08:00
Ray ec342c4e28 fix: post-merge review fixes for v0.2.10
- Conformant responses() envelope: official output (model items only)
  + items (full transcript for round-trip), usage details aggregated
  across turns; verified against real OpenAI API
- Python floor corrected to >=3.10 (litellm stable requires it)
- Three stale anthropic>=0.84.0 hints updated to 0.108.0
- Image-stub behavior disclosed on tool-builder docstrings
- Non-essential comments trimmed
2026-08-13 14:57:08 +08:00
Ray 4e41acdc68 feat: agent tools and local chat for the PageIndex SDK (v0.2.10) (#396)
* feat: agent tools — the cloud MCP tool contract on the client

Four new client methods make PageIndex documents available to agent
frameworks, in both modes, with the mode decided solely by the client
constructor:

- agent_tools(): plain functions (browse_documents, get_document,
  get_document_structure, get_page_content) matching the PageIndex cloud
  MCP server's tools/list — same names, schemas, descriptions, and JSON
  response envelopes — so agent prompts port unchanged between the cloud
  MCP connection and these in-process tools. Tools never raise; errors
  come back in the same envelope. remove_document ships behind
  include_management=False.
- as_openai_tools(): the same tools wrapped for the OpenAI Agents SDK.
- as_claude_mcp(): one mcp_servers entry for the Claude Agent SDK —
  cloud clients get the remote MCP config (the framework connects to
  api.pageindex.ai/mcp and discovers the full cloud tool set), local
  clients get an in-process SDK MCP server.
- agent_instructions(doc_id=None): orchestration guidance for the
  agent's system prompt; doc_id (same shape as chat_completions) appends
  the target documents.

submit_document() gains wait=True: poll get_document status until
completed, raise on failed or after 30 minutes — the manual polling loop
every cloud caller writes today spins forever on a failed document.

Neither framework becomes a dependency: imports happen at call time with
actionable errors, and pageindex[openai] / pageindex[claude] extras are
floor-only pins. tests/data/cloud_mcp_contract.json freezes the tool
contract; a parity test guards against drift. 36 new tests (95 total),
plus a live OpenAI Agents SDK run over a seeded local store verifying
the structure-first navigation flow end to end.

* fix: agent tools review — next_steps order, resolve caching, error semantics

- Large-doc next_steps now says structure-first, consistent with tool
  descriptions and agent instructions
- _remove_document fetches document list once instead of per-name
- call_tool returns error envelope for unknown names instead of raising
- _not_ready_error timed_out flag reflects actual wait outcome
- openai_agents.py docstring corrected to match default (FunctionTools)
- Removed unused ModelSettings import from demo

* fix: agent tools review 2 — bridge thread safety, browse paging, metadata merge

- McpBridge reads session/protocol headers under the lock (now RLock:
  _ensure_initialized posts while holding it). openai-agents runs sync
  tools on threads and executes parallel tool calls concurrently, so
  bridge functions genuinely race; a torn read sent a new session id
  with a stale protocol header. Measured: one session expiry under 8
  threads cost 4 initializations before, minimal 2 after.
- Session-expiry retry also resets the negotiated protocol version, so
  the re-handshake carries no stale MCP-Protocol-Version header.
- browse_documents time sort pages list_documents natively instead of
  fetching the whole library to slice one window (relevance still needs
  the full list for scoring).
- _await_completion: a status refetch that nulls out metadata no longer
  clobbers the listing's copy (setdefault was a no-op on existing None).
- Structure tool reads the raw stored tree via a named LocalAPI
  raw_tree() seam instead of reaching into _api._store internals; drop
  the redundant deepcopy before _format_structure (store re-reads from
  disk, formatting builds fresh containers).
- Shared pageindex/_version.py replaces _sdk_version duplicated in
  mcp_bridge and the Claude integration.

Left as-is after source verification against the cloud MCP: first-page
budget bypass, pageNum falsy-zero, and the page-gap fallback text are
letter-for-letter cloud behavior — parity wins over local repair.

* fix: agent tools review 3 — page-span cap, duplicate names, wait resilience, contract drift

- _parse_page_spec bounds the requested span arithmetically (10k pages)
  before materializing it; pages="1-1000000000" previously expanded to a
  billion integers inside the caller's process.
- Local submit_document uniquifies document names the way the cloud
  upload does (taken name -> _1.._99, then reject with the cloud's own
  message). Same-name duplicates broke name-addressed tools: resolution
  always picks the newest, so older duplicates were unreachable.
- agent_instructions(doc_id=...) now fails loud when the pinned doc's
  name is shadowed by a newer same-name document (legacy stores predate
  the rename) — it previews resolution with the same _resolve_document
  the tools use, so the check cannot drift from actual behavior.
- submit_document(wait=True) tolerates transient network errors, not
  just API errors; a dropped connection at minute 25 of a 30-minute
  wait no longer kills it. Third strike wraps into PageIndexAPIError
  per the documented contract.
- The live contract-parity test compares full per-param schemas, not
  just names and descriptions. It immediately caught real drift the
  shallow check had been passing: the server now emits nullables as
  anyOf unions and stamps MAX_SAFE_INTEGER maxima on offset/part.
  Contract and snapshot updated to the served wire form; _annotation_for
  learned anyOf so bridge signatures stay Optional[str] instead of
  degrading to Any.

Adjudicated, not changed: the allowed_tools wildcard example stays
(docstring advice covers scoping; Ray's call), and raw-length response
accounting stays (letter-for-letter cloud behavior, parity wins).

* feat: surface the stored document name from submit_document

Compute PR #558 makes /doc/ return {"doc_id", "name"} carrying the
post-dedup-rename name. Mirror it end to end: local submit returns the
stored name, the client warns when it differs from the uploaded file
name (read via .get so older cloud servers stay compatible), the local
name-exhaustion check runs before indexing instead of after the LLM
spend, and the demo caches doc_id in a file instead of name-matching —
a renamed document made the name lookup re-index on every run.

* fix: add missing page_list kwarg in duplicate-name test mock

* revert: keep README.md unchanged from main — SDK section deferred

* feat: serve cloud agent instructions live from the MCP server

The cloud MCP server publishes its agent instructions in the initialize
result, adapted to each key's tool set. agent_instructions() previously
returned the SDK's local-subset text in both modes — a silently forked
copy that lacks the guidance for cloud-only tools (search_documents
escalation, folders, images) and drifts as the server's prompt evolves.

Cloud clients now serve the server's live instructions, captured from
the initialize handshake on a per-client bridge shared with
agent_tools() (one session, no extra request). An empty server response
raises instead of silently substituting the subset text — same posture
as the annotation-regression guard. The local constant stays as the
honest subset for the in-process tools, with its provenance noted and a
consistency test that every tool it names exists in the local registry.

* fix: local relevance sort answers honestly instead of imitating

sort="relevance" is cloud-side semantic ranking; the local substring
imitation could satisfy the letter of the interface while silently
missing semantically relevant documents. Per the honest-subset rule
(same treatment as folders), local now returns the "not available
here" envelope for sort="relevance" or a stray query, and the local
instructions steer discovery through name/description matching plus
full-library paging instead of prescribing a capability that does not
exist here. The tool schema keeps the cloud contract verbatim, like
folder_id: honesty lives in the runtime answer, not a forked contract.

* docs: note the cloud+Claude instructions duplication trade-off in as_claude_mcp

* fix: unsupported-capability envelopes say local-mode-yet, point to cloud

"Not available here" read as a broken feature; the honest framing is
that folders and semantic ranking exist on PageIndex cloud and are not
in local mode yet. Both envelopes now say so and name the cloud client
in next_steps, so agents relay an accurate story to the user.

* fix: local tool descriptions pre-announce cloud-only capabilities

The cloud-verbatim browse_documents description invites
sort="relevance" and folder drilling, so a local agent's first semantic
search attempt was a guaranteed dead end discovered only from the
runtime error envelope. Local registration now appends a LOCAL MODE
note to the description — the agent learns what is cloud-only before
calling; the runtime envelope stays as the backstop for prompts that
ignore descriptions. The cloud-facing contract stays byte-verbatim.

* refactor: localized tool guidance replaces the appended LOCAL MODE note

Appending a retraction to the cloud-verbatim description left the model
parsing an instruction and its negation — and kept the cloud text
recommending search_documents and get_folder_structure, tools that are
not registered locally (get_page_content likewise pointed at
get_document_image). Guidance now adapts to the local surface the way
AGENT_INSTRUCTIONS already does: schema structure stays byte-identical
to the contract (mechanically asserted by a strip-descriptions test),
while local description strings teach only what works here and point to
PageIndex cloud for the rest. A dead-reference test forbids local
guidance from naming tools outside the local registry, so a contract
refresh that reintroduces a cloud-only reference fails loudly.

* feat: hide cloud-only parameters from the local tool surface

folder_id, sort, query, and recursive were exposed locally with
localized "cloud-only" descriptions, leaving the dead-end calls
expressible and discovered at runtime. Schema constraints beat
guidance: the local surface now serves the contract minus these
parameters, so strict-schema frameworks make the calls inexpressible
and a prompt that insists on sort="relevance" degrades to the bare
call (the correct local behavior) instead of an error round-trip.

The implementations still accept the hidden parameters and answer with
the guided "works on PageIndex cloud" envelope — the backstop for
direct call_tool callers and hosts without schema enforcement.
wait_for_completion stays: seeded or torn stores can hold documents
that are genuinely not completed. The structural guard now asserts the
local schema equals the contract minus the documented hidden set,
descriptions aside.

* fix: incremental-review findings — bridge cache, guards, envelope drift

Three independent review passes over the agent-instructions increment
surfaced six fixes:

- The per-client bridge moved off the instance into a weak-keyed,
  lock-guarded module cache: cloud clients stay picklable
  (threading.RLock no longer rides on the client) and concurrent first
  calls can no longer construct duplicate bridges/sessions.
- Blank or non-string initialize.instructions now hit the same honest
  error as a missing one — a whitespace-only or structured value could
  previously become the system prompt (or crash the doc_id append with
  a raw TypeError).
- The invalid-sort envelope no longer prescribes sort="relevance" — the
  one error text that still taught the cloud-only value it would then
  reject.
- "Page through the rest of the library" is emitted only when has_more
  is true; a fully-listed library no longer instructs a pointless call.
- The mandatory full-library paging step now says limit: 50 — 6 calls
  instead of 30 on a 300-document library.
- Docstrings and comments rescoped to what is actually true: the
  never-raise contract covers invocations the signatures accept
  (unknown params fail at the Python boundary; call_tool answers them
  with the guided envelope), recursive is accepted as the identity
  rather than errored, lenient framework arg models drop hidden params
  pre-call, and the module header no longer claims full schema parity.
  The capability-phrase guard now covers every local docstring, not
  just browse_documents.

* chore: keep the demo's doc_id cache file out of the repo

* test: live envelope field-parity guard against cloud response drift

The frozen contract guards tools/list, but the response envelopes the
local tools emit were hand-built to mirror the cloud's and had no drift
detector. A key-gated live test now asserts every field local emits
exists in the live cloud response for the analogous call (top-level
keys, next_steps, document entries, structure nodes, content entries).
Guidance wording is deliberately localized and not compared. Verified
green against the live server: local and cloud field structures
currently match exactly.

* feat: local chat — three protocol surfaces over the agent tools (v0.2.10)

Local mode gains managed document QA: an agent over the #393 local tool
set, reachable through three wire protocols, each 1:1 with the backend
and with no translation layer.

- chat_completions(): standard chat.completions semantics on any
  OpenAI-compatible backend (openai-agents engine). Final answer only,
  cross-turn aggregated usage, streaming as text pieces or chunk dicts
  (the existing cloud signature, now implemented locally; model and
  max_turns are local-only additions).
- responses(): the agentic surface — OpenAI Responses format, the tool
  process is standard output items, streaming forwards native events
  (tool outputs emitted as response.output_item.done, the way the
  platform streams its own server-side tools). Round-tripping output
  into the next input keeps provider prompt-cache prefix continuity and
  the agent's memory — live-verified: the follow-up call answered from
  round-tripped tool output with zero new tool calls.
- messages(): Anthropic-native via the SDK's own tool runner (new
  pageindex[anthropic] extra, floor 0.68.0 verified for
  tool_runner/beta_tool(input_schema)). tool_use/tool_result round-trip
  is the format's native behavior; the envelope is the final message
  with aggregated usage plus the full new-turn sequence; the managed
  system blocks carry cache_control breakpoints.

Shared skeleton: thin chat header + the local AGENT_INSTRUCTIONS
(caller system content is appended, not rejected), the doc_id targeting
block as a leading context item (factored out of
build_agent_instructions), read-only toolset, structural-only
validation (no arbitrary caps — backend limits govern), sampling params
passed through, per-run tracing disabled, enable_citations rejected as
cloud-only. Design basis is industry-standard formats rather than the
cloud chat endpoint; responses()/messages() raise on cloud clients
until the cloud converges.

Tests run the real engines against scripted backends (a Model fake for
openai-agents, a mock HTTP transport under the real anthropic SDK) with
real tool execution against a seeded store, including the round-trip
prefix-extension assertions on both engines.

* fix: local-chat review findings — truncation, serialization, streams

Three independent review passes (bug scan, claims-vs-code, adversarial
runtime probes) over the local-chat increment; every fix below was
reproduced before being fixed.

messages():
- A max_turns cut no longer duplicates the final assistant turn: the
  runner has already appended it when iterations exhaust, so the
  round-trip history carried a duplicate tool_use id and ended on an
  unanswered tool_use — a guaranteed 400 on continuation. The append
  now keys on stop_reason, and truncation reads natively as
  stop_reason: "tool_use" with a continuable history.
- The envelope is JSON-serializable end to end: runner-stored turns
  carry pydantic content blocks; everything is dumped to plain dicts,
  excluding SDK-internal __api_exclude__ fields (parsed_output) that
  the API rejects on round-trip.
- Bounded by default (max_iterations 10, like the OpenAI surfaces);
  usage aggregation now preserves the final turn's native fields and
  sums the token counters None-safely; empty caller system strings are
  skipped; non-dict message entries and bad doc_id types raise
  PageIndexAPIError; anthropic < 0.68 gets an actionable version error;
  the doc block no longer spends a cache_control breakpoint.

chat_completions()/responses():
- MaxTurnsExceeded wraps into PageIndexAPIError on all four run paths.
- responses(stream=True) is one logical response: per-turn backend
  lifecycle events are collapsed (a canonical consumer previously
  stopped at turn 1's response.completed and never saw the answer),
  sequence numbers are reassigned monotonically, and the synthesized
  tool-output event carries output_index/sequence_number.
- The responses envelope carries the real request surface
  (instructions, the actual function tool definitions, tool_choice,
  parallel_tool_calls, error/incomplete_details).
- RunConfig(group_id) pins a stable prompt_cache_key: openai-agents
  otherwise stamps each run with a fresh key, tagging round-tripped
  prefixes as different cache groups and defeating the feature the
  round-trip exists for.
- Abandoning a stream now cancels the run: a watchdog task lets the
  cancellation land even while the pump awaits the backend, and the
  per-call AsyncOpenAI client is closed before its loop ends (fixes
  "Task exception was never retrieved" noise). The opening role chunk
  is emitted even for empty outputs; empty responses() input and
  enable_citations-before-extra ordering fixed.

Docs rescoped to what is true: finish_reason/status reflect loop
completion on the OpenAI surfaces (the engine does not surface per-turn
backend reasons); chat streaming yields visible narration including
pre-tool text; messages(stream=True) forwards the Anthropic SDK's
native event objects (not wire-verbatim); the doc block is a leading
conversation item on OpenAI surfaces and a system block on messages().

Tests: 25 in the file (11 new), with per-extra skip sections so a
machine with only one framework still covers the other surface;
without-frameworks matrix re-verified; live smoke re-run green with a
clean exit.

* feat: as_anthropic_tools — Anthropic tool-runner export, both modes

Fills the last cell of the agent-connection matrix: users driving their
own anthropic tool_runner loop get runnable tools directly. Cloud wraps
the live MCP tool set with input schemas passing through verbatim (MCP
inputSchema is the Messages API schema shape); local exposes the same
set messages() runs internally. The beta_tool wrapping moves from
local_chat into integrations/anthropic_sdk.py, parallel to
openai_agents.py, and messages() now consumes the shared builder.
agent_tools grows _bridge_invoker/_read_only_tools so the plain-function
and beta_tool cloud paths share invocation containment and the
read-only gate.

* fix: as_anthropic_tools review findings — async flavor, schema isolation

Adversarial + best-practice review of 4590dd8 (three independent passes)
surfaced two holes. The export was sync-only: AsyncAnthropic's runner
accepts only BetaAsyncFunctionTool and splices anything else into the
request body unserialized, so the first call died with an opaque
TypeError — asynchronous=True now builds beta_async_tool runnables
(present since the 0.68.0 floor) that run the blocking bridge/store call
in a worker thread, keeping I/O off the caller's event loop. And
beta_tool stores input_schema by reference, so cloud tools aliased the
bridge's cached metas while the local path deep-copied — the builder now
copies, and the passthrough test asserts equal-but-not-aliased so it can
no longer compare an object with itself. Docstring fixes from the same
round: the MCP-connector pointer now carries the full live-verified
shape (authorization_token was missing — following it literally gave a
401), and the manual messages.create loop's to_dict() serialization is
documented. Tests pin the runnable flavor both ways (isinstance), which
existing tests could not distinguish.

* docs: doc_id is per-call table-setting — keep it identical across a conversation

The targeting block doc_id adds is re-set on every call and sits in the
cached prompt prefix, so a round-trip that drops (or changes) doc_id
silently diverges the prefix and loses the cache continuation. State the
rule on all three chat surfaces' doc_id docs, and pin it with a prefix
test that passes the same doc_id on both calls.

* feat: every chat surface takes a bare query string

query + doc_id is the minimal PageIndex contract, so it now works
uniformly: chat_completions and messages accept a plain string (one
user message), as responses always did per its wire format. The wrap
is input sugar at the SDK surface, not a translation layer — the
outgoing wire is unchanged, and managed agent surfaces taking strings
is the ecosystem convention (Runner.run, claude_agent_sdk.query).
Cloud chat_completions gains the same acceptance; blank strings raise
on every path.

* feat: messages() defaults max_tokens to 4096

The Messages API requires a per-turn output budget on the wire, but
that is table-setting, not a PageIndex-layer user obligation — the
simple call is now a question + model + doc_id. The knob stays
overridable (passthrough intact); model stays required because no
cross-vendor default is honest to guess.

* fix: raise messages() max_tokens default to 8192

max_tokens is a cap, not consumption, so the default should be the
highest universally safe value: 4096 could truncate long-form answers
(whole-document summaries), while 8192 is the output ceiling every
non-EOL Claude model accepts and stays under the SDK's non-streaming
long-request threshold.

* fix: restore per-extra skip markers the string-input tests displaced

Inserting tests above decorated ones absorbed their @needs_agents
markers, so two tests ran (and failed) in the without-frameworks CI
job. Both simulated-bare and full runs are green again.

* fix: close 17 findings from the v0.2.10 max review

Tool layer:
- anthropic adapter: failed tool calls raise ToolError so the runner
  emits tool_result is_error:true; McpBridge.call_tool returns
  (text, is_error) and surfaces the server's MCP isError marking
- as_openai_tools builds FunctionTool with the contract/server schema
  verbatim (strict off) — function_tool() regenerated schemas from
  signatures, dropping items/enum/pattern/bounds and aborting the whole
  list on object-typed params; shared _tool_specs() feeds both adapters
- remove_document validates every name before deleting anything; call_tool
  classifies only bind-time TypeErrors as INVALID_INPUT
- unknown-tool envelope formatted with _dumps like every other envelope

Local chat:
- doc_id is enforced at the tool layer (allowlist threaded through
  call_tool and the adapters), not just prompted; the shadow check runs
  inside the scope
- _openai_model routes litellm/ and provider/ paths via LitellmModel and
  strips openai/ — the normalized retrieve_model 404'd as a raw wire name
- responses() reports the backend's real terminal status (recorded at the
  transport client; the framework discards Response.status) and wraps
  framework exceptions in PageIndexAPIError
- chat_completions streaming yields its opening chunk inside try, so an
  abandoned iterator still cancels the run and closes the backend
- prompt-cache group_id is per-conversation (model+instructions+first
  item) instead of one global constant pooling every user
- messages() max_tokens default resolves per model (claude-3 caps at 4096)

Packaging / surface:
- __init__ registers the 0.2.10 modules in _SUBMODULES; unknown names
  raise AttributeError instead of eagerly importing page_index_classic
- anthropic floor 0.84.0: first release with ToolError whose runner also
  executes the final turn's tools on a max_iterations cut
- client docstrings caught up with local chat landing

Claude Agent SDK gate:
- claude_allowed_tools(mcp_servers) derives mcp__<key>__<tool> entries
  from the caller's own registration map (live server annotations on
  cloud, the contract locally) — no name is ever spelled twice
- claude_agent_config() bundles the three slots as one-call sugar over
  the explicit form

Examples:
- demo runs against cloud again (getattr for local-only attrs) and finds
  an existing indexed copy by name before re-indexing

Tests: monkeypatches replace the consuming module's binding instead of
mutating the shared time/requests modules; 185 -> 211.

* feat: one-call config bundles for every bring-your-own-framework surface

claude_agent_config() gets two symmetric siblings, so each framework's
front door is a single splat over the same explicit primitives:

- openai_agent_config(): Agent(**...) kwargs — instructions, tools, and
  the local retrieve_model (cloud omits model for the framework default)
- anthropic_runner_config(): tool_runner(**...) kwargs — system, tools,
  and the messages() defaults (per-model max_tokens, 10-iteration bound);
  only the user's messages remain

Bundles stay pure sugar: doc_id rides agent_instructions, no extra
semantics over the explicit form, docstrings point both ways. The demo
agent shrinks to Agent(**client.openai_agent_config(doc_id=...)).

Construction is pinned against the real frameworks in tests (Agent and
tool_runner both built offline), so an upstream kwargs rename fails
loudly; 211 -> 215 tests.

* fix: three more review findings — partial-read reporting, reply correlation, output_index axis

- get_page_content: the summary is additive, not either/or — a call that
  both truncates for size and has out-of-range pages reported only the
  latter, telling the agent every in-range page was returned (#2)
- McpBridge._extract_result: strict request-id correlation only; the
  eager fallback could hand back a stale or mis-correlated JSON-RPC
  message as this call's reply (#16)
- responses() streaming: output_index now addresses the logical
  response.output — backend per-turn indexes are re-based past prior
  turns' items and the SDK-injected tool outputs take the next slot on
  that axis, instead of reusing the event-sequence counter (#15)

215 -> 217 tests.

* feat: gate the config-handoff surfaces by the read-only MCP endpoint

pageindex-chat#448 adds /mcp?tools=read — the server registers only
readOnlyHint-annotated tools — so the URL itself becomes the gate for
every surface that hands a config to a third party:

- as_claude_mcp: include_management now picks the endpoint on cloud;
  the parameter is real in both modes
- as_openai_tools(hosted=True): OpenAI connects to the read-only
  endpoint by default and require_approval simplifies to "never" — the
  approval-flow middle ground becomes hard absence, matching every
  other surface's default
- claude_allowed_tools() retired before ever shipping: with the server
  gated, allowed_tools degenerates to whole-server pre-approval, which
  claude_agent_config emits as the constant ["mcp__<name>"] — no
  setup-time bridge round-trip remains
- in-process surfaces (agent_tools, as_openai_tools, as_anthropic_tools
  over the bridge) keep bare /mcp + client-side annotation filtering:
  they materialize tools locally and hand no URL to anyone

Release ordering: 0.2.10 must ship after pageindex-chat#448 deploys —
an older server ignores unknown query params and would silently serve
the full set behind a URL that promises read-only.

* fix: chat_completions wraps framework exceptions like responses()

The AgentsException -> PageIndexAPIError wrap from the responses() fix
covered only that surface; a backend stream dying without a terminal
event (or any engine failure) still escaped chat_completions as a raw
openai-agents exception type on both its paths.

* fix: same-name documents in different folders no longer refuse doc_id targeting

The shadow check in doc_targeting_block compared names across the whole
library, so agent_instructions(doc_id=...) hard-raised for a legal cloud
layout — one file name in two folders — with advice (rename/remove) that
contradicts the contract, whose folder_id parameter exists precisely to
disambiguate this case.

Shadowing is now judged per folder: only a newer same-name document in
the SAME folder makes the name unreachable and raises. A same-name
document in another folder serves the call, and the targeting block adds
a directive to pass folder_id on every tool call — dropping the raise
alone would have traded a loud refusal for the agent silently reading
the newer document.

Local mode (folderId always None) and the scoped chat path (allowlist
resolution, fixed with the doc_id enforcement) are behaviorally
unchanged.

* Revert "fix: same-name documents in different folders no longer refuse doc_id targeting"

This reverts commit 9d16dbf6, whose premise collapsed on verification
against the cloud upload paths. Review finding #10 inferred from the
folder_id tool description that one file name in two folders is a legal
cloud layout; both upload paths actually dedup names per USER SPACE with
no folder dimension — chat's getSignedUploadUrl queries fileName +
sourceName + mode + owner (file-access.service.ts), and compute's
get_upload_url probes the S3 key (user, source, file_name) — so own
documents cannot share a name across any folders. The only legitimate
same-name source is shared mounts (shared-with-me/following), which the
api-proxy surface this SDK talks to never carries.

Cloud and local therefore share one invariant — names unique per space,
server-enforced — and the original global shadow check was the right
shape: a duplicate is an anomaly worth refusing loudly, not a layout to
accommodate with per-folder adjudication and conditional prompt notes.
The invariant is now stated in doc_targeting_block's docstring so the
finding does not get re-raised.

* docs: state the name-uniqueness invariant in library terms

* docs: trim doc_targeting_block docstring to the contract

* fix: compress out-of-range page lists in get_page_content

Two message strings enumerated every out-of-range page number one by one
while the payload fields beside them already used _format_page_spec.
Against a 2-page document, pages="3-10000" produced a 59,310-character
error whose own requested_pages field expressed the identical set as
"3-10000"; the mixed case pages="1-10000" produced 59,479. Both now
render through the helper: 431 and 600 characters.

This was inherited behaviour, not a local slip — the cloud MCP server
enumerated at the same two sites, so local reproduced it verbatim. The
cloud fixed it first (pageindex-chat #449), and this follows to keep the
strings byte-identical; the error message now matches the served one
character for character. A differential run of the two compressors over
94 inputs (empty, single, unsorted, duplicated, 10k spans, 80 random)
agrees on every one, separator included.

The new test pins all three shapes, including the non-contiguous case
("5,9" must not collapse into a range) that the compressor had no direct
coverage for.

* docs: messages() marks only the managed prefix with cache_control

The method docstring claimed the doc targeting block carries a
cache_control breakpoint too, and the doc_id note called that block part
of the cached prompt prefix. _anthropic_system deliberately marks only
the stable managed prefix — the API allows four breakpoints and the
varying doc block must not consume one — and the block is appended after
the sole breakpoint, so it is never cached.

a45b554 added the block with cache_control, making the claim true when
written; daac9d2 removed it without touching the docstring, and adb2f1f
then added the "cached prompt prefix" sentence after the fact. The same
phrase at the chat_completions and responses docstrings is correct —
there the block is a leading conversation item inside the auto-cached
prefix — so only the Messages surface is reworded.

The doc_id advice itself stands: the block is per-call table-setting and
should stay identical across a conversation. Only the caching rationale
was wrong.

* docs: _stream_sync cancels on close, not on abandonment

The docstring promised that closing "or abandoning" the iterator cancels
the run. Abandoning only works when refcounting collects the generator:
a caller that breaks out of the loop while keeping the reference never
runs the finally that sets the cancel event, so the pump thread stays
parked on the full queue and the backend client is never released.

Closing is correct and is what the dedicated test exercises. Narrowing
the promise to the behaviour the code actually provides is the honest
fix; a watchdog or finalizer would be machinery bought for a shape the
sync surface is not meant to serve, and the async client planned for
0.2.11 gets native task cancellation instead.

* fix: raise the openai-agents floor to 0.14.0

_conversation_group_id feeds RunConfig.group_id into OpenAI's
prompt_cache_key so a round-tripped prefix stays in one cache group.
That wiring first appears in openai-agents 0.14.0: 0.8.0 through 0.13.x
have no prompt_cache_key at all, and group_id there is a tracing group
id only — inert, since tracing is disabled on the line above. An install
resolving to the declared floor lost the cache continuity that the
responses() docstring sells, silently and with no test able to catch it.

The old floor's rationale (0.8.0 offloads sync tools to a thread) is
subsumed by the new one. Every symbol the package imports predates
0.14.0, so nothing else constrains the bound.

* fix: enforce doc_id at the tool layer in the framework config helpers

openai_agent_config / anthropic_runner_config / claude_agent_config
accepted doc_id but built unscoped tools, so the parameter that is a
structural allowlist on chat_completions() was prompt-only advice here —
the agent could read every document in the store regardless.

- as_openai_tools / as_anthropic_tools / as_claude_mcp take a doc_id
  tail parameter and thread it to the existing _allowed_ids channel;
  the config helpers pass it through in local mode
- cloud config helpers keep prompt-level targeting (tool scoping is
  server-side there, documented); explicit as_*(doc_id=...) raises on
  cloud instead of silently dropping the allowlist — including the
  hosted branch, which returned before _tool_specs' existing guard
- _require_local_scope consolidates the cloud rejection that was
  inlined in _tool_specs
- doc_id=[] is an empty allowlist, not "unscoped": dropped the
  `or None` at the three local chat surfaces

* fix: two chat findings — final-turn append and cache-key seeding

run_messages keyed its re-append guard on stop_reason, but the anthropic
runner executes tools whenever the turn's content carries tool_use
blocks (refusal excepted) — a max_tokens turn with complete tool_use
blocks was already appended by the runner, so the guard re-appended it,
duplicating tool_use ids and 400ing the documented verbatim
continuation. The guard now checks whether final's tool_use ids already
sit in the appended history; unexecuted tool_use blocks (refusal turns)
are stripped from the appendable history, as the SDK itself does when
rebuilding params around an unresulted turn.

_conversation_group_id seeded on items[0], which is the doc-targeting
block whenever doc_id is set — byte-identical across every conversation
about a document, so all of them pooled under one prompt_cache_key and
evicted each other's prefixes. Seed on the conversation's own first
item instead: continuations keep their key, unrelated conversations
never share one.

Also drop the dead pytestmark_openai assignment (pytest's magic name is
pytestmark; the section gate it implied never existed).

* fix: six review findings — pagination, compat, and containment

- _all_documents advances by what actually arrived and treats `total`
  as an optimization: absent/null totals and short pages silently
  truncated the library behind every name resolution
- _make_bridge_function survives description: null (the parallel
  _tool_specs path already did)
- as_openai_tools answers a malformed argument string with the guided
  error envelope instead of raising through the caller's whole run
- the pre-0.2.10 package attributes (ConfigLoader, count_tokens, ...)
  resolve again: main's underscore-guarded fallthrough is restored —
  dunder probes stay lazy, a non-underscore typo pays one classic
  import before its AttributeError
- _split_structure chunks are always lists: the structure field no
  longer changes JSON type between parts of one paginated response
- the bridge replays only session-carrying 404s (the spec's expiry
  status); 400 raises instead of re-running side effects, and the
  reset double-checks under the lock so concurrent retries cannot
  clobber a freshly re-initialized session

* fix: five secondary review findings — containment and guards

- the bridge maps content blocks individually: base64 payloads
  (image/audio) become metadata stubs instead of handing the model the
  raw blob, text blocks pass verbatim, anything else keeps the JSON
  dump (revisit if tool results become real multimodal input)
- cloud proxy annotations keep array item types (list[str], not bare
  list) so strict function calling accepts the round-trip; a type-array
  in items degrades to bare list instead of crashing the build
- run_messages raises when set_messages_params stops delivering params
  instead of silently dropping every tool turn from the envelope
- call_tool drops None-valued arguments (None ≡ omitted, the
  contract's semantics) — adapters that forward the model's nulls
  verbatim no longer trip parameter validation
- client._parse_pages bounds the span arithmetically before
  materializing it, like the tool layer: "1-999999999" raises instead
  of allocating a billion integers

* fix: three review findings — protocol honesty, model echo, containment

- responses() promised the Responses protocol ("no translation layer")
  but _openai_model ignored protocol on the LiteLLM branch: provider-
  prefixed models silently ran chat.completions under a responses-shaped
  envelope, and with no transport hook to record status (LitellmModel
  has no _client.responses) a turn truncated at the output cap reported
  status "completed". The branch now raises for protocol == "responses"
  — at agent-build time, before any backend call — naming the routes
  out: chat_completions(), messages() for Anthropic models, or
  OPENAI_BASE_URL + a bare/openai/-prefixed name for backends that
  genuinely speak /responses. Refusal, not emulation: most providers
  have no /responses endpoint to drive.
- chat_completions envelopes echoed retrieve_model verbatim, which
  carries the SDK's litellm/ routing marker after normalization — a
  name no provider catalog contains, and a different string than the
  same model passed per-call. The envelope and every streaming chunk
  now report the name the provider actually serves; routing and the
  prompt-cache group key keep the prefixed form. responses() needs no
  change (post-refusal the prefix cannot reach its envelope), and the
  user-typed openai/ prefix stays echoed as typed.
- _remove_document caught only PageIndexAPIError around the per-doc
  delete, so a bare OSError (local_store re-raises them) or a transport
  error (cloud delete_document wraps nothing) escaped mid-batch,
  discarded the entries for documents already irreversibly deleted, and
  surfaced as a generic INTERNAL_ERROR envelope inviting a retry — which
  then reports the destroyed document as not_found. The loop now catches
  Exception, keeping the per-document results the contract promises.

* fix: config bundles use the scoped shadow check their tools earned

9f67fdd made the three config helpers enforce doc_id at the tool layer
but left their instructions on doc_targeting_block's unscoped default,
so a bundle refused any doc_id whose name a newer library-wide
duplicate shadows — a raise whose message ("the tools address documents
by name and would read the newer one") had just become false: the
bundle's own tools resolve names inside the allowlist and read the
targeted document correctly. chat_completions() accepted the same
doc_id via _doc_block's scoped=True.

Each helper now computes scope = _local_doc_scope(doc_id) once and
derives both slots from it — scoped=scope is not None for the
instructions, doc_id=scope for the tools — so the check mode and the
tool allowlist come from one fact and cannot drift apart again.
build_agent_instructions grows a scoped passthrough; cloud stays on the
whole-library check (scope is None there and the tools are genuinely
unscoped), and the public agent_instructions() keeps its unscoped
default for the same reason. An in-set duplicate still raises — and in
that case the message is true on every surface that emits it.

* docs: as_openai_tools' remote-MCP note moves to the Cloud paragraph

The MCPServerStreamableHttp alternative sat in the Local: paragraph
pointing at bare {BASE_URL}/mcp — a cloud-only route (BASE_URL is the
hosted API; local has no HTTP MCP server) that as written would connect
unauthenticated to the full tool set. Now stated where it applies, in
the as_anthropic_tools connector-note form: Cloud paragraph, Bearer
auth spelled out, ?tools=read default with the drop-it escape.

* fix: nine review findings — argument coercion, scope, and honest envelopes

- call_tool coerces string booleans per the TOOL_CONTRACT schema
  ("false"/"no"/"0" read as False, not a truthy 3-minute wait) and
  survives arguments: null (json.loads("null") reaches the seam as None)
- _local_doc_scope raises on an explicitly empty doc_id on cloud: with
  no tool-layer allowlist there, dropping it silently widened an empty
  scope to the whole library
- both page-spec caps count distinct pages instead of summing parts, so
  overlapping ranges (a parent section plus its children) within the
  10k union pass again as they did in 0.2.9; the per-part arithmetic
  bound still rejects billion-page specs before materializing anything
- _remove_document deduplicates doc_names: a repeated name is one
  deletion, not a second "failed" row with an internal error string
- doc_targeting_block merges the user's metadata tags from the listing
  (local get_document keeps the 7-key cloud detail wire shape, which
  carries none) so the block delivers the metadata it promises
- _wait_until_ready folds its two raise branches into one that carries
  the doc_id: a poll that dies no longer discards the handle to an
  uploaded, billed document
- _reported_model strips both routing prefixes (litellm/ and openai/)
  and responses() now reports it too, instead of echoing a model id the
  provider never served
- _openai_model wraps AsyncOpenAI() construction so a missing backend
  credential surfaces as PageIndexAPIError like every other gate on the
  chat surfaces (and builds the client once for both protocols)
- _browse_documents advances its cursor by the rows that actually
  arrived and guards a null/absent total — the same hazards
  _all_documents already guards — and an empty window ends pagination
  instead of freezing the cursor

* fix: two chat findings — protocol terminal states, provider error types

- responses(stream=True) raised PageIndexAPIError when the backend
  ended the response with response.failed / response.incomplete:
  openai-agents yields the terminal lifecycle event, then re-raises it
  as ModelBehaviorError, so the generic AgentsException wrap
  short-circuited the emit the agen's tail was built for — its
  failed/incomplete terminal mapping was dead code against the real
  engine, and the caller lost both the partial output and the real
  status. The wrap now steps aside when the recorded terminal state is
  failed/incomplete, and the stream ends with the honest terminal
  event (committed output, real status, error/incomplete_details) —
  the backend's terminal state is a protocol event, not an engine
  failure. Non-stream was already honest for incomplete via the
  transport recorder; a failed response arrives there as an HTTP
  error, covered below. Known ceiling: the truncated final turn's
  partial text was already streamed as deltas but is not reconstructed
  into the terminal event's output (the engine commits items only on
  turn completion).
- Provider exceptions (network, auth, rate limit) leaked as raw
  openai/anthropic types through every chat surface, against the
  layer's own "never raw engine types" contract. Every engine boundary
  now wraps its vendor's base exception into PageIndexAPIError
  (chained): the four OpenAI-engine sites catch openai.OpenAIError —
  LiteLLM's exception types subclass openai's, so one handler covers
  both routing paths — and messages() catches anthropic.AnthropicError
  around the batch drive and the stream generator.

* fix: guided failure for unknown LiteLLM providers, non-object call_tool args

- _openai_model pre-checks the first path segment against
  litellm.provider_list (fail-open if the attribute ever disappears):
  a HuggingFace repo id like Qwen/Qwen2.5-7B-Instruct on an
  OpenAI-compatible server now fails at build time with the escape
  spelled out — 'openai/<id>' plus OPENAI_BASE_URL — instead of at
  request time inside LiteLLM with "LLM Provider NOT provided". The
  slash-means-provider routing convention itself is unchanged; the
  retrieve_model and chat_completions docstrings now document it where
  they promise "any OpenAI-compatible server works"
- call_tool answers a non-dict arguments value (a JSON array or scalar
  from a misbehaving caller) with the guided INVALID_INPUT envelope
  instead of raising AttributeError through the agent loop, matching
  the openai adapter's own non-object guard

* fix: wrap litellm import in PageIndexAPIError when not installed

* fix: silence CodeQL findings — merge implicit string concat, drop unused vars

* fix: two external review findings — init-notification race, SDK floor

notifications/initialized moves inside the bridge lock: a concurrent
first use could send tools/list between the handshake and the
notification, which strict MCP servers reject with a 400 the bridge
never replays. Regression test races two threads through a stalled
notification window.

claude-agent-sdk floor rises to 0.1.53 — below it, string prompts with
SDK MCP servers (the documented local-mode flow) hit invisible
registration (#597) and a deadlock (#780).

* refactor: drop the unused exc parameter from _wrap_max_turns

The parameter was dead from the moment it was introduced (daac9d2):
the body reads only max_turns, and every call site already carries the
cause via `raise ... from exc`. The signature implied the helper
inspected the engine exception, which it never did.

No behavior change — message text and __cause__ chaining verified
identical across all four call sites (chat_completions and responses,
stream and non-stream).

* fix: raise the anthropic and openai-agents floors past broken releases

Both declared floors named a version that cannot work, and CI never
caught either because it installs the latest.

anthropic >=0.84.0 -> >=0.108.0. Probed against a mock transport: on a
turn with stop_reason="refusal" carrying a tool_use block, 0.84.0,
0.92.0 and 0.100.0 all execute the tool and post the tool_result back;
0.108.0 and later stop at the refusal. test_messages_refusal_with_
tool_use_stays_appendable asserts the latter, so that test was false at
the floor. messages() is unaffected in practice (it never passes
include_management, so remove_document is not registered), but
as_anthropic_tools(include_management=True) hands it to a caller's own
runner.

openai-agents >=0.14.0 -> >=0.18.1. 0.14.0 and 0.16.0 raise pydantic
ValidationError on InputTokensDetails.cache_write_tokens before any
request reaches the transport when paired with openai 2.54.0 — and they
declare openai <3,>=2.26.0, so pip resolves exactly that pair. 0.18.1 is
clean. The 0.14.0 rationale (RunConfig.group_id -> prompt_cache_key)
still holds above the new floor.

The three extras' floor comments are cut to the binding constraint; the
reasoning lives here.

* test: cover max_turns wrapping on every chat surface

test_chat_completions_max_turns_wrapped only drove chat_completions, so
the two responses() call sites had no coverage, and no test asserted
that the engine exception survives as __cause__. Parametrized over both
surfaces and both stream modes; the non-positive max_turns rejection
splits out, since it is input validation rather than wrapping.

* fix: four review findings — envelope size honesty, contained tool errors

- _dumps drops indent=2: emission now matches _serialized_size's compact
  accounting, so the pagination budget bounds what is actually sent
  (indented parts measured under 95k but emitted ~1.8x the 100k cap)
- call_tool builds the _allowed_ids frozenset inside the guarded block:
  a non-iterable doc_id returns the INVALID_INPUT envelope instead of
  raising into the agent loop; same move for _bridge_invoker's
  arguments normalization
- next_steps strings qualify submit_document() as
  PageIndexClient.submit_document() (three sites), matching the one
  already-qualified site — it is a client method, not a registered tool
- tests: import httpx at module scope (guaranteed via the hard openai
  dependency) so agents-gated tests survive an install without the
  anthropic extra; formatting assertion follows the compact envelope
2026-08-13 07:32:11 +08:00
Ray d375c00a5a feat: local mode for the PageIndex SDK (v0.2.9) (#389)
* feat: add local mode to the PageIndex SDK client

One PageIndexClient, two backends. With api_key: the 0.2.x cloud SDK,
request for request (with the reviewed fixes: bounded timeouts on JSON
endpoints, none on uploads, URL-encoded ids, 401 key hint, empty DELETE
body tolerated). Without api_key: the same methods run locally —
page_index builds the tree in submit_document (mode="flash" uses
PageIndex Flash), documents are stored as plain JSON per doc under
storage_path, submit_query is LLM tree search with retrieve_model, and
chat_completions answers over the retrieved nodes with OpenAI-style
responses and streaming.

Local responses mirror the cloud wire shapes verified against the server
source: tree nodes rename start_index to page_index and drop end_index,
a non-leaf summary becomes prefix_summary, and the metadata/list/delete/
retrieval envelopes match key for key. Cloud-only features (folders,
beta_headers, enable_citations) raise instead of pretending.

Replaces the demo-only workspace client (index/get_document_structure/
get_page_content had no real users) and its retrieve.py helpers.
page_index_main gains an optional logger param so the SDK can keep
./logs out of the caller's working directory; pymupdf import is now lazy
(only the optional PyMuPDF parser path needs it).

* chore: package pageindex 0.3.0.dev4 for PyPI

Poetry packaging for the combined SDK + local pipeline: every production
import is a declared dependency (openai and requests join requirements.txt
for the same reason), config.yaml and the flash data tables ship in the
wheel, the benchmark PNG does not. pymupdf drops to an optional note now
that its import is lazy. dev4 follows the already-published 0.3.0.dev1-3;
pip still resolves plain 'pip install pageindex' to 0.2.8 until a final
0.3.0 — install with --pre.

* docs: add SDK section to README; move the agentic demo onto the SDK

The demo keeps its flow and the post-cutoff demo paper, swapping the
removed workspace client for PageIndexClient local mode (list_documents
for the doc-id cache, get_tree/get_ocr behind the agent tools). The old
examples/workspace JSONs demoed the removed format and go with it.

* refactor: rebuild cloud_api on the 0.2.8 client text

Ray's rule for the cloud half: his 0.2.8 code is the base; a Kylin-lineage
change survives only when strictly better — invisible on healthy traffic
while fixing a real failure mode. Kept under that bar: request timeouts
(dead connections hung forever; uploads still pass none, exactly like
0.2.8), the upload handle closed via with (leaked on request errors),
URL-encoded path ids (a crafted id could reroute the URL), the
empty-DELETE-body guard, and stream hardening (choices guard,
response.close in finally). Reverted as not strictly better: the
lowercase summary param (the server accepts both spellings), the 401
message hint (visible text change; AUTH_HINT dropped from errors.py), and
the _request/requests.request reorganization — every method body,
docstring, and section comment is 0.2.8's text again.

diff -w against ../pageindex_sdk/pageindex/client.py now reads as that
surgical patch plus plumbing: CloudAPI reads BASE_URL/api_key through the
owning client, PageIndexAPIError comes from errors.py, and
is_retrieval_ready lives verbatim on PageIndexClient shared by both
modes. A mocked-requests harness driving 0.2.8 and this file through 19
identical calls shows the only remaining request-level difference is
timeout.

* refactor: drop the local retrieval endpoints — cloud-only, deprecated

Ray: the cloud already marks POST /retrieval/ and GET /retrieval/{id}/
deprecated in favor of chat completions, so local mode should not grow a
fresh implementation of a retiring surface. submit_query/get_retrieval
now raise in local mode with a pointer to chat_completions; cloud mode is
untouched (the endpoint still works there and 0.2.8 code keeps running).
The tree search that backed them stays as chat_completions' retrieval
engine; the retrievals/ storage goes away. This also closes the one real
cross-mode parity gap — the retrieved_nodes inner shape — by removing
its local half.

* feat: manifest.json — one-file document listings for the local store

Ray wanted a metafile that shows every document in one place instead of
per-directory reads. It is a cache, never a second source of truth:
writers update it best-effort after save/delete (atomic replace, no
locks), and list_metas trusts it only while its id set matches the docs/
directory names — documents are immutable, so matching names imply valid
content. Any mismatch (lost concurrent update, crash, corrupt or deleted
manifest) rebuilds it from the doc.json files, reading only the missing
entries. Incomplete dirs (no doc.json) stay invisible and are never
recorded, so a save that completes later is still picked up.

1000-doc listing: 37ms of per-dir reads -> 2.3ms warm (scandir names +
one manifest read); one-time rebuild 160ms.

* fix: align local doc_id prefix and createdAt format with the cloud

Local doc ids now carry the cloud's pi- namespace prefix (random token
stays uuid4 hex — nothing parses cuid internals, the prefix is the
contract; chat ids already mirrored chatcmpl-). createdAt now matches
the server byte for byte: the cloud emits the DB datetime's bare
isoformat — naive UTC, second precision — while we emitted microseconds
plus +00:00.

* feat: optional metadata tags on submit_document (both modes)

Ray's call after the alignment review: the metadata key in tree/OCR
envelopes and list entries should carry real data, and the only honest
way is exposing the field the server already accepts. Cloud mode
forwards it as the existing metadata form field; local mode validates it
early (a JSON-serializable dict, checked before any LLM spend), stores
it in doc.json, and returns it from the same three places. get_document
still omits it, mirroring the server, whose metadata-endpoint SQL never
selects that column. Scope stays deliberately narrow: set at submit and
read back — no metadata_filter, no update API.

* docs: state that createdAt is UTC and show how to localize it

The value is naive UTC in both modes (the cloud column is timestamp
DEFAULT CURRENT_TIMESTAMP on a UTC server, emitted via bare isoformat).
Wall-clock display is the consumer's layer: emitting local time under
the same format would silently change meaning per machine, and adding an
offset marker would break both the byte-format parity and string-order
sorting.

* fix: createdAt carries milliseconds, matching the cloud's datetime(3)

The earlier second-precision alignment was reasoned from the postgres
schema file, but production is MySQL (DATABASE_BACKEND defaults to
mysql) and its FilePageIndex.createdAt is datetime(3) DEFAULT
CURRENT_TIMESTAMP(3) — so the server isoformat()s a millisecond-
precision naive-UTC datetime, emitting .XXX000 fractions (bare seconds
only when the millisecond happens to be zero). Local now generates
through the same mechanism. The docs' 2024-01-15T10:30:00.000Z sample is
a JS-style string Python isoformat cannot produce — not evidence.

* fix: createdAt at millisecond precision, matching the timestamp(3) column

The cloud column is timestamp(3); its datetimes render through bare
isoformat() as six fractional digits ending in 000 (or no fraction when
the millisecond is exactly zero). Truncate to the millisecond and render
the same way, replacing the second-precision guess from cd44ef0. Worth
one confirmation against a live cloud response when a key is around.

* chore: trim non-essential comments

Alignment narration (mirrors the cloud, matches 0.2.8, trailing-period
notes) moves out of the code — that rationale lives in the commit
history. Kept only constraints the code can't show: the storage commit
protocol and manifest trust rule, the id-escape guard, the silent
logger's reason to exist, the client-reference indirection, and the
tree-node reshape spec.

* fix: close the local_store crash and corruption holes found in review

Adversarial review reproduced three edge holes in the no-lock design:
a torn delete (doc.json unlinked, rmtree unfinished) left a ghost the
manifest kept listing forever; one truncated doc.json crashed every
listing with a raw JSONDecodeError; and two concurrent deletes of the
same id could crash the loser's rmtree.

doc.json is now the existence marker in both directions — written last
on save, unlinked first on delete — and list_metas serves an entry only
after confirming it still exists, instead of trusting the manifest/dir
name-set proxy. Atomic writes fsync before replace, closing the
power-loss truncation window. An unreadable JSON file is logged and
treated as absent; a document meta additionally falls back to the
manifest copy (immutable docs make the cache a valid replica). rmtree
runs with ignore_errors: the commit point has passed and cleanup is
best-effort, which also absorbs the double-delete race.

Also restores the reversed-page-range error in the demo's page parser
(silently empty since the retrieve.py removal). Warm 1000-doc listing
goes 2.3ms -> 8.3ms for the per-doc existence check; full parse remains
37ms.

* docs: correct two docstring claims and the pymupdf note

is_retrieval_ready reports only API errors as False — transport errors
propagate; delete_document may return {} on an empty cloud body; pymupdf
is also used by tree_optimize's page loading, not just get_page_tokens.

* chore: target 0.2.9 for the local-mode release

Ray's call: 0.2.x stays the no-collections line, so local mode ships as
0.2.9 and 0.3.0 stays reserved; plain pip installs never see the
0.3.0.devN pre-releases, and the cookbooks' existing 'pip install
--upgrade pageindex' will deliver the new SDK without any --pre
instructions.

* ci: publish to PyPI on version tags

Tag-driven releases: the pushed v-tag is the single source of truth for
the version — validated as PEP 440, injected into pyproject, built, and
published via OIDC trusted publishing with no stored credentials; a
GitHub Release with the artifacts is created alongside.

* chore: tighten the store docstring to essentials

* fix: contain invalid-UTF-8 corruption; fail loud on unreadable data files

The corruption guard caught JSONDecodeError but not the
UnicodeDecodeError a torn multi-byte write produces — the exact scenario
the guard targets whenever names or descriptions carry non-ASCII text —
so that flavor crashed listings and gets raw, and regressed the old
manifest read's broader ValueError guard. _read_json now catches
ValueError, which covers both.

Unreadable tree.json/pages.json under an intact doc.json previously
served an empty tree with retrieval_ready true — a silent lie; those
paths now raise 'stored document data is unreadable' and
is_retrieval_ready honestly reports False. delete_document survives a
doc.json tampered into a directory (cleans it, reports not-found)
while real unlink failures such as permissions stay loud.

* feat: LocalClient and CloudClient for explicit mode selection

PageIndexClient(api_key=os.getenv(...)) with an unset variable gets None
and silently falls into local mode — methods keep working against local
storage on the caller's own LLM bill, the silent mode flip the
empty-string guard can't see. The explicit classes close it at the type
level: CloudClient raises on a missing key, LocalClient has no api_key
parameter at all. Names per Ray.

* feat: explicit-mode clients PageIndexCloudClient and PageIndexLocalClient

PageIndexClient(api_key=os.getenv(...)) with an unset variable yields
None and silently lands in local mode — with both modes fully working,
that's a silent mode flip onto the user's own LLM bill. The explicit
classes pin the mode at construction: the cloud one refuses a missing or
empty key, the local one has no key parameter at all. Names follow the
package's PageIndex- prefix convention.

* fix local mode edge cases

* fix: keep the litellm/ prefix normalization the demo depends on

a108c02 added _normalize_retrieve_model because the agentic demo hands
client.retrieve_model straight to the OpenAI Agents SDK, which routes a
non-OpenAI provider only when the name carries a litellm/ prefix. The
normalization sat in client.py, so the demo line never had to change --
and this rewrite dropped the helper while leaving that lone consumer
untouched. A retrieve_model like anthropic/claude-sonnet-4-6, the form
config.yaml documents, then died at Agent() with "Unknown prefix".

Local mode is indifferent to which form it gets: llm_completion and
_chat_llm both removeprefix("litellm/"), count_tokens returns the same
count either way, and _is_openai_model classifies both as LiteLLM, so
_require_llm_key still asks for no OPENAI_API_KEY. Models without a
provider path (the packaged gpt-5.4 default) pass through untouched.

* fix: restore the published 0.2.8 helper signatures the cookbooks call

pyproject.toml makes this repo the source of the PyPI pageindex package,
so this utils.py replaces the published 0.2.8 one -- whose helpers are
the documented surface of the cookbook notebooks. Three had drifted:
remove_fields lost max_len, create_node_mapping lost
include_page_ranges/max_page, print_tree lost exclude_fields. Both
README-linked notebooks open with `pip install --upgrade pageindex` and
pass exactly those kwargs, so tagging v0.2.9 as-is would TypeError
every Colab run of vision_RAG_pageindex.ipynb.

remove_fields and create_node_mapping readopt the 0.2.8 bodies, strict
supersets of the current ones (no in-repo caller passes the new params).
print_tree keeps the outline view as its default and routes an explicit
exclude_fields= to the 0.2.8 pprint view -- the two versions disagree
on what the second positional means (indent vs exclude_fields), and the
notebooks pass it by keyword. call_llm, the fifth published name, stays
out: nothing imports it -- the one notebook using a call_llm defines
its own, with a different signature.

* perf: resolve the indexing stack lazily from pageindex/__init__

`import pageindex` eagerly pulled page_index, flash, and tree_optimize
-- 0.73s warm, numpy and pypdfium2 in-process -- while the published
0.2.8 package imported in 0.08s on requests+openai alone. A cloud-only
SDK user upgrading to 0.2.9 would pay for an indexing stack they never
call, on every interpreter start.

The eager surface shrinks to client and errors (2ms warm); everything
else resolves on first attribute access via PEP 562 and is cached in
the module namespace. `pageindex.page_index` stays the function, still
shadowing its submodule as the old star-import had it. __all__ now
names the public surface, so `from pageindex import *` binds the same
working set as before instead of 124 names including stdlib modules.
A TYPE_CHECKING block keeps real signatures visible to IDEs.

The test suite's sys.modules lookup assumed the eager import chain;
it now imports pageindex.page_index explicitly.

* fix: wrap PDF read failures in PageIndexAPIError on local submit

_extract_page_texts sat one line outside the try that wraps everything
else in submit_document, so a corrupt PDF surfaced as a raw
PyPDF2.errors.PdfReadError and a password-protected one as
FileNotDecryptedError -- while a blank PDF, checked on the very next
line, got a clean PageIndexAPIError. Callers handling the SDK error
type crashed on exactly the malformed downloads and encrypted files
an ingest loop sees most.

The extraction gets its own wrap rather than joining the indexer try
below, whose except would re-prefix the blank-PDF error into "Failed
to submit document: Failed to submit document: ...". FileNotFoundError
stays native, asserted by test_submit_rejections as cloud parity.

* fix: address code review findings across local mode and publish workflow

- _parse_json_reply: switch from extract_json to _reply_json (dead
  try/except, global None→null substitution, wrong error messages)
- client.py: config override filter uses `is not None` instead of
  truthiness, so empty-string model args are no longer silently dropped
- llm_completion/llm_acompletion: raise RuntimeError after retries
  exhausted instead of returning empty string
- _require_llm_key: extend to anthropic/gemini/mistral providers
- _chat_llm: reuse _openai_sync_client singleton from utils
- _tree_search: drop redundant deepcopy before non-mutating remove_fields
- _stream_chunks: move final chunk inside try so usage data is reachable;
  close stays in finally
- _validate_chat_messages: accept system messages, merge them into the
  internal system prompt for cloud/local parity
- _build_chat_context: accept pre-read metas and pass structure through
  to _tree_search, eliminating double get_meta and double tree.json reads
- _index_standard: reuse ConfigLoader from construction; reject empty
  structure (parity with flash mode)
- extract_json: bare except: → except Exception: (no longer swallows
  KeyboardInterrupt)
- Replace all str.removeprefix() with _strip_prefix() helper to restore
  Python 3.7 compatibility; pyproject.toml back to python >= 3.7
- publish.yml: add test job (py3.10 + py3.13) gating the publish job
- Remove unused bare `import pageindex` from tests

* fix: tighten CI permissions, timestamp format, import style, and docstring

- publish.yml: add permissions: {contents: read} to the test job
- local_api: emit 3-digit ms timestamps via isoformat(timespec="milliseconds")
- local_api: use relative import for pageindex.utils in _chat_llm
- local_api: note the node_summary gate in _format_tree_node docstring
- test_client: update createdAt regex to match the new 3-digit format

* fix: accept GOOGLE_API_KEY for Gemini; mark pre-releases in GitHub

- _require_llm_key: Gemini provider now accepts either GEMINI_API_KEY
  or GOOGLE_API_KEY (LiteLLM supports both)
- publish.yml: set prerelease flag on rc/dev/alpha/beta tags so they
  are not shown as regular releases on GitHub

* fix: preserve print_tree backward compat with 0.2.8 positional call

Detect list passed as second positional arg (old 0.2.8 signature) and
treat it as exclude_fields instead of indent.

* refactor: print_tree param order — exclude_fields second for 0.2.8 compat

Move exclude_fields back to the second position (matching 0.2.8) instead
of detecting list-as-indent. Recursive call uses indent= keyword arg.

* refactor: remove _require_llm_key pre-check entirely

Let OpenAI SDK and LiteLLM report their own missing-key errors instead
of maintaining a parallel provider-to-env-var map. The except Exception
wrapper in chat_completions already converts these to PageIndexAPIError.

* fix: let LLM provider errors propagate instead of wrapping them

OpenAI/litellm auth, rate-limit, and other provider errors now reach the
caller as their original type (e.g. openai.AuthenticationError) instead
of being wrapped in PageIndexAPIError. PageIndexAPIError stays reserved
for PageIndex's own errors (bad doc_id, invalid params, indexing failures).

* refactor: catch only RuntimeError instead of isinstance check on openai

Only our own _tree_search logic raises RuntimeError (bad JSON, missing
node_list). Provider errors and unexpected bugs propagate naturally.

* fix: restore createdAt to 6-digit .177000 format matching cloud DATETIME(3)

Revert the timespec="milliseconds" that a linter introduced — it output
.177 (3 digits) while the cloud server's isoformat() outputs .177000
(6 digits). Verified against pageindex-compute server/lib/db/planet.py:
DATETIME(3) column + bare isoformat() = .177000.

* fix: resolve 15 review findings from PR #389

- Rename page_index.py → page_index_classic.py to fix __getattr__
  shadowing (function permanently replaced by submodule after import)
- Add return_exceptions=True to verify_toc, generate_summaries, and
  summarize_tree gathers so one LLM failure doesn't abort the batch
- Gracefully degrade generate_doc_description to "" on failure instead
  of discarding the entire completed index
- Normalize file_path with str()/expanduser/abspath to support
  pathlib.Path and tilde paths
- Strip text from tree.json at save time; reconstruct from pages.json
  on read via _load_tree_with_text
- Filter empty-string model overrides in PageIndexClient constructor
- Guard empty choices list before indexing response.choices[0]
- Wrap streaming iteration errors as PageIndexAPIError inside the
  generator
- Catch PermissionError in _read_json alongside FileNotFoundError
- Set max_retries=3 for OpenAI and num_retries=3 for litellm in
  _chat_llm to match the retry behavior of the indexing path
- Replace _SilentLogger with logging.getLogger(__name__)
- Add .github/workflows/tests.yml for PR and push-to-main test runs

* fix: close publish workflow injection and detect flash silent summary failure

1. publish.yml: pass VERSION via os.environ instead of shell interpolation
   into python -c, eliminating the command injection vector.

2. summarize_tree: raise RuntimeError when every node's summary generation
   fails (e.g. missing LLM credentials), instead of silently saving a
   document with all-empty summaries marked as completed.

* refactor: make chat_completions cloud-only until agent-based local chat lands

The local implementation was a tree-search RAG engine (retrieve prompt +
context stuffing) that diverged from the cloud /chat/completions design,
where an agent navigates documents through MCP tools. Rather than ship
the divergent engine in 0.2.9, remove it; local chat returns as an agent
loop built on the agent-tools layer in the 0.2.10 line.

Also from the PR #389 review:
- get_tree fails loud on unreadable pages.json instead of silently
  serving textless nodes
- drop the wasted deepcopy in get_tree (_format_tree_node builds new dicts)
- generate_doc_description catches only RuntimeError so provider errors
  (bad key, unknown model) propagate instead of storing ""
- _read_json treats IsADirectoryError as unreadable
- list_metas skips directory names that fail _is_safe_id
- constructing with api_key plus local-only args raises PageIndexAPIError
  (was ValueError) to match the empty-api_key path

* fix: verification follow-ups for the chat removal commit

- retrieve_model docstring no longer promises the deleted tree-search
  machinery; the param is reserved for the coming agent-based local chat
- empty pages.json ([]) fails loud like unreadable pages: a stored
  document can never legitimately have zero pages, and get_tree's
  "nodes always carry text" promise held only for the None case
- README: chat example moved to its own cloud-only block so the local
  quickstart no longer ends in a raise
- tests: pin IsADirectoryError loud-fail, list_metas unsafe-name skip,
  empty-pages loud-fail, and generate_doc_description's swallow-vs-
  propagate boundary

* fix: detect classic-path silent summary failure like flash already does

generate_summaries_for_structure absorbed every per-node error to
summary: "" with no systemic check, so a bad key or model name produced
an all-empty-summary tree with zero errors on the default CLI config
(if_add_node_summary: yes, if_add_doc_description: no). Mirror flash's
summarize_tree guard: partial failures still absorb, all-failed raises.

* fix: accurate wording for retrieve_model docstring and empty-pages error

- retrieve_model is not "unused": the agent demo drives its model from
  client.retrieve_model today; say so instead
- an empty pages.json is invalid, not unreadable — dedicated
  _require_pages raises "stored document has no page content" so the
  operator reindexes instead of hunting for file corruption

* chore: trim non-essential comments, fix three docstring issues

Review fixes:
- client.py: remove unverified "pending" from get_document status enum
- cloud_api.py: update stale "kept line-for-line" module docstring
- cloud_api.py: add node_summary to get_tree Args

Comment trimming across __init__.py, client.py, cloud_api.py, errors.py,
local_api.py, local_store.py — module docstrings shortened, explanatory
inline comments removed, multi-line class docstrings collapsed to one line.

* feat: add get_page_content and get_tree include_text parameter

- client.get_page_content(doc_id, pages): convenience method wrapping
  get_ocr + page filtering; _parse_pages moved from the demo into client.py
- client.get_tree(..., include_text=False): skips node text for
  structure-only views; local reads the stored tree directly (no text
  to begin with), cloud strips client-side via remove_fields
- demo simplified: tools now call the new client methods directly

* simplify: drop redundant try/except in demo get_page_content tool

* feat: add get_document_structure convenience method

* test: cover get_page_content, get_document_structure, include_text=False

* fix: guard against four edge-case crashes found in PR #389 review

- Demo: use getattr for retrieve_model so cloud clients don't AttributeError
- utils: return empty string when start/end page index is None instead of TypeError
- local_store: clean up temp file on _write_json_atomic failure
- local_api: re-raise PageIndexAPIError before the catch-all Exception block

* chore: trim verbose optional-dep comments in requirements.txt

* revert: restore README.md to main — SDK section deferred to next version

* fix: eliminate double PDF parse in local standard indexing

_extract_page_texts already reads the PDF via PyPDF2; pass the
pre-extracted texts as page_list to page_index_main so it skips
its own get_page_tokens call. Token counts are computed once via
litellm.token_counter in _index_standard.

* test: assert page_list is passed and correctly shaped

* fix: wire include_text to cloud API and guard get_page_content on processing docs

cloud_api.py now sends include_text as a query param so the server can
omit node text from tree responses, saving bandwidth.  Backward-compatible:
old servers ignore the param, client-side remove_fields still strips.

get_page_content raises PageIndexAPIError instead of TypeError when the
document is still processing (get_ocr returns result: null).
2026-08-11 21:52:03 +08:00
Ray b723c9f0a7 ci: run the test suite on every push and pull request
Publishing was the only automation touching code: a tag builds and ships
to PyPI without ever running a test, and pull requests get no checks at
all. This runs pytest on a small matrix — Python 3.10 (the floor) and
3.13, each with and without the agent frameworks installed, so the
lazy-import contract (the package must work with neither framework
present) is enforced rather than assumed.
2026-08-10 16:32:56 +08:00
Ray d5c4e62c20 Exclude common words from roman numeral conversion (#387)
Update PageIndex Flash
2026-08-04 19:15:44 +08:00
Ray 933df35b2f Add a Flash summary flag (#386)
* Update PageIndex Flash

* Add a summary flag
2026-08-04 18:51:27 +08:00
Ray 502b763876 Update README (#385) 2026-08-04 18:38:29 +08:00
Ray 3c4c7e8353 Use embedded PDF bookmarks in Flash trees (#384)
* Update PageIndex Flash

* Add embedded ToC CLI flag and README note
2026-08-04 18:35:23 +08:00
Ray fb3f6e441a Fail fast on rejected keys and unknown models (#381) 2026-08-03 05:19:17 +08:00
Ray 03cf0fa0ef Sync PageIndex Flash from private branch (#380) 2026-08-03 04:28:14 +08:00
Ray d409566464 Sync PageIndex Flash from private branch (#379) 2026-08-03 03:00:59 +08:00
Ray 1b2fdadf51 Keep merge inside --optimize (#377) 2026-08-02 21:52:32 +08:00
Ray 2a29ac5aa0 Merge pull request #375 from VectifyAI/fix/default-merge
Refine the Flash tree with merge
2026-08-02 20:51:23 +08:00
Ray 9e789babf3 Add benchmark to Flash README (#376) 2026-08-02 19:10:48 +08:00
Ray e17dac2053 Strip internal marker from optimize output 2026-08-02 18:23:46 +08:00
Ray 4a116daa2f Replace thinning with merge 2026-08-02 17:56:54 +08:00
Ray 3f33a53b50 Add tree optimization and node summaries for Flash 2026-08-01 17:56:04 +08:00
Ray afb3100232 Merge pull request #374 from VectifyAI/fix/flash-requirements
Add Flash dependencies to requirements
2026-07-31 20:24:10 +08:00
Ray 6e248f8f05 Add Flash dependencies to requirements 2026-07-31 20:22:46 +08:00
Ray d178bc557d Merge pull request #372 from VectifyAI/openai-sdk-fast-path
Use OpenAI SDK directly for OpenAI models
2026-07-31 02:38:45 +08:00
Ray 548d3201b7 Use OpenAI SDK directly for OpenAI models 2026-07-31 02:32:58 +08:00
Ray 66d5b9ac9b Merge pull request #371 from VectifyAI/fix-flash-readme
Add --flash CLI flag and update README
2026-07-30 20:51:41 +08:00
Ray faf3c57fc8 Add --flash flag to CLI 2026-07-30 20:49:16 +08:00
Ray a77301e5c2 Merge pull request #370 from VectifyAI/fix-flash-readme
Update Flash README links
2026-07-30 20:42:40 +08:00
Ray 165fcea3db Update Flash links to main branch 2026-07-30 20:42:19 +08:00
Ray c9a01fad13 Merge pull request #369 from VectifyAI/pageindex-flash
Add PageIndex Flash
2026-07-30 20:36:33 +08:00
Ray fffc2c5512 Add PageIndex Flash 2026-07-30 20:35:19 +08:00
Ray 0e4b68c92d Merge pull request #367 from VectifyAI/add-summary-model-config
Add summary_model config
2026-07-30 20:21:54 +08:00
Ray acfcde2c58 Remove summary_model from page_index() signature 2026-07-30 20:20:43 +08:00
Ray 820e8bf103 Use summary_model for doc description, expose in page_index() 2026-07-30 20:19:32 +08:00