mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-10-01 23:35:08 +08:00
main
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6d23caf416 |
Unify the document tree across local and cloud (#541)
get_tree returns one node shape in local and cloud mode: {title, node_id,
start_index, end_index, summary, text, nodes}. page_index and
prefix_summary no longer appear. The SDK only renames fields on the way
out, so a document indexed before keeps its own ranges and summaries.
New local indexes, standard and flash:
- A parent whose first child starts on a later page gets a first child
"<parent title> (intro)" that holds those pages.
- A parent's range covers its whole subtree, and its summary is written
from its children's summaries. Standard mode now summarizes with
summarize_tree, as flash does.
- A node the model leaves unsummarized falls back to its subsection titles
or its own text.
- The standard large-node split acts on leaves only.
A node's text is its own pages. A parent's runs onto the page its first
child starts on, and is empty when its intro holds those pages.
|
||
|
|
f0d67c1224 |
test: summary concurrency test waits for the expected overlap (#547)
The mock held each call for a fixed 20 ms and expected all three to overlap inside it. On a slow CI runner the three summaries start further apart than that, and the control run read (3, 2) instead of (3, 3). Each call still holds 20 ms, so a lane with no cap still lets the others in. It then holds until its lane reaches the expected peak, with a 5 s deadline. An ignored cap and too little overlap both still fail. |
||
|
|
60c8ff8a6f |
feat(sdk): node navigation helpers get_node, get_node_parent, get_node_path, get_node_map (#542)
They take the tree, not a doc_id: fetch it once (get_document_structure) and navigate locally, instead of one request per step. get_node / get_node_parent / get_node_path share one depth-first walk, O(n) per call. They return the tree's own nodes rather than copies (unlike get_nodes / get_leaf_nodes), so node["nodes"] keeps working. get_node_map is create_node_mapping plus the input check; create_node_mapping stays as the 0.2.8 surface. Wrong input raises TypeError naming get_document_structure(doc_id) / get_tree(doc_id)['result'] instead of reading as "not found": the whole get_tree() response, a None tree, or a non-string id. create_node_mapping (so get_node_map too) walks past a node whose nodes is None instead of raising TypeError, as get_node already did. |
||
|
|
f279431eb4 |
Restructure repo layout (#539)
- Remove docs/naming-rules.md - Move examples/documents/results/ to examples/results/ |
||
|
|
d2693d8079 |
ci(publish): gate releases on the live cloud tests again (#538)
Reverts #537. |
||
|
|
619cbd89f6 |
ci(publish): run the release gate without the live cloud tests (#537)
The CI PageIndex account is over the 1,000 active-page free tier, so every test_live_* call returns HTTP 403 and blocks the release. Without PAGEINDEX_API_KEY the live tests skip; the offline suite still gates the publish. |
||
|
|
037a7dbacf |
perf: summaries run deepest-first and start while expand is still deciding (#432)
Flash indexing spends most of its wall time in summaries, and until now that stage waited for expand to finish and then ran its calls in whatever order the tree recursion produced. This branch makes the summary stage run deepest node first and start while expand is still deciding, so the LLM channels never sit idle waiting on the expand chain. **What changes** - `_PriorityGate`: the summary semaphore admits the queued call with the most work still above it (depth = calls left on the node's path to the root, its own included), FIFO within a depth. Cancellation-safe like `asyncio.Semaphore`. - Tasks are created deepest node first, so the first admissions are the deep leaves rather than whichever shallow leaves the recursion reached first. - `summarize_tree` becomes a thin wrapper over `SummaryScheduler`: `mark_final(nodes)` says those nodes will not gain, lose or swap children and starts their subtrees; `finish()` awaits the roots. Same task order, gate and error semantics as before. - `optimize(on_final=...)` reports which nodes are final as it goes: after each round's merges, at each expand candidate's decision (together with what it grew), and for the whole tree at the end. A node is final when it is collapsed under the trigger, collapsed and already judged by expand, or has children — the cost merge cannot fire on a surviving node after the first round (see the commit message for the argument). - Same-page fusion moves to where duplicates arise (right after a collapsing merge, right after expand attaches children) instead of the next round's start, so no node waits a round for it. The nine corpus PDFs produce byte-identical merge-only trees; SpaceX just stops after two rounds instead of a third that did nothing. - `page_index_flash` runs expand and summaries on one event loop when both are on; every other combination keeps the old path. **Measured** (same hour, end to end via `submit_document`) | | before | after | |---|---|---| | fed-2023 (222 p) | 97.9 s | 72.6 s | | PRML (758 p) | 174.3 s | 136.8 s | Summary-stage only (fed, 182 calls, 64 wide): FIFO 58–62 s → gate 50–57 s → gate + deepest-first 45 s. Same calls, same prompts; outputs are order-independent. Peak in flight is now the expand cap plus the summary cap (32 + 64). **Tests** cover the ordering, cancellation, scheduler, final-node reporting, immediate-fusion and one-loop overlap cases, and every knob's path from the client and the CLI to the model calls. **Summary prompt and indexing knobs** The summary prompts no longer ask for the `points` list that `parse_summary` discarded, and cap the summary at `summary_max_words` (default 150). Measured on gpt-5.6-luna, mirror A/B, summary stage only: per-call latency 9.7 → 5.3 s (−45%), fed-2023 47.5 → 30.7 s (−35%), PRML 71.1 → 38.1 s (−46%), output tokens −65%. Summaries come out ~1160 chars instead of ~670 and carry the specifics that used to sit in the discarded list; a blinded pairwise judge (claude-sonnet-5, source in view) prefers them 21-1-0 over the old ones. Deleting the list without a cap is not enough: the model then pours it into the summary (3× longer) and parents slow down more than the leaves gain. Four indexing knobs are settable from the SDK (flat arguments or the `index=` slot) and the CLI: `summary_max_words`, `summary_concurrency`, `use_embedded_toc`, `optimize` (`"full"` / `"merge"` / `"off"`). `summary_concurrency` bounds both lanes: expand's gate becomes min(32, the cap), so one knob lowers the whole indexing lane on a tight quota (the lanes overlap, so up to cap + min(32, cap) calls run at once). Defaults are unchanged. The two summary knobs are flash-only: `submit_document(mode="standard")` refuses them rather than index without the cap, as the CLI already does. Both must be positive integers, checked before the PDF is opened; a direct `page_index_flash` call that passed `0` (read as the default until now) or a whole-number float such as `8.0` now raises `ValueError`. |
||
|
|
8eff773368 |
Anthropic lanes: send the model as written; the prefix picks the route (#443)
- chat(protocol="messages"), anthropic_runner_config() and claude_agent_config() honor chat_model: a model you set carries over; the SDK's stock default is never sent. - Model names are sent as written; only the routing prefix is read. bedrock/, vertex_ai/ and azure_ai/ select AnthropicBedrock, AnthropicVertex and AnthropicFoundry (and the matching CLAUDE_CODE_USE_* switch for Claude Code); anything else goes to Anthropic directly. - Bedrock: InvokeModel rejects top-level cache_control for Opus 4.6 and earlier, so the Messages lane moves an explicit breakpoint onto each turn's newest tool result; anthropic_runner_config() omits the key on that route. - The max_tokens default is 8192, or budget_tokens + 8192 with thinking; no LiteLLM ceiling lookup, no claude-3 list. - The anthropic extra requires >=0.122.0 (the Bedrock/Vertex tool runner). |
||
|
|
9c4c3ff2dd | test(sdk): cover list_documents regression boundaries | ||
|
|
ed06ec69ed |
fix(sdk): list_documents review followups
- Forward recursive value instead of hardcoding True (#3) - Raise limit cap from 100 to 10000 to match server (#4) - Accept folder_id='root'/'' in local mode (#6) - Fix wire test to assert presence not identity (#6, #9) - Delete restating comments (#15) Addresses PR #522 review findings 2-7. |
||
|
|
91238b33b6 | README banner: drop fixed height so it scales on PyPI (#528) | ||
|
|
71714e86d9 | get_document_id accepts a path: strip the folder prefix since names are unique | ||
|
|
6674d6c4b4 |
resolve_citations: paired cite tags leave no closing tag behind (#525)
<cite doc= page=>quoted</cite> and <cite doc= page=></cite> are tag shapes the parser accepts and the chat renderer matches, but the rewrite replaced only the opening tag, so the display text kept a raw </cite>: "[[1]](#pageindex-citation-01)quoted</cite>". The chat UI hides that because its HTML pipeline drops a stray end tag; this text goes to any host. _CITE_TAG_RE now takes an optional "text</cite>" tail and the link keeps the text. A closing tag is consumed only together with the citation it closes, so a tag left unresolved, or a </cite> in quoted HTML, stays as written. The parser reads group 1 as before: 0 diffs on 30,000 random tag inputs, and 1 MB bodies scan in about 15 ms. The Returns section said every block-level citation carries bbox, block_type and text. get_citations() adds them only when the block could be read, so a local document or a 403/404 block has none, and c['bbox'] raised KeyError for a reader of this docstring alone. The highlight test drew on a 1000 px square at scale 1000 and probed one pixel on the corner of the transposed rectangle, so an x/y swap, an ignored scale and a whole-page fill all passed. It now draws on a 500x2000 page and checks one pixel inside the region and one outside; all three mutants fail it. |
||
|
|
0445ee80bd |
feat(sdk): align upload and folder names with shared naming rules (#524)
* feat(sdk): align upload and folder names with shared naming rules * fix(sdk): normalize filenames before multipart upload |
||
|
|
8e386ff626 |
Citation display helpers: numbered links, region highlight, image URLs, folder paths
BREAKING: resolve_citations(), released in 0.2.17, is renamed
get_citations(); same arguments, same list. resolve_citations() now returns
{'answer', 'citations'}: the answer with each citation tag replaced by
[[i]](#pageindex-citation-0i), and the get_citations() entries led by
'anchor' and 'index'. Numbers follow the tags as written, so a repeated
citation reuses its number and a block cited under the wrong page still
gets its link while its entry carries the block's real page. One shared
_citation_key() parses a matched tag for both the parser and the rewrite.
highlight_region(image, bbox, scale=1000) draws the cited region on a page
image (PIL.Image or bytes in, PIL.Image out). Pillow becomes a dependency.
get_page_image(doc_id, page) and get_document_image(doc_id, img_id) return
short-lived URLs. Cloud-only. They need the compute routes that return
presigned URLs; against an older server they raise PageIndexAPIError.
get_document_path, get_folder_path and get_folder_id translate between ids
and readable paths. Paths are built from list_folders() parent links,
because the server's `path` field exists only on /docs listing rows. A path
shared by two folders raises instead of picking one.
|
||
|
|
2a7363692f |
list_documents: pass through the API's recursive flag
Listing a folder returned only its direct contents, with no way to reach the documents in its sub-folders; the endpoint has taken recursive all along, the SDK just never sent it. Listed documents also carry a folder path, served by the cloud. Local mode accepts the flag and ignores it — there are no folders to descend into — and reports the same keys a cloud listing has, path among them, always None. |
||
|
|
bd72efe09d |
Tighten instruction parity alarm (#516)
- Markers: 'folder' → 'folder_id', 'get_folder_structure' → 'get_folder_structure(',
'recursive' → 'recursive='; the browse-folder routing line moves to exact-match
- Section order: shared headers must appear in the same relative order
- None guard: _assert_instructions_local_parity(None) → readable failure
- Format enumeration: live test reads prompts/list and asserts the server's
cited_answer format set matches LOCAL_CITATION_PROMPTS
|
||
|
|
a64bc4f7fd |
ci: pass PAGEINDEX_API_KEY to publish.yml test gate (#514)
#512 added the secret to tests.yml but missed publish.yml — the release gate still skips every live parity test. |
||
|
|
c840b69b25 |
Sync browse_documents query: "Required when sort=relevance" (#515)
Sync browse_documents query description with server (#486) pageindex-chat #486 fixed 'Only used' → 'Required when sort=relevance'. Sync the local contract and snapshot to match. |
||
|
|
18eb5c9b3c |
Review fixes: unify format validation, pin key set, self-check exception list
- Hoist format guard above the api_key branch so both lanes reject
unknown formats (cloud was silently returning the markdown prompt)
- Docstring: 'any other value raises' (the server does not reject)
- assert set(LOCAL_CITATION_PROMPTS) == {markdown, cite} pins keys
- assert set(_LOCAL_ONLY_LINES) <= frozen catches stale entries
|
||
|
|
fd3c12ea04 | Detect deleted shared agent instructions in parity checks | ||
|
|
d059a8b378 |
Sync browse_documents query description with server
The server (PR #484) changed the query parameter description from "must be omitted when sort='time'" to "Only used when sort='relevance'". Update the local contract and snapshot to match. |
||
|
|
bf0645cc28 |
Drop the footnote citation format; alarm on instruction drift
pageindex-chat removed `footnote` from the cited_answer prompt
(2fc3d298), leaving the SDK advertising a format the cloud now rejects:
`citation_prompt("footnote")` raised a raw ProtocolError against the
cloud while local mode happily returned the frozen copy — the same call,
two behaviours.
Also adds the drift alarm the instructions never had. The tool contract
and the citation prompts each have a live parity test; AGENT_INSTRUCTIONS
only had a non-empty check, so a cloud edit to shared guidance could
leave local agents on stale instructions unnoticed. The new test lets a
cloud line through only when it names a tool or parameter local mode does
not have.
Verified against the live server: 8 live tests pass, 167 offline.
|
||
|
|
d44799f33f |
ci: enable live parity tests in CI (#512)
ci: pass PAGEINDEX_API_KEY to pytest so live parity tests run The 6 live tests (drift alarms that verify the local contract against api.pageindex.ai) are gated on PAGEINDEX_API_KEY. Without the secret they silently skip, making the CI green but the guards ineffective. The secret must still be added in repo Settings → Secrets → Actions. Fork PRs never receive it (GitHub default for `pull_request` trigger). |
||
|
|
8fede179c6 |
Parity test: Sources line no longer served by the cloud (#508)
Chat PR #480 removed the Sources trailer from the MCP server's cited_answer prompt, so the only dropped line is get_document_image(). |
||
|
|
ac0eed7444 |
v0.2.18: flash improvements, prompt alignment, get_document_id (#507)
* Flash: drop the size and landscape bail, fall back to page nodes extract_toc refused any PDF under 300 text weight, or 200 on its densest page, or with mostly-landscape pages, and returned an empty outline. Both rules threw away documents the detector handles: three to five pages of heading plus a line yield every heading, and an ordinary slide deck yields one node per slide. Only the no-alphabetic-script condition stays, and its result now says so with toc_source="unreadable" instead of borrowing "detected"; toc_source is set on every result (detected, bookmarks, pages, unreadable). When detection yields nothing for a document that has text, the pages are the tree: page_index_flash emits one node per text-bearing page, titled by its first block and labelled toc_source="pages". A flat tree larger than FLAT_TREE_MAX_NODES is returned without the optimize and summary passes, since the managed pipelines refuse it: flash_rejection_reason() gives the local client and the CLI one policy, pointing at standard mode for an oversized flat tree, and for an unreadable PDF naming the missing text (standard mode stays the hint there too, since a PDF of nothing but numbers lands in the same bucket and the model can still read it). The nine example PDFs extract byte-identically before and after: the removed rules never fired on ordinary documents. * Flash: layout decides, never script; the page fallback covers every page extract_toc refused a PDF whose dominant script fell in Scholar's "other" family, which is not a script but a bucket: Arabic, Hebrew, Persian, Urdu, Devanagari, Bengali, Tamil, Thai, Khmer, Georgian, Armenian, Amharic, and any document with no letters at all, each reported as "no alphabetic text". Three more rules keyed on script inside the estimator: an unnumbered heading in a script other than the body's was dropped, so a Chinese report lost its English section titles; a kana-majority Japanese document had every detected heading discarded; a mostly-landscape document picked its title from page one without the body-paragraph check the portrait path applies, so a slide deck's title became slide one's body text. These are Scholar's precision-over-recall scope limits for an index of Latin and CJK papers; on PageIndex's default local mode they were silent refusals and silent losses. They are deleted here, in this repo's copy of the port, together with the Cyrillic-only density threshold; the private scholar/ tree stays a faithful port and the new tests guard the fork. With the rules gone the same layout yields the same three headings in Japanese, Hindi, Arabic, Hebrew, Thai and mixed Chinese/English, and English is unchanged. Language now only decides which cues are available: case, keyword tables, numbering styles. The page fallback emitted a node only for pages carrying text, each spanning one page, so a partly-OCR'd 20-page PDF indexed as six islands with fourteen pages in no node's range and unreachable by retrieval, accepted without an error. Titles came from the page's first raw line, on a real document the running header or the page number. Every page is now a node titled "Page N", FLAT_TREE_MAX_NODES bounds document pages, and toc_source="unreadable" means exactly that no page carries text. The refusal says what was measured, no text layer, run OCR, instead of "no alphabetic text ... try standard mode", which asserted a cause never checked and pointed at a mode fed the same bytes. The non-Latin fixtures are PyMuPDF-generated with open-licensed font subsets embedded (FiraGO, Droid Sans Fallback); tests/data/flash/make_fixtures.py regenerates them byte-identically. * Flash parser: a multi-code-point glyph no longer crashes the RTL sign A ToUnicode value can be several code points: a Devanagari conjunct, a Thai cluster, an Arabic ligature. _rtl_sign passed the whole value to unicodedata.bidirectional, which takes exactly one character, so every real Hindi and Thai PDF raised TypeError in the char merge, before any rule ran. The first code point now decides the direction, as _reverse_if_rtl already does for the same values; the empty string is LTR. * Flash: title scoring stops looking at script; the docs describe page nodes as emitted The title scorer halved any candidate whose dominant script differed from the document's, the last place flash changed a decision on script. Without the factor the nine example PDFs and the four fixture PDFs extract byte-identically. get_leaf_nodes reads `nodes` with .get like its siblings, so a flat page tree, whose nodes carry no `nodes` key, walks instead of raising KeyError. The README output block and the page_index_flash docstring now say what a node actually carries: node_id always, `nodes` only with children, `summary` only when summaries ran, and a flat tree past FLAT_TREE_MAX_NODES pages coming back unsummarized and unoptimized. The local client's page-fallback test feeds that real node shape. * Flash: a hierarchy that starts after page 1 gets a Preface node, as in standard mode Retrieval reaches a page only through a node range. When the first heading or the first bookmark sits on a later page, everything before it was in no node: a memo's first section once its heading became the document title, a title slide's body, a report's cover, contents and letter to shareholders, the cover pages of a bookmark tree. Standard mode has always inserted a Preface node for exactly this (utils.add_preface_if_needed); flash now does the same, next to the page fallback and before the optimize and summary passes, and renumbers the node ids. Of the nine example PDFs, the two Federal Reserve reports and Four Lectures gain the node (pages 1-4, 1-2 and 1); the other six and every extract_toc result are unchanged. The four fixtures gain a Preface over their title page. * Align local prompt copies with the chat MCP server The frozen citation prompt carried an invented "Sources" footer and the agent instructions had a pagination tutorial chat never shipped; the persistence protocol was missing the rephrase-and-retry step chat has had since April. * get_document_id(name): look up a document ID by its display name One API call via the new server-side name filter on /docs, so no client-side pagination. Cloud and local both covered. * Fix citation parity test for the Sources line removal The test asserted the frozen copy matched the cloud exactly minus get_document_image(); the Sources line is now a second deliberate omission in the cite format. |
||
|
|
bfbd4b305c |
Flash: layout decides, never script; the page fallback covers every page (#502)
Flash returned an empty structure, and `submit_document(mode="flash")` and the CLI a hard error, for any PDF under 300 text weight, under 200 on its densest page, or with mostly-landscape pages. Both rules threw away documents the detector handles. Four more rules keyed on the document's script: the "other" script family (Arabic, Hebrew, Persian, Urdu, Devanagari, Bengali, Tamil, Thai, Khmer, Georgian, Armenian, Amharic, and numbers-only text) was refused as "no alphabetic text"; an unnumbered heading in a script other than the body's was dropped, so a Chinese report lost its English section titles; a kana-majority Japanese document had every detected heading discarded; a mostly-landscape document picked its title from page one without the body-paragraph check, so a slide deck's title became slide one's body text. These are Scholar's scope limits for an index of Latin and CJK papers; on PageIndex's default local mode they were silent refusals and silent losses. **What changes** - Layout decides, never script. The size and landscape bails, the script gate, the cross-script heading drop, the Japanese outline nullifier, the landscape title branch, the Cyrillic-only density threshold and the title scorer's cross-script penalty are deleted from this repo's copy of the port; the private `scholar/` tree stays a faithful port and the new tests guard the fork. Language now only decides which cues are available: case, keyword tables, numbering styles. - When detection finds no hierarchy, `page_index_flash` returns one node per page titled `Page N`, covering every page, labelled `toc_source="pages"`. A flat tree over `FLAT_TREE_MAX_NODES` (10) pages comes back without the optimize and summary passes and is refused by the local client and the CLI through one shared `flash_rejection_reason()`, pointing at standard mode. - Every page is in some node. A hierarchy that starts after page 1 (a memo whose first heading became the document title, a title slide, a report's cover and contents, a bookmark outline that begins on page 3) is preceded by a `Preface` node covering the pages before it, the node standard mode has always inserted for the same case; until now those pages were reachable from no node. - `toc_source="unreadable"` means exactly that no page carries text; the refusal says so and points at OCR, not at standard mode, which would receive the same bytes. - The character-level parser no longer raises on a glyph whose ToUnicode value is several code points (a Devanagari conjunct, a Thai cluster, an Arabic ligature); real Hindi and Thai PDFs used to fail with a `TypeError` before any rule ran. - `toc_source` is present on every result: `detected`, `bookmarks`, `hybrid`, `pages`, `unreadable`. The README and the `page_index_flash` docstring list them, and describe a node as emitted: `node_id` on every node, `nodes` only on entries with children, `summary` only when summaries ran. - `get_leaf_nodes` walks a flat page tree instead of raising `KeyError` on a node without a `nodes` key; it was the one tree helper reading the key unguarded. **Behaviour change** Small documents, slide decks, and Japanese, Arabic, Hebrew, Indic, Thai and mixed-script documents that used to fail flash indexing or lose headings now index; with the rules gone the same layout yields the same headings in every one of those scripts, and English is unchanged. A garbage text layer that still has layout structure now indexes as a garbage-titled tree instead of being refused. A Chinese-body report whose cover sets an English title over a Chinese subtitle now picks its title by layout; the deleted penalty could hand `doc_title` to a body paragraph. `extract_toc` yields the same nine example trees, node for node, before and after; `page_index_flash` adds the `Preface` node to the three whose hierarchy starts late (the two Federal Reserve reports, pages 1-4 and 1-2, and Four Lectures, page 1), the node standard mode already gives them, and leaves the other six identical. **Tests** Fixtures for Japanese, Chinese with English headings, Hindi and Arabic under `tests/data/flash/`, PyMuPDF-generated with open-licensed font subsets embedded; `make_fixtures.py` regenerates them byte-identically. Green on all three CI legs locally (with and without agent frameworks, pypdfium2 4 and 5). |
||
|
|
a30a805151 |
Block citations resolve to bounding boxes: get_block() and resolve_citations() (#503)
* Block citations resolve to bounding boxes: get_block() and resolve_citations()
resolve_citations(answer, doc_ids=None) reads the <cite doc= page= block=/>
and <doc=…;page=…;block=…> tags of a cited answer and returns one entry per
distinct citation with the document id and, for block-level citations on
cloud documents, the block's page, bbox, type and text from the new
GET /doc/{doc_id}/block/{block_id}/ route, wrapped as get_block(). Local
mode stays page-level; a name shared by two documents raises instead of
guessing; a block the document lacks keeps its entry without a bbox.
* resolve_citations: paired quotes, tag-edged names, tolerant block lookups
_CITE_ATTR_RE pairs its quotes, so an apostrophe inside a quoted name
(Moody's Outlook.pdf) stays in the name instead of ending it.
_OLD_CITATION_RE stops a name at < or >, so a <doc=name> tag without its
own ';' no longer swallows the text up to the next tag's ';' and that
citation with it.
The block lookup keeps its bbox-less entry on a status-less raise (local
mode, where pages have no blocks) and on 403, as doc_targeting_block
does; a transport failure still propagates with its status.
The library listing is _all_documents(), which tolerates an absent
'total' and ends on the empty page; ids are deduplicated before the
collision check, so a document re-served after an upload shifted the
window is one document, not a collision with itself.
doc_ids=[] fails loud like chat(doc_id=[]) instead of resolving nothing.
Docstrings: the bbox coordinate space on both get_block(); the
shared-document promise dropped, since the metadata route is owner-only.
* Citation tag scan is linear
_CITE_TAG_RE's second alternative (the paired form) was unreachable: the
first alternative matches the opening tag whenever a '>' follows, and
without one neither alternative can match. Dropping it, stopping the tag
body at '<' and letting the group absorb the whitespace makes an
unterminated '<cite ' scan linear: '<cite ' plus 4000 spaces went from
2 minutes to 0.1 ms, 99 KB of unterminated tags from 27 minutes to
about a millisecond. Output is identical on every shape tried, the
paired form included.
* Finish the tag-edge and listing-entry sweep
Two holes of the kinds already closed, left open one line and one loop
away from their siblings.
The old format's block group still crossed tag boundaries the way its
name group used to: "<doc=a.pdf;page=3;block=p3_text_5 <cite doc=..." put
the whole following citation inside block_id and, since a block lookup
that 404s is now kept without a bbox, lost it silently. It stops at the
tag's edge like the name beside it; a block id has never contained < or
>, the server writes p{page}_{type}_{seq}.
The library listing reads entries with .get() like the rest of the
codebase does: _all_documents() tolerates a listing whose shape varies,
so reading its result with raw subscripts put the KeyError back one
level down. An entry without a name or an id cannot answer a citation
either way, so it is skipped.
* resolve_citations: block ids trimmed, linear attribute scan, doc_id= like chat()
add() trims both values itself: the name, which each caller used to
trim, and the block id, which neither did. "block=p3_text_5 >" and
block=" p3_text_5 " now look up p3_text_5. A padded id 404s, and a 404
is kept without a bbox, so the padding lost the bbox silently and handed
the padded id back to the caller.
_CITE_ATTR_RE anchors the attribute name at a word boundary. Without
\b the greedy (\w+) restarted at every character of a long word inside
a tag body: '<cite ' + 64 KB of x + '>' cost 19 s on the caller's
thread. Anchored, 64 KB is 1.6 ms and 1 MB 27 ms. Output is identical:
the leftmost match of (\w+)= always begins at a word boundary (0 diffs
on 30,000 random inputs). The linear-time test now reaches this
expression; its inputs had no '>' and never did.
The scope parameter is doc_id, the name chat(), chat_completions() and
document_context() give the same str-or-list argument, and the one the
docstring already used to describe it. The integration builders keep
their doc_ids.
|
||
|
|
c75800506c |
Three docstring truths the 0.2.16 merges outran (#501)
- chat(instructions=) called the managed system prompt the carrier of "the tool guidance and the document context". Targeting moved to the first user message, so it carries the tool guidance alone. - chat(citations=) said the managed endpoint's resolved citations come back "only from chat_completions". They ride the response envelope, which protocol="chat_completions" returns whole as well; what cannot reach them is the answer lane, which returns the answer string. - Rejoins a line as_claude_mcp()'s paragraph lost to a merge. Each clause was written on a branch cut before the change that voided it, and survived the merge because neither side conflicted. |
||
|
|
fc2dc41c16 |
Client instructions: a standing persona for every answer surface (#497)
* Client instructions: a standing persona for every answer surface PageIndexClient(instructions=...) sets standing guidance for the answering agent — persona, language, format — appended after the managed system prompt wherever an answer is produced: chat() and chat_completions() on both engines, the Responses and Messages protocol lanes, agent_instructions() and the three *_agent_config() bundles. It is a client-level argument like mode=, not a chat-side spelling: it combines with any chat= and never selects own-model chat on its own. Blank configures nothing, as chat(instructions="") does. chat(instructions=) adds to it per call; the prompt order is managed base, client, call, history system rows. One insertion point serves every own-model surface (_base_instructions). The managed cloud chat takes exactly one system message, first: the client's instructions, the call's, and the history's system rows now fold into it, in that order — so chat(instructions=) works on a managed client (refused since #460, although the endpoint has accepted custom instructions since Sep 1), and a system row anywhere in the history no longer 400s. The answer lane's messages contract is the same on both engines: text history only — the endpoint refuses tool rows and structured content itself; the SDK says so first, with the protocol-lane pointer. Live-verified against the production managed chat and the live MCP instructions. Claude-Session: https://claude.ai/code/session_01J8fbpdM5pz2JLjiNY11usy * Managed fold: always send the canonical history The managed payload is now the same shape every time — one leading system row when there is any system text, then the role/content history — instead of forwarding the caller's list untouched when nothing folded. That branch let a blank system row sit mid-history and reach the endpoint's "only one system message, first" refusal; blank text now configures nothing, like a blank instructions= does. Claude-Session: https://claude.ai/code/session_01J8fbpdM5pz2JLjiNY11usy * Managed fold: lift system rows only, forward the rest verbatim The managed chat_completions lane ran the whole payload through the own-model validator: tool rows, structured content and tuples were refused before the wire, and every field beyond role/content was stripped (tool_calls, name, the endpoint's own citations). The SDK is a thin skin over the cloud. It folds the client's instructions and the history's system/developer rows into the one leading system row the endpoint takes, and sends everything else as given; the endpoint decides what it accepts. Also: - _managed_instructions drops blank system texts, as the managed fold and _anthropic_system already do, so both engines build the same prompt for the same input (and share one prompt-cache key). - An end-to-end test drives chat() on the own-model lane and asserts the persona and the call's system text reach the model; the white-box builder tests alone stayed green with the persona removed from the answer lane. - as_claude_mcp(): the MCP-instructions channel carries the tool guidance only; the client's instructions ride system_prompt. - Comments that restated code or an assertion removed; the chat= combination test asserts chat_model too. Blank instructions configure nothing at the constructor, as chat(instructions="") does. Claude-Session: https://claude.ai/code/session_01TgdXZx63aMwrpaKFTQFFKx * Chat history: normalize the container, validate only what the SDK reads Both lanes now take any iterable of message dicts. chat()'s instructions prepend expanded only lists, so a tuple or generator on the managed lane silently dropped the call's instructions once _require_own_chat no longer refused it; own-model rejected the same shapes outright. list() once at each reader instead. None and other non-iterables fail with Python's own TypeError, as in the OpenAI SDK. _system_text refuses a system row carrying non-text parts instead of keeping the text parts and dropping the rest: the SDK folds that row into its own system text, so it is the reader and must say what it could not read. The endpoint stays the authority on every row the fold leaves in place. Restore the messages docstring sentence 108875b rewrote: tool-role turns are rejected on both engines (the endpoint 400s them), so the managed lane does not take them "verbatim"; rename the test that carried that claim. Cover the blank-text filter, which no test guarded, and drop a truncated comment. Claude-Session: https://claude.ai/code/session_01UBpLu7TJUvrvLvFacs8WYT * fix: ctor type error names the answer, not a lane managed clients cannot use A plain managed client is the caller most likely to pass Messages-style blocks as instructions=, and the old message sent it to chat(protocol="messages"), which that same client refuses. Say "pass text", and make the block pointer the exact working call: model= is required on that lane, and it needs a chat_model= client. Claude-Session: https://claude.ai/code/session_018od4o3YBqrny2vzGhrX4q5 * Rebase onto the merged lanes: two pins the lift and the targeting move void The chat_completions protocol lane's guard test still pinned the managed refusal of instructions=, which this branch lifts on purpose; its own managed-fold tests cover what the lane now does. And _anthropic_system lost its doc_id argument when targeting moved to the first user message, so the prompt-order assertion calls it with what it takes. |
||
|
|
a3364c8a1b |
Document targeting moves to the first user message; folders join chat() (#495)
* Document targeting is conversation content: document_context() replaces doc_id on the agent surfaces
The doc-targeting block ("The user has specified document: ...") was
appended to the system prompt by agent_instructions() and the three
*_agent_config bundles, and placed as its own system block by the Messages
chat lane. A per-request target in the system prompt breaks the cached
prefix, sticks across turns, and on cloud splices SDK text onto the
server-served instructions. The cloud's managed chat puts the same block at
the head of the conversation, so every surface now does the same.
- chat(): all three lanes prepend the block as the first user message
(the Messages lane moved off its system block).
- New client.document_context(doc_id): the block text for callers who own
the conversation (the framework routes), to lead their first message.
- BREAKING: doc_id removed from agent_instructions(), openai_agent_config(),
anthropic_runner_config(), claude_agent_config(), agent_tools(),
as_openai_tools(), as_anthropic_tools() and as_claude_mcp(). The
tool-layer allowlist was local-only and raised on cloud; it stays
internal to local chat(doc_id=).
- The block is one get_document per id: names are unique per library
(uploads suffix a taken name), so the shadow check and its listing sweep
went. The cloud endpoint carries metadata, so local get_document() gained
the key for parity. Rendering matches the cloud's: one document is an
object, several a list.
Live-verified on a local store and the cloud library across chat() (answer
lane, chat_completions, responses, messages) and the OpenAI Agents,
Anthropic tool_runner and Claude Agent SDK routes.
Claude-Session: https://claude.ai/code/session_01VxguPoTv2BmzS9d3erzTrS
* Agent surfaces: keyword-only tails so a stale positional doc_id raises
doc_id sat first in the positional list on agent_instructions,
openai_agent_config and claude_agent_config, and second on
anthropic_runner_config and as_claude_mcp. With it removed, a v0.2.15 call
such as openai_agent_config("pi-a") no longer failed: the id landed on
include_management (truthy, so the management tools came along and
targeting silently vanished) or on server_name. A bare * after the
surviving leading positional turns those calls into an immediate
TypeError; every in-repo caller already passes keywords.
Claude-Session: https://claude.ai/code/session_017FumozBm2xbT2SG6WBxjMe
* Fix the delegate build_claude_mcp missed by the keyword-only tails
|
||
|
|
d1a4477c29 |
chat(protocol="chat_completions") and the compat note on chat_completions() (#493)
* chat(protocol="chat_completions") and the compat note on chat_completions() chat() gains the third protocol value: the answer lane's own engine with its Chat Completions envelope kept (chunk dicts when streaming, instructions as a leading system row). It is the one protocol the managed cloud chat serves, so that lane opens without a chat model; the own-model knobs still refuse there. chat_completions() stays, unchanged in signature, with a docstring that marks it as kept for existing code and points new code at chat(). Error strings that steered callers to it now name the protocol lane; the cookbook's two cells use chat(). The managed endpoint takes extra_body as its own request fields (temperature, enable_citations), merged last under the skeleton refusal, so new code reaches them without the old door. Claude-Session: https://claude.ai/code/session_015b7YGZ8LvYp2Q3oGNN8Lfc * Review fixes: type and document the chat_completions protocol value The protocol overloads name "chat_completions" so the literal narrows to the envelope dict and the chunk-dict iterator; the protocol arg doc lists it and notes the managed chat serves it; the show_process texts no longer promise a transcript the Chat Completions lane has none of. Claude-Session: https://claude.ai/code/session_015b7YGZ8LvYp2Q3oGNN8Lfc * Review fixes: name the lanes in error text, finish the chat() steer sweep - _split_chat_messages rejected tool rows and structured content with "use chat(protocol=...)", which is the lane the caller just used now that chat_completions is a protocol value; name responses / messages. - submit_query / get_retrieval still steered to chat_completions(), the door this branch demotes to compat; steer to chat() like the rest. - Three docstring claims narrowed to where they hold: the compat note's "everything is chat(protocol=...)" excepts the text-only stream (that is chat(stream=True, show_process=False)); system rows join the managed prompt only with your own chat model, the managed endpoint forwards them verbatim; extra_body is verbatim on Responses, Messages and the managed endpoint, while own-model chat_completions splits it like the answer lane. - Tests: drop the warnings guard around a deprecation that does not exist and the comments restating assertions; the managed-lane skeleton test now calls the lane it is named for. Claude-Session: https://claude.ai/code/session_014bdfnYbWsejWhuphPt1aSv * Fix managed chat stream overrides in extra_body * Clarify chat completions protocol parameter documentation * chat(): keyword-only after messages, type the protocol value chat(messages, doc_id, stream, ...) and chat_completions(messages, stream, doc_id, ...) disagree on positions 2 and 3, and the compat note now tells existing callers the two are interchangeable. A positional rewrite bound doc_id=True and stream="pi-1": an unscoped answer, billed, no error. From messages on, chat() takes keyword arguments only, so that rewrite is a TypeError. Nothing in the repo passed chat() a positional after messages. BREAKING: chat(messages, doc_id) and chat(messages, doc_id, stream) by position no longer bind; spell doc_id= and stream=. Also: the implementation and catch-all overload now carry the same Literal as the protocol overloads, so a misspelt protocol fails type-checking instead of only at runtime; _refuse_skeleton names the managed lane's door (a system row), since instructions= is refused there; _require_own_chat's comment no longer claims to be the gate for every protocol. Claude-Session: https://claude.ai/code/session_01DCBitejpuoNpD8cmQ9HR7j * fix: prevent managed chat extras from overriding document scope * extra_body: one gate for every lane, before any I/O The managed endpoint's refusals of stream and doc_id sat in cloud_api.chat_completions, one call site of four. The other lanes took the same keys and broke worse: an own-model OpenAI backend sent stream:false under an SSE parser and returned an empty answer, and LiteLLM raised a bare KeyError. The merge had no Mapping check either, so a list splatted into the payload as fabricated fields. _refuse_skeleton now owns all of it: a non-dict is refused, and stream / doc_id join the refused keys beside the skeleton. chat() and chat_completions() call it first, so the own-model half no longer runs doc targeting and the MCP initialize fetch before refusing. cloud_api is a plain transport again. The skeleton remedy no longer splits by client type. A system row works on every chat-shaped lane and instructions= becomes one, so the two labels were inverted for chat_completions() callers. Managed chat(instructions=) opens when #497 lands. Tests: non-dicts and the argument keys at the gate, both public doors refusing before any lane is entered, a foreign managed key pinned as forwarded. The keyword-only test builds its client outside pytest.raises and matches the positional message. Claude-Session: https://claude.ai/code/session_01RTqXsk6Y3iGn9iXZHrzpv6 |
||
|
|
5f7a39e175 |
Tool-path rate limits: retry at the bridge, then fail the run fast (#492)
* Tool-path rate limits: retry at the bridge, then fail the run fast A PageIndex cloud 429 (or 5xx) on a tool call used to reach the model as an INTERNAL_ERROR envelope saying "try again": the model re-called once with no wait, then wrote the failure into its answer, and chat() returned normally with no status anywhere. The same 429 before the loop (the doc_id targeting lookup) already propagated raw. - McpBridge mounts a urllib3 Retry: 429/502/503 and connection failures, three attempts, 0/2/4 s apart or as Retry-After says; read timeouts are never replayed (240 s each, and the server may have acted); a Retry-After past a minute is a quota, not a blip, so the backoff runs instead of sleeping it out. Exhausted, the last response falls through to the existing >= 400 branch, so the status_code survives. - _bridge_invoker re-raises 429/5xx alongside 401/403. The frameworks turn a raised tool exception back into model-visible text, so each chat() door gets its own escape: the in-process MCPServer's failure_error_function lets a PageIndex-caused failure propagate and _translate_run_error unwraps it from the framework's wrapper (which also un-flattens the 401 case); the Messages lane runs each turn's tools through the runner's public generate_tool_call_response() and raises before the next model call. - _model_backend_error keeps the provider's status_code. Claude Agent SDK tools cannot fail fast: the SDK MCP server converts handler exceptions into JSON-RPC errors for Claude Code by design. Claude-Session: https://claude.ai/code/session_014S88dcSz7jykegAWyWZk8E * Tool-path fail-fast: cover unreachable servers and all 5xx The bridge retry is now a plain urllib3 Retry: 429 and every 5xx retried three times at the fixed 0/2/4 s backoff, Retry-After ignored. That drops the _Retry subclass, whose get_retry_after raised InvalidHeader on a non-integer header (turning a 429 into "could not reach the server"), honoured a 60 s Retry-After three times over, and let a 413 carrying Retry-After replay. 500 and 504 join the forcelist so the invoker's "what survived the bridge's retries" holds for every status it re-raises. Retry is imported from requests.adapters, the declared dependency. The invoker re-raises transport failures too: once the bridge's own connection retries fail, the model cannot reach the server either, and the envelope only sent it round the retry loop. The handshake error blames the API key only on 401/403: a rate-limited handshake is now a run-terminating error and was telling users to rotate a working key. Docstrings on agent_tools()/build_agent_tools and the Anthropic adapter state the real raise set: 401/403, post-retry 429/5xx, unreachable server. Claude-Session: https://claude.ai/code/session_013xk3xt9KgHNTjsYmFLKxbu * fix: surface Messages tool failures before advancing runner * fix: require urllib3 1.26 for MCP retries * fix: keep the Messages fail-fast quiet and single-path Raise ToolError from the tools chat(protocol="messages") runs instead of the raw PageIndexAPIError: the Anthropic runner log.exception()s any other exception, so every fail-fast printed a 20-line traceback from anthropic's internals before the SDK raised its own error. The lane still records the failure and raises it right after the runner's tool batch, so the ToolError content never reaches the model. Drop the two post-loop _messages_fail_fast calls: the runner executes tools only through the public generate_tool_call_response (0.108.0 through 1.4.0), which checked_tool_response wraps, so they could never fire. Annotate _pageindex_cause for the py.typed package. Claude-Session: https://claude.ai/code/session_01CfYSeq8kM7HjfF79TGbsiT * fix: fail fast on the cloud's account-limit tool errors; one 5xx list RATE_LIMITED / USAGE_LIMIT_REACHED (pageindex-chat #472) arrive as a normal tool error inside HTTP 200, already retried server-side; the invoker re-raises them as 429 / 402, the way a post-retry status escapes, so every lane fails fast without a per-lane change. The bridge retries the whole 5xx range, the same range the invoker re-raises; a non-JSON 200 body (a JSONDecodeError is a RequestException too) stays a model-visible envelope, since the server was reached. The MCP stub tests keep the machine's proxy out of 127.0.0.1. Claude-Session: https://claude.ai/code/session_01F7vqmZdnWeKC9SUBytDrdf * fix: raise the account-limit escape outside the invoker's own try _raise_account_limit fired inside _invoke's try, so its escape depended on 429 and 402 also appearing in the except's re-raise tuple: two lists in one function that had to agree. The check now runs after the try, where the except cannot swallow it, and 402 leaves the tuple (the MCP route never answers HTTP 402; it was there only to let the raise through). The mcp_stub fixture also sets the lowercase no_proxy: requests reads that spelling first, so a machine with no_proxy set still routed the stub requests through its proxy despite NO_PROXY. |
||
|
|
ebf19b21db |
Process display elides an image's base64 payload (#489)
The woven [tool_result] line clipped the framework's image item to its
first 200 characters, which is the head of a data URL. The line now shows
the item as it is with the payload elided: {"type": "image", "image_url":
"data:image/png;base64,..."}. The item itself is unchanged.
Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
|
||
|
|
d0d6aa39e8 |
Tool results reach the frameworks as MCP content; the frameworks render it (#487)
* Tool results reach the frameworks as MCP content; the frameworks render it get_document_image returns an MCP image block, and the SDK flattened every tool result to text on its way to a framework, so an own-model chat received "[image/png content omitted: ~256 KB]" where the hosted lanes (HostedMCPTool, the http MCP config, the managed chat) received the page. The rule now: the SDK carries MCP types and renders nothing. Each framework gets the tool set the way it consumes MCP, and converts the content itself. - McpBridge.call_tool returns the content blocks untouched; the invoke contract behind _tool_specs is (content blocks, is_error) on the cloud and local paths alike (local tools send one text block). - openai-agents: an in-process MCPServer over the tool specs, and the FunctionTools are MCPUtil.to_function_tool over it. The framework's own MCP conversion builds the tools (schema verbatim, strict off) and renders results (text items, image data URLs); its failure pipeline answers malformed arguments. The hand-built FunctionTools are gone. - Anthropic tool runner: results validated as CallToolResult and rendered by the Anthropic SDK's mcp_content (text blocks, base64 image blocks), on the error channel too. - Claude Agent SDK: the content passes through, it is MCP already. - Plain functions (agent_tools): render_text, the former flattening, keeps the size stub for binary content; a string cannot carry an image. Wire consequences: tool results on the openai-agents lanes are the framework's structured items (a single text item for a text result) rather than a bare string, and the Anthropic tool_result content is a block list. ChatStream tool_result events carry that structured output; the woven display shows its text. mcp becomes a declared dependency: it already arrives with openai-agents, and both adapters use its types. Verified live on a 15-page paper, page 1, across chat() on the LiteLLM lane (gpt-5.6-luna), chat(protocol="responses"), and chat(protocol="messages", claude-sonnet-4-6): each answer described the red attribution text and the arXiv stamp rotated along the left margin. Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW * Relay test asserts the wire aliases: mcp 1.x and 2.x differ in attribute names CI installs mcp 2.x, where CallToolResult exposes is_error and content blocks exposes mime_type as attributes, with the wire names (isError, mimeType) as aliases; mcp 1.x uses the wire names as attributes. The adapters only ever validate from and dump to the wire shape, so they run on both; the test now does the same. Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW |
||
|
|
2858384a75 |
docs: the chart images load by absolute URL so PyPI renders them (#475)
pyproject publishes README.md as the PyPI description, and PyPI does not resolve a relative path against the GitHub repo — the five charts render as broken images on the project page (verified on 0.2.14). Every chart now loads from raw.githubusercontent.com; the ten other images in the file already did. PyPI's sanitizer drops <source>, so only the <img> fallbacks matter there; the srcset URLs change too, to keep GitHub's dark variants spelled the same way. |
||
|
|
0400366d35 |
litellm terminal noise: quiet as a default, and the no-fetch stamp on every lane (#473)
* fix: flip litellm's banner switch at the provider lookups, not after them
_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.
Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.
Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9
* refactor: the preload thread only imports — every litellm entry quiets itself now
Since
|
||
|
|
85180eeacd |
fix: extra_body refuses the skeleton keys; sampling fields ride ModelSettings on LiteLLM-routed models (#472)
* fix: extra_body refuses the skeleton keys extra_body merges last on every lane, so a caller's system / instructions / input / messages / tools replaced the managed prompt, the conversation or the doc tools wholesale. The run succeeded and answered without tool guidance or document scoping, silently. Those are the SDK's on every lane (the three-layer rule: skeleton, named knobs, caller extras); the door for the prompt is instructions=. Refused at the two seams where extra_body meets the wire, before the Anthropic transport exists on the Messages lane. Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc (cherry picked from commit |
||
|
|
c8c8293aa7 |
ChatStream moves to a light module so chat()'s type hints resolve at runtime (#471)
* fix: ChatStream lives in a light module, so chat()'s type hints resolve at runtime ChatStream was importable in client.py only under TYPE_CHECKING (a real import would have dragged local_chat's asyncio stack into `import pageindex`), which left chat()'s return annotation a dangling string: typing.get_type_hints(PageIndexClient.chat) raised NameError, and so did anything that introspects signatures — agents' function_tool (client.chat) died on it before looking at a single parameter. The class touches neither asyncio nor the agent frameworks, so it moves to pageindex/chat_stream.py, client.py imports it for real, and the package exports it directly instead of lazily. `import pageindex` still leaves local_chat unloaded. Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY * fix: the ChatStream move keeps its own invariants — future annotations, a guarded eager path, the old import path pinned Review of #471 found the move's guards thinner than they look: - chat_stream.py had no `from __future__ import annotations`, unlike every sibling module, which made its `-> "ChatStream"` quotes load-bearing: unquoting them — the very edit this move made in client.py, and what `ruff --select UP037 --fix` does — broke `import pageindex` outright. - The import in client.py is the whole fix and reads like a typing-only one; a comment says why it must stay real. - The lazy-import test's denylist named no framework, so agents, litellm, openai or anthropic could join the eager path with a green suite — the cost the module was split out to avoid. - The type-hints walk had no floor, so it could silently stop covering anything, and nothing pinned `pageindex.local_chat.ChatStream`, the path the class shipped under in 0.2.11-0.2.14. - local_chat's module docstring still claimed the class. All four guards mutation-checked red. * style: the walk-collapsed assertion message fits the 79-col convention * style: no rationale comments — the guard is the test, the why is the commit |
||
|
|
564bc97c0f |
Chat process follow-ups: a mid-stream error raises, .events survives partial reads, bad show_process chokes first (#470)
* fix: .events survives partial reads; hidden call lines still label results; bad show_process chokes first - ChatStream.events delegated with `yield from`, so a dropped handle (next(stream.events), for ... break) closed the shared run on GC and the rest of the run silently vanished. A plain loop leaves it alone. - _weave filled call_args only past the tool_call visibility guard, so with call lines hidden the standalone result lines never carried the arguments they promise. - show_process is validated before the stream check: an invalid value is refused as such instead of being told to add stream=True and then refused again; the managed lane's duplicate choke goes with it. - Docstring: show_process is not own-model-only. Claude-Session: https://claude.ai/code/session_016M3qaQedSK7L4DwysFRmk2 * fix: mid-stream error chunk raises instead of ending as a short answer; stream docstring says show_process is on by default - The managed endpoint reports a server-side failure as a final {"error": ...} chunk after the partial answer (api.py refunds the credits, then yields it). Neither chunk decoder looked at it, so chat(stream=True) and chat_completions(stream=True) in both modes ended as an apparently complete short answer with no exception. One guard in each decoder raises PageIndexAPIError; the partial answer is still delivered first. - The `stream:` arg and the Returns block still described the pre-PR contract (bare text chunks); only the show_process paragraph said it is on by default. Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY |
||
|
|
1ccb31bf9d |
fix: Responses effort rides extra_body; envelope reports metadata; "" is unset on every lane (#469)
* fix: Responses effort rides extra_body; chat() docs say what the wire does
The rule chat() follows, written down: chat() names PageIndex's own
parameters plus a subset of LiteLLM's unified vocabulary; anything
vendor-specific rides extra_body under the lane's wire names. Three
layers meet on the wire — the skeleton (managed prompt, conversation,
tools) is the SDK's, the named knobs are translated per lane, and
extra_body is the caller's, merged last so it wins. A key the SDK writes
into a nested object must go in through extra_body too, so the caller's
other keys in that object survive — the SDK merge is shallow.
- Responses lane: reasoning_effort joins the caller's extra_body
"reasoning" object (their keys win) instead of riding a separate
reasoning= that extra_body's object replaced whole on the wire.
Mirrors the Messages lane's output_config. The envelope's
given.get("reasoning") now reports the merged object for free.
- reasoning_effort="" is unset on both protocol lanes, like model=""
and instructions="".
- extra_body docstring: drop the invitation to override the SDK's
system/input — that is the skeleton, and a fixed input breaks the
tool loop; name the wire fields per lane instead of one mixed list.
- messages docstring: every system row joins the managed prompt on the
answer lane, not only a leading one (_split_chat_messages hoists all).
- max_turns docstring: the OpenAI lanes raise at the cap, the Messages
lane returns the truncated run — the divergence was documented only on
the private door.
- Migration message: parameters go by keyword; the doors' sampling and
thinking fields ride extra_body.
- chat_model docstring: responses is no longer a chat surface.
- _split_chat_messages refusals no longer name chat_completions, a
method the chat() caller never typed.
- Tests: the Responses door equivalence takes the extra_body shape;
effort/extra_body collision, key survival and "" pinned on both
protocol lanes; the ModelSettings spy asserts the values reached the
wire, not only the envelope.
Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE
* fix: Responses envelope reports the metadata sent
metadata is a Responses request field the caller sets through
extra_body; the envelope hard-coded None. Same source as the other
caller-set fields: given.get().
Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE
* fix: "" is unset on the answer lane; managed-cloud gate reads falsy as unset
reasoning_effort="" reached LiteLLM as a literal empty effort on the
answer lane while the protocol lanes already treated it as unset. And
chat_completions' managed-cloud own-model gate, which chat() routes
through, refused model="" / reasoning_effort="" / {} as knobs the caller
never set. Non-numeric knobs now read falsy as unset, as the local lane
always has (model or chat_model, if extra_body, backend or {}); numeric
ones keep `is not None`.
Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit
|
||
|
|
a7995a4d22 |
Ignore the default results/ output directory (#468)
`run_pageindex.py` writes its trees to `./results` (`output_dir = './results'`, lines 142 and 198), so running the CLI leaves untracked JSON at the repo root. That output has already been committed by accident ten times — 51 files, 3.2 MB, across commits from "first commit" through "improve tree optimization". `examples/documents/results/` is unaffected: this only ignores the top-level directory the CLI writes to. Claude-Session: https://claude.ai/code/session_01KXmNuHrkQEkefNS12V7apV |
||
|
|
5189cf2892 |
chat(protocol=): the protocol doors move behind the front door (#460)
* feat: chat(protocol=) — the protocol doors move behind the front door responses() and messages() become _responses()/_messages(): the same engines, reachable as chat(protocol="responses"|"messages") with the protocol's own input and output shapes (transcript items / content blocks in, the envelope or native stream out). The old names are the vendor SDKs' own, and an agent-written client.messages(...) now fails fast with the way in — a runtime-only __getattr__, invisible to the type checker so attribute typos on the client still get flagged. chat() gains the knobs that lost their public home: instructions (appended after the managed prompt — a string on every lane, Messages system blocks with protocol="messages"), max_turns, backend, extra_headers, extra_body. reasoning_effort lands natively on each lane: LiteLLM's kwarg, Responses reasoning.effort, Anthropic output_config.effort. show_process stays the answer lane's view. Six overloads keep the return types narrow for py.typed consumers; the answer lane's contract is unchanged. The managed cloud chat rejects the own-model knobs as before. Claude-Session: https://claude.ai/code/session_017uvMD9eatgdaupLGFUjuKM * fix: chat(protocol=) honors extra_body thinking; honest envelope and remedies - Messages lane: through the public door the thinking budget rides extra_body, so the default max_tokens lift reads it there (extra_body wins on the wire, so it wins in the lift); the documented extra_body={"thinking": ...} no longer 400s on max_tokens < budget. - Responses lane: temperature/top_p/reasoning/max_output_tokens reach the wire via extra_body only; the envelope now reports what was sent instead of the never-set locals. - One _require_own_chat refusal for chat(protocol=...), the doors behind it, and instructions. The doors' own copies had already drifted from chat()'s text, and every copy pointed a local client with chat_model blank at the managed chat it does not have; a door reached directly on such a client fell through to a raw AttributeError. The check is the one chat_completions already makes. - Messages lane treats model="" as unset, like every other model check. - The protocol refusal for show_process runs before the stream=True hint, so the first remedy offered is the right one. - instructions="" configures nothing (no empty system row, cache key unchanged), matching the protocol lanes. - Responses validation names messages, the public parameter, not input. - Stale docstring pointer to messages() fixed. - A protocol=None, stream: bool overload restores str | ChatStream for a runtime-variable stream; the catch-all had widened it to a 4-way union. - Tests: door X refuses like chat(protocol=X) on managed and blank-local clients; the protocol gate and the non-stream show_process order are asserted by distinct messages; chat(stream=True) joins the max_turns matrix; each fix carries a red-verified assertion. Claude-Session: https://claude.ai/code/session_01M9GdjnuDHHwDKqMzvPWCcj |
||
|
|
91b6c365e0 |
Chat process display: labels are the event type names (#459)
* style: label woven lines by their event type names [tool_call] and [tool_result] replace the "[tool]" label and the "->" arrow, so the text view's labels are exactly the .events type names (config keys stay plural — they switch a class of lines; each line is one instance). Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP * style: tool_result lines flush left — no nesting indent Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP * style: one vocabulary — show_process keys are the event type names tool_call / tool_result (singular) everywhere: event types, text labels, and now the config keys, which select event types by name. Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP |
||
|
|
ad7956d3c3 |
Keep litellm's terminal noise out of answers: stdout banners, logger chatter, retry print (#455)
* fix: mute litellm's stdout Provider List banner in the litellm lanes litellm's OpenRouter adapter probes supports_reasoning() with the provider-stripped model name on every completion, so any model missing from its static map (e.g. openrouter/z-ai/glm-5.3-flash) makes get_llm_provider print a red "Provider List:" banner straight into stdout — interleaved with the streamed answer, once per agent turn. suppress_debug_info is litellm's own embedder switch (its Router sets it too) and gates only this banner and the "Give Feedback / Get Help" one; errors still raise with their full text. Applied at the same lazy hook points as the existing litellm repairs, plus the background preload, so merely importing pageindex still leaves the host's litellm untouched. Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL * fix: retry notice rides logging, not the caller's stdout llm_completion/llm_acompletion printed '* Retrying *' straight into stdout on every retried request — the channel that belongs to answers and CLI output. The notice moves to logging.warning beside the error line that already accompanies it. Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL * fix: gate litellm's stderr WARNING chatter alongside the stdout banners The banner mute grows into _quiet_litellm: litellm's own logger sprays WARNING records (remote-map fetch fallbacks, cost hiccups) onto stderr from inside requests — not actionable for SDK callers, whose real failures raise as exceptions. The preload stamps LITELLM_LOG=ERROR before litellm's import initializes its logger (setdefault, so an explicit caller choice wins, and litellm honors a chosen level itself); the hook's setLevel covers litellm imported before us, plus the dotted litellm namespace its adapters log under. Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL |
||
|
|
555c370f66 |
Show the chat run: show_process weaving and the ChatStream events view (#454)
* feat: show the chat run — show_process weaving and the ChatStream events view
chat(stream=True) now returns a ChatStream: iterating it yields the
answer text with the run woven in by default — "[thinking] " sections,
one "[tool] name arguments" line per call with its clipped result —
and .events yields the run as typed dicts (thinking/answer deltas,
tool_call with parsed arguments, tool_result with the full output).
One run serves one view; close() kills it like a closed generator.
- show_process: on by default ("on where available"); False for the
bare answer stream; a dict (ChatProcessOptions: thinking, tool_calls,
tool_results, max_chars) selects the parts. Explicit True without
stream=True raises.
- Managed clients weave what the endpoint serves: tool-call lines
parsed from its block_metadata chunk tags (that wire carries no
thinking and no tool results); old-wire chunks stay plain answer
text, and the bare answer view no longer leaks tool-argument JSON.
- Engine: one typed-event primitive (_chat_events_agen /
_cloud_chunk_events) with _weave as a pure renderer over it; the
chat lane's prologue is shared via _chat_agent, behavior unchanged
on chat_completions/responses/messages.
453 tests green (16 new, red-verified), no-openai-agents leg simulated,
pyright flat vs main.
Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
* fix: pair woven tool results with their calls; drop empty deltas; validate show_process before the managed request
Three review findings on the show_process weave, all red-verified:
- _weave nested every tool result under the most recent [tool] line,
which misattributes results when a turn makes parallel calls (the SDK
streams all calls, then all results). A result now nests only when the
line above is its own call (by call_id); otherwise it stands alone
with its call's clipped arguments echoed, so same-name parallel calls
stay tellable apart. Hidden-call mode falls out unchanged (no stored
arguments, no echo).
- Empty deltas now stop at the event source. chat_completions and the
managed chunk lane both filter them; the new local lane did not, so a
mid-stream "" (litellm forwards annotated/provider-field empties)
leaked into the bare view and flipped _weave sections, splitting one
thinking burst into repeated labels.
- The managed streaming lane sent the billed request before
_process_options ran, so a config typo cost a real chat call. chat()
now chokes on bad show_process before dispatching, matching the local
lane's validate-first order.
Two docstring truths: close()'s "a run never consumed never starts"
holds only for own-model chat (the managed request is already on the
wire), and the module docstring now covers the managed chunk weave.
456 tests green; the three new ones red-verified; managed-lane paths
re-run with the agents package blocked; pyright adds nothing on touched
lines.
Claude-Session: https://claude.ai/code/session_01XeeD2214z6Vd6qKcAi9ZvJ
* fix: seal open tool blocks from the answer; typed chokes; chat() stream overloads
Four review fixes plus two coverage gaps, each red- or
mutation-verified:
- _cloud_chunk_events treated any non-tool_use tag inside an open tool
block as answer text, so argument JSON leaked into the
show_process=False answer — the one meant to be appended back as
conversation history. Inside an open block nothing is answer:
argument chunks now accumulate under any tag. And non-string
argument pieces stringify at the join instead of killing the whole
stream with a raw TypeError.
- _process_options sorted unknown keys before repr-ing them, so
mixed-type keys ({1: True, "foo": 1}) raised a bare TypeError past
the caller's `except PageIndexAPIError`; sorting the reprs keeps the
single error type.
- chat() gains @overload on stream, so the docstring's own `.events`
usage type-checks for py.typed consumers (previously pyright ruled
`Cannot access attribute "events" for class "str"` on the exact
documented snippet). pageindex/ error count unchanged (234).
- Coverage: the streaming lane's whole `finally` could be deleted with
the suite still green — the new abandonment test pins the teardown
(pump exits, turn 2 emits nothing, _aclose_backend closes the
per-call client). And FakeModel emitted only the reasoning event
production never sends (litellm folds reasoning into
reasoning_content, which arrives as summary deltas); it now
alternates variants, so dropping either from the isinstance tuple
goes red.
459 tests green; managed-path tests re-run with the agents package
blocked; flake8 parity on every touched file.
Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
* fix: reading .events is inert — consuming claims the view; honest falsy show_process message
ChatStream.events was a property whose getter latched the stream's one
view on mere attribute access: a debugger variable pane, hasattr, or
getattr(stream, "events", None) — which PageIndexAPIError escapes, as
getattr only swallows AttributeError — was enough to make a later
`for chunk in stream:` refuse, with nothing consumed. The getter now
returns a lazy generator: the managed refusal, the view claim and the
run start all happen on first consumption, so introspection is
side-effect free and the text view stays usable after a probe.
And the stream=False guard's message told falsy-but-not-False values
("show_process=0", "") that they passed show_process=True; the check
itself is the ruled falsy-{} trap and stands, but the message now
names the off values and echoes what was got.
461 tests green (2 red-verified new: inert read on both lanes, plus
the falsy-message case); changed managed-path tests re-run with the
agents package blocked; pyright pageindex/ 234 -> 234.
Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
|
||
|
|
7d18cc033c |
docs: drop stray blank lines in the cloud quickstart snippet
Claude-Session: https://claude.ai/code/session_01Wh75QagUhr5DPVATjmGrCe |
||
|
|
9fee239b17 |
docs: correct what the index model does (#441)
* docs: correct what the index model does The index model does not build the tree structure — Flash extracts it from the document layout without an LLM. The model only summarizes and refines the tree. Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD * docs: name PageIndex Flash in the submit_document note Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD |
||
|
|
21f2b1018e |
docs: FinanceBench chart as local light/dark assets (#440)
* docs: swap the FinanceBench chart for local light/dark assets Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF * docs: replace the FinanceBench chart with light/dark assets Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF * docs: narrow the FinanceBench chart to 60% Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF * docs: set the FinanceBench chart to 65% Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF * docs: set the FinanceBench chart to 70% Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF |
||
|
|
688ce200fa |
docs: hide the FinanceBench accuracy chart
Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF |
||
|
|
d947ab6c93 |
docs: README structure and layout pass (#438)
* docs: the benchmark charts render at 70% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the benchmark charts render centered at 80% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the benchmark charts render at 85% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the benchmark charts render at 90% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the collapsed sections drop the spacer <br> Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the header link row drops Discord Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Updates entry comments out the MCP/API pointer Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the one-line summary stands without the Why it works heading Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: TEMP four punchline variants side by side Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the one-line summary reads as a centered pull quote Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: TEMP alert-style variants next to the pull quote Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the pull quote sits under an In one sentence heading Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the pull quote sits under a TL;DR heading Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the TLDR quote bolds the claims and drops the vector DB and chunking Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the TLDR quote reads no vector DBs or chunking Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the TLDR quote leaves retrieval unbolded Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Retrieve step leads with agentically Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the TLDR quote reads left-aligned Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the TLDR quote wraps on its own Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations example uses a plain system-message string Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations example keeps the system prompt inside the code block Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations example inlines the system prompt Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Cloud example drops the doubled blank line Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench case study returns to Benchmarks Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench paragraph speaks of PageIndex directly Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench sentence names the corpus and the margin plainly Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench sentence drops the corpus aside Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench sentence keeps the benchmark description Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Benchmarks splits into the local open-source run and FinanceBench Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: top-level sections use single-hash headings again Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the local benchmark part is headed Local mode Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the local benchmark part is headed PageIndex Local Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the benchmark charts render at 80% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench chart stays at 90% and the benchmark gloss goes in parentheses Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench chart returns to its original 70% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the local benchmark charts render at 70% width Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the local benchmark part is headed Running PageIndex locally, charts at 75% Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations example indents the prompt continuation Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations prompt breaks before the cite tag Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the citations prompt breaks before using Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the collapsed usage guides follow the Quickstart Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: a Usage section holds the two collapsed guides Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage guides sit at h3 with their steps one level down Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage guides sit at h2 so their steps keep their levels Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage guides sit at h3, inner levels untouched Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: each usage step collapses on its own under Usage Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Usage keeps its two guides as headings with collapsed items inside Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the *_config note folds into Other agent frameworks Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Usage and the Detailed Usage Guide open with a line of orientation Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in drops the happy-path idiom Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in names the two guides directly Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the guide and integration lead-ins say what each block holds Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in contrasts direct use with integration Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in lists its two guides; their intros stay general Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in is one line with (i) and (ii) Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in says SDK client and your own agent Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in uses parallel verbs Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in says access rather than call Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in drops the verb in (i) Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the integration lead-in reads as one plain sentence pair Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the client guide is headed Use PageIndex through the SDK client Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in lists its two ways as bullets Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage lead-in returns to one (i)/(ii) line Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Ready to Try It links cover only the noun Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Usage section is headed Usage Guide Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the summary heading reads tl;dr Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: collapsed items use bold summaries so they sit close together Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: expanded items get a line of air under their summary Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: collapsed items try h4 summaries again Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: collapsed items settle on bold summaries with a spacer Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: a line of air between neighbouring collapsed items Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Step 2 anchor lives inside its summary so item gaps match Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the two guide headings rely on their generated anchors Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the SDK client guide opens with a plainer line Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the SDK client guide lead-in drops workflow Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the SDK client guide lead-in names its three steps Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the SDK client guide lead-in, polished Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the SDK client guide lead-in points below Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: a line of air before the integration heading Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: no spacer before the integration heading Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the two usage guides are labelled (a) and (b) Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench closing line names the evaluation Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench closing line links only the results Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench claim links straight to its benchmark subsection Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the indexing-cost line names the model in prose Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench subsection is headed Leading on FinanceBench Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the FinanceBench heading reads Leading accuracy on FinanceBench Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Updates lists the PageIndex File System again Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the File System update entry links in plain weight Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the header link row gains Blog Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the summary heading reads TL;DR Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Model Recommendations points at the Usage Guide Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Model Recommendations names the SDK client usage guide Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the usage-guide pointer promises more than models Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the usage-guide pointer, one word shorter Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the integration lead-in says each example covers one framework Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the integration lead-in says a different framework Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: each framework fold shows the one-call and explicit forms end to end Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the streaming example uses a literal and the quickstart drops trailing spaces Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Flash update entry says where the structure comes from Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Flash update entry, tighter Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Flash update entry says built by an LLM Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Flash update entry, Ray's wording Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: the Flash update entry keeps extracted heuristically Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k * docs: Updates dates read [Aug '26] Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k |
||
|
|
d4ce6ee65f |
docs: the run_messages comment keeps the principle, drops the version numbers
The rationale to preserve is that execution policy belongs to the vendor and only the runner's history is ground truth; the exact 1.0/1.1 switch point lives in the previous commit's message. |
||
|
|
070293bdca |
test: the messages max_tokens test follows anthropic 1.1's stop policy
anthropic 1.1.0 replaced the runner's refusal special case with an explicit stop-reason table: every non-tool_use stop is terminal and its tool_use blocks are never executed (1.0 executed a max_tokens turn's complete blocks). The envelope already keys on the runner's history, so the product adapts by design; only the test had the 1.0 behavior baked in. It now asserts the version-appropriate shape on both sides, and the run_messages comment describing the old behavior is reworded to name the policy split. Verified: 434 green on anthropic 1.1.0 (CI's failing config) and on 0.120.2 (the pre-1.1 branch); no-frameworks collection stays clean. |
||
|
|
f47f27b7b3 |
docs: README restructure with quickstart, benchmarks, and a collapsed usage guide (#431)
Quickstart on the v0.2.11 index=/chat= spellings, a citations example, indexing-time benchmarks with new charts, the own-agent integration in its own collapsed block, a rewritten PageIndex Cloud section, the reasoning positioning restored, and prose em dashes replaced with plainer punctuation. Earlier snapshots of this branch landed via #418/#419; this squash carries the tail. |
||
|
|
31f01910cd |
docs: API-key URLs follow the dashboard move to developer.pageindex.ai
dash.pageindex.ai/api-keys now 307-redirects to developer.pageindex.ai/api-keys, and the README (PR #431) already standardized on the developer domain — the client docstring and the CloudClient keyless error message were the last SDK references to the old host. |
||
|
|
57252e2686 |
fix: the LiteLLM lanes hide litellm's bridge usage warning (#430)
Streaming chat() on an OpenAI gpt-5.4+ model with function tools makes litellm re-route chat.completions through the Responses API. When the stream ends, litellm's logging stores a chat-shaped usage dict inside a ResponseAPIUsage field and model_dump()s it, so pydantic prints "Expected `ResponseAPIUsage` - serialized value may not be as expected" once per streamed turn. litellm does this on purpose (litellm_logging._get_assembled_streaming_response, 1.97 and 1.98) and the answer is unaffected. Both seams that hand a LiteLLM-routed model to openai-agents — the chat lanes' model builder and openai_agent_config()'s litellm/ lane — register one warnings filter matching exactly that message; every other warning still surfaces. A process-wide filter is the only placement that works: litellm emits the warning from its async success handler on the worker thread, out of catch_warnings' reach. Claude-Session: https://claude.ai/code/session_01APJbbp3jpRxJhvchcU2Aed |
||
|
|
174f95f35b |
fix: the slots accept any Mapping at runtime; eight docstrings stop calling cloud tool scoping server-side (#429)
fix: the slots accept any Mapping at runtime, as their annotation admits; eight docstrings stop calling cloud tool scoping server-side |
||
|
|
b9a9a3b5aa |
fix: the .env search ends at the cwd tree; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike (#428)
* fix: .env stays unset when the cwd tree has none; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike find_dotenv(usecwd=True) returns '' when nothing is reachable from the cwd, and `or None` turned that into load_dotenv's own upward walk from utils.py — the install-dir leak the cwd search was added to replace. A pip-installed SDK could load another project's .env from above site-packages, silently. _local_chat treats a blank chat_model as "managed chat", which a client without an api_key does not have: chat_completions() then reached for LocalAPI.chat_completions and raised a bare AttributeError. The managed branch now refuses as a PageIndexAPIError naming chat_model. py.typed made the annotations authoritative while storage_path was typed str; _ARG_TYPES accepts os.PathLike, so Path(...) ran fine and failed the user's type check. Both signatures and LocalIndexConfig now say so. Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb * fix: the exported config shapes pass into index=/chat=; a comment and two docstrings stop overclaiming The slots were annotated dict[str, Any]. A TypedDict is consistent with Mapping[str, object], never with dict (PEP 589: a dict-typed receiver could write arbitrary keys through it), so the four shapes types.py exports — and py.typed advertises to installed callers' checkers — could not be passed to the one place they describe. pyright on a probe that does exactly that: 9 errors before, 0 after. The constructor only reads the slot (items(), then a fresh conf dict), so Mapping is the honest bound; a plain dict is a Mapping, and TypedDict instances are plain dicts at runtime, so nothing moves at runtime. The _ARG_TYPES comment said "every value" is shape-checked; api_key is not in the table (its empty check is separate, its type check stays unchecked by ruling), so the comment now speaks for the table only. _local_doc_scope and _require_local_scope still explained the cloud drop as "scoping is server-side" — true of the managed chat, which never reaches either function. What reaches them on a cloud client is own-model chat and the config helpers, whose cloud tools take no allowlist: targeting there is prompt-level only, as the error message between them already said. 434 passed; pyright on pageindex/ unchanged at 235 (0 in the touched files, before and after). Claude-Session: https://claude.ai/code/session_01TxG8u8x29XRnK4yscZVCch * test: the install-dir .env test is named for what it asserts Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb |
||
|
|
920db2b1b1 |
feat: the client grows two sides — documents and chat each pick their home (#424)
* feat: the client grows two sides — documents and chat each pick their home
One client, two independent switches: api_key decides where documents
live (the PageIndex cloud, or the local store); a configured chat model
decides who answers (your own model in your process, or the managed
cloud chat). Their free combination opens the bridge — cloud documents,
your model — and the fourth cell stays unspellable.
- index=/chat= slots: string shorthand or grouped dict, 1:1 with the
flat arguments; one spelling per side, sides mix freely
- optional "type" everywhere (top-level and in either dict): always
omittable, checked against the content, meaningful alone —
type="cloud" is a keyless cloud spelling
- PAGEINDEX_API_KEY is read only when the code explicitly says cloud
(PageIndexCloudClient(), type="cloud", "pageindex-cloud",
{"type": "cloud"}); a bare PageIndexClient() stays local
- bare mode words ("cloud", "local", …) are reserved: they error
with the real spellings instead of silently parsing as model names
- bridge chat runs the in-process agent over the live cloud MCP tools
and instructions; doc_id targets at the prompt level; citations stay
managed-only; an auth-shaped backend failure explains whose
credentials run the model
- typed shapes (IndexConfig, ChatConfig) ship as optional annotations
Every previously working program is byte-for-byte unchanged: the only
behavioral deltas are error paths — reworded guidance, and the
api_key+chat_model combination graduating from an error into the
bridge.
* fix: the constructor refuses empty and mistyped values on every spelling
- .env keys reach all four keyless-cloud spellings: utils' import-time
load_dotenv now runs before every PAGEINDEX_API_KEY read
- an empty chat-side value ("", {}) errors instead of silently selecting
own-model chat on the default model; None-valued slot keys mean absent,
exactly like the flat arguments
- _local_chat derives from chat_model, so a post-construction assignment
switches the whole client, never half of it
- model= beside a slot gets the split guidance (index_model=/chat_model=)
instead of "two spellings of the same thing"
- the messages door wraps provider failures through _model_backend_error,
and 401s count as auth-shaped even without "api key" in the text
- keyless-cloud hints name the spelling that actually combines; slot
strings are stripped; wrong-typed values raise PageIndexAPIError
- retrieve_model/chat_backend docs drop the stale "Local mode only";
the local-scope refusal no longer claims bridge tools are server-scoped
* fix: type= cross-checks the index slot; the cloud pinned class frees its chat side
- type= beside index= now does what the docstring promises: agreement
passes, disagreement errors, and a mistyped value reports the
vocabulary error instead of a spelling collision
- PageIndexCloudClient grows the chat-side arguments (chat=, chat_model,
retrieve_model, chat_backend), so "pin the index side" is literally
true and the chat surfaces' construct-with-chat_model guidance is
followable on it
- the four chat doors' doc_id entries carry the enforcement split the
config helpers already state (local: tool-layer allowlist; cloud:
prompt-level / server-side)
- types.py stops claiming slot keys share the flat names — the side
prefix is factored out, index={"model"} is index_model=
* docs: bridge-reachable wording — dependency errors say own-model chat, hints name a chat= model
- the three framework-missing errors said "in local mode", which is
wrong on a bridge client (cloud documents + own model) — they now
explain the dependency the way the surfaces do: your own chat model
- the construct-with guidance reads "(or a chat= model)": a bare
chat="pageindex-cloud" is also chat= but selects the managed side
- the mechanical Local-only → Own-model-chat-only substitution left
orphan fragments and two overlong lines; those paragraphs re-flowed
* fix: managed chat reads None; the bridge stops paying per-turn tool lists
- a managed-chat cloud client stores chat_model/chat_backend as None, so
the documented attribute reads instead of raising AttributeError;
_local_chat derives from "is a chat model configured"
- McpBridge caches tools/list per session — every chat turn rebuilds the
tool set, and the round trip was pure latency; the 404 session-expiry
reset drops the cache with the session
- run_messages builds tools before the transport: on a bridge client
that build is network I/O, and a failure there stranded a per-call
anthropic client ahead of the try/finally
* refactor: the side declaration is spelled mode=, not type=
"type" is Python's own word — a builtin, and "data type" beside the
TypedDict shapes; "mode" is what the SDK already calls the two sides
("local mode", "cloud mode"). Same grammar everywhere the declaration
appears: the top-level argument, the index dict, the chat dict, the
typed shapes. The rename also frees the builtin inside the constructor,
so the shape-check error names the offending class through type() again.
"type" in a slot dict is now an ordinary unknown key.
* fix: the reserved-word errors stop calling "cloud" not a mode word
With the declaration key spelled mode=, 'index="cloud" is not a mode
word' contradicted its own remedy, index={"mode": "cloud"} — "cloud" is
exactly a mode value. The four bare strings are reserved words; the
message now says so.
* fix: a blank tools/list is not cached; the auth note's managed exit is chat-lane only
- McpBridge.list_tools caches only a non-empty list — a transient blank
(a deploy blip, a gate misconfiguration) would otherwise run every later
turn with zero tools while the instructions still name them, and only a
404 session reset could clear it
- the 401 architecture note appends "drop the chat model configuration"
only on the chat lane: responses() and messages() refuse a client
without an own model, so on those lanes the exit sent the caller in a
circle
- CloudIndexConfig says api_key is omittable only while mode: "cloud"
stays — index={} refuses as an empty dict rather than reading the env
* fix: the bridge fetches tools/list per call again; .env resolves from the cwd
- McpBridge.list_tools no longer caches: the tool set is built once per
SDK call (Agent(tools=...) ahead of Runner.run; build_anthropic_tools
ahead of tool_runner), not per model turn, so the cache saved one round
trip per later call while a mid-pagination 404 replayed a dead cursor
into a duplicated (and cached) list, and the list went out by reference
across a lock dropped between miss and store
- utils.load_dotenv searches upward from the cwd: a bare load_dotenv()
walked up from utils.py, which is site-packages for an installed SDK,
so the four keyless-cloud spellings never saw a project-root .env; the
package-relative walk stays as the fallback
- the emptiness guard strips strings: chat_model=" " selected own-model
chat, the silent flip the guard's own comment rules out
- the _local_chat comment stops advertising post-construction assignment
as a full mode switch
* fix: the pinned classes take index=/chat=; "cloud"/"local" are mode words; a blank chat_model stays managed
- PageIndexLocalClient takes index= and chat=, PageIndexCloudClient takes
index= — the grouped spelling of the flat vocabulary each already took;
their refusals name the class and an exit that class can take, and the
mode cross-check runs before any environment read
- "cloud" and "local" are accepted wherever "pageindex-cloud" was (index=,
chat=, mode=, {"mode": ...}), case- and whitespace-insensitive; "hosted"
and "managed" still refuse, pointing at the real word
- every spelling strips its strings, and the slot spellings' type/empty
errors name the slot key (index["model"]), not the flat argument
- _local_chat treats a blank chat_model as managed: the constructor
refuses "", so assignment agrees instead of opening the bridge on a
nameless model; openai_agent_config carries no model then either
- an empty MCP tools/list raises like empty instructions does — a
zero-tool agent would answer from the model's own knowledge silently
- enable_citations names the real gate (managed vs own chat), not
"cloud-only", on a cloud own-model client
- pageindex/py.typed: the exported config TypedDicts reach installed
type-checked callers
* test: the two framework-door tests skip without openai-agents
as_openai_tools() and openai_agent_config() need the agents package, which
the "without frameworks" CI legs do not install — the same importorskip
every other test on those doors already carries.
|
||
|
|
416e304f51 |
perf: expand schedules dependency-exact at thirty-two concurrent proposals (#422)
* perf: expand proposes a wave of nodes concurrently The expand loop awaited one propose_children at a time — 20-30 nodes at ~3s each put 1-3 minutes of pure round-trip latency on every default local submit. Nodes waiting in a wave are all frontier leaves whose decisions cannot affect each other, so the model half now runs concurrently (EXPAND_CONCURRENCY = 8) while the apply half stays serial in wave order: decisions, log entries, and child ids land exactly as before, and children attach into the next wave. A fatal classification still aborts the run right after the wave's gather. Benchmarked on real PDFs with a fixed-latency fake model: 408 pages 21.1s -> 3.0s, 758 pages 28.2s -> 3.5s (7-8x); final trees byte-identical to the serial pass on both. The cap stays low on purpose: expand treats an exhausted retry ladder as fatal, and a wide burst on a rate-limited account would trip exactly that — 8 already collapses minutes to seconds. * perf: expand schedules dependency-exact instead of in waves A child's only prerequisite is its own parent's apply, so each kept node gathers its children directly rather than waiting for its whole generation to finish. Same recursive shape as summarize_tree; the semaphore still caps in-flight proposals at 8; trees are unchanged. * perf: expand admits thirty-two concurrent proposals Cap sweeps on six real documents put the speed plateau at 32: the ready frontier tops out at 21-28 nodes on few-hundred-page PDFs, so 64 buys nothing while doubling the burst. Live runs at 32 cut the expand phase 24-30% on the two documents wide enough to feel it, with zero ladder retries anywhere - and summaries already burst twice as wide through the same ladder. |
||
|
|
8289729aff |
fix: indexing failures surface loud, chat envelopes append verbatim (#421)
Indexing: dead credentials or a missing model fail the run instead of storing a document with blank summaries; a 400 (context_length_exceeded) skips the retry ladder — the prompt will not shrink — and stays a per-prompt failure the run absorbs; all-empty model replies can no longer store a retrieval-ready document; the one-sentence doc description absorbs its own context overflow instead of discarding a fully indexed document; the heading-less flash refusal points at mode='standard'. Chat: messages() output is append-verbatim clean — unset response-only defaults are dropped (no "caller": null the request schema rejects); Claude cache marks follow the wire routing; model_settings and name are openai_agent_config parameters; one Anthropic client per backend; lifted thinking defaults are clamped to the model's output ceiling from LiteLLM's capability map. Store and inputs: lone surrogates are scrubbed from page text and the stored basename, so the returned name is byte-for-byte the stored name and the rename warning fires; NaN/Infinity metadata is rejected at the gate; every cloud error now carries its HTTP status. CLI: the flash lane resolves the summary model through ConfigLoader like the standard and markdown lanes; an empty flash structure errors like the SDK instead of writing "structure": [] with exit 0; --summary-model reaches the markdown lane; the SDK page-spec surface keeps 0.2.10's whitespace tolerance while the tool layer stays strict. pypdfium2 stays on the 5.x line for every install; the 4.x code paths are tested compatibility insurance with their own CI leg; process-pool construction failure falls back to the sequential parse; a py3.10 GC flake in text extraction is fixed. Port of feat/local-chat 0667e3b..1993740 (28 commits); README and assets untouched. |
||
|
|
418d1549bb | docs: weave the original README's highlights back into the SDK rewrite | ||
|
|
b628190529 |
ci: dev tags publish to PyPI only — skip the GitHub Release (#414)
Dev builds need an explicit ==pin to install, so their GitHub Releases carry no install value and double the feed next to the same-day stable. |
||
|
|
ba0ef02d78 |
fix: Aug 18-19 review hardening — five max-effort rounds across tools, chat, and flash (#413)
Squash-port of feat/local-chat's post-dev5 wave (03ffab3..49a24e1): five review rounds of reproduced-then-fixed findings. Highlights: standalone agent_instructions back on the strict shadow check (the *_agent_config bundles keep the relaxed in-set check they can prove); caller-owned http_client survives the per-call backend closes; get_tree keeps key_items under the flash merge default; OpenAI-protocol classification follows litellm's own routing (azure/openrouter/deepseek/groq/xai); thinking-aware max_tokens default shared by messages() and anthropic_runner_config(thinking=); 401/403 re-raise instead of retry-coaching envelopes, bridge errors carry status_code; cloud discovery + instructions ride the ?tools=read endpoint matching the tool gate (live-verified); the chat lane's litellm model resolution deduplicated into utils._litellm_model; explicit optimize= wins over the deprecated optimize_expand (DeprecationWarning added); mcp 2.0 compatibility; spawn-worker and doc-scope guard fixes; the suite is .env-independent and pins the LitellmModel._fetch_response seam. Per-finding rationale in the ported commit messages on feat/local-chat. |
||
|
|
ae2a5b49b5 |
feat: backend connection overrides; extra_headers on every local door (#411)
* feat: backend connection overrides; extra_headers on every local door
Two clients, two configs — the gap this closes. index_backend /
chat_backend on the constructor (and per-call backend on the chat
doors, mirroring per-call model) carry connection params in each
lane's own vocabulary, verbatim: the indexing gateways take the dict
as LiteLLM call kwargs (a contextvar scopes it per operation, and the
env-var key pre-check yields to it), the chat lane lifts api_key /
base_url into LitellmModel's two pinned constructor slots and rides
the rest as call kwargs, responses() and messages() hand the dict to
their SDK client constructors. Per-call keys win over the client's;
the openai-SDK fast path normalizes LiteLLM's api_base spelling.
Config bundles deliberately don't carry it — you run those in your
own environment (docstring says so).
extra_headers lands on all three protocol doors, each engine merging
caller headers verbatim (anthropic-beta wire-proven on messages).
Wire-probed exception, documented on the chat door: LiteLLM's
anthropic adapter owns the anthropic-beta header and drops the
caller's value — Anthropic beta flags belong on messages(). Cloud
rejects the new knobs like the other local-only params, and the
docstrings now say credentials belong in backend, never extra_body
(on the bare lane they would leak into the JSON body without
touching auth — wire-probed).
* docs: constructor Args list gains index_backend / chat_backend
The class docstring documents every constructor argument; the backend
wave added two without entries. Same phrasing as the per-call docs:
index lane is LiteLLM vocabulary verbatim, chat_backend reaches
whichever door runs (api_key/base_url portable across all three).
* chore: three audit leftovers
The classic pipeline's ThreadPoolExecutor import (unused since the
0.2.9 merge) goes. The bare-model missing-key error now also names
the backend={'api_key': ...} route, which satisfies the same check.
messages() joins the uniform: a bad backend dict wraps as
PageIndexAPIError like the responses door, instead of leaking the
SDK's raw constructor error.
* chore: narrow the anthropic wrap to TypeError; key advice names chat_backend
Anthropic's constructor raises only TypeError in our supported range
(unknown kwargs, conflicting credentials) — a missing key defers to
request time, so catching AnthropicError there guarded an impossible
case. And chat() has no backend parameter, so the missing-key advice
now names chat_backend alongside the per-call route.
|
||
|
|
08ea1d975c |
fix: py3.10 litellm type repair; chat() gains reasoning_effort (#410)
* fix: rebuild litellm's Message/Delta types on Python 3.10 litellm 1.97.0 ships Message and Delta annotations whose nested forward refs (ChatCompletionReasoningSummaryTextBlock et al) do not resolve on 3.10, so every completion() dies constructing its response object — non-stream and stream alike (upstream BerriAI/litellm#36384, open, no patch release; 1.96.2 is clean, so the floor raise surfaced it, and pydantic 2.12/2.13 both reproduce). The repair rebuilds the two models once with their defining modules' namespaces at our three completion gateways; version-gated to <3.11 and best-effort, so it is a no-op on healthy interpreters and future fixed litellm releases. Verified on a 3.10 venv: the previously failing anthropic wire test and the full suite pass (250 green, matching CI's matrix leg). * feat: chat() takes reasoning_effort — the front door's one thinking knob Ruled in as a business-level control alongside model: who answers, and how hard it thinks. Same name, values, and verbatim semantics as chat_completions underneath (LiteLLM's cross-provider tier string); unset sends nothing so each backend's own default behavior applies. Sampling and wire-level knobs deliberately stay off the front door. |
||
|
|
bc1c1740ff |
feat: model knobs, LiteLLM-verbatim chat lane, per-door passthrough params (#409)
* docs: drop the demo's install step — openai-agents ships with the SDK now
* feat: the chat lane routes every model through LiteLLM — bare names included
The direct-OpenAI special case existed to dodge LiteLLM's import cost,
and it made OpenAI's own Responses-first models fail on the front door:
gpt-5.6-sol 400s on chatcmpl+tools while reasoning is on (server-side
policy — wire-captured with no reasoning_effort in our request).
LiteLLM 1.97 translates such calls onto /v1/responses; 1.84 does not,
so the sol-class 400 now carries its two exits (upgrade litellm /
responses()).
Routing after the flip: chat protocol — bare names are OpenAI-compatible
shorthand (wire form openai/<name>; OPENAI_API_KEY / OPENAI_BASE_URL
still select the backend, and the missing key stays a build-time
failure), litellm/ strips, openai/ opts out to the OpenAI SDK directly;
responses protocol unchanged (OpenAI-SDK native, LiteLLM refused).
The import cost is handled instead of dodged: local clients preload
litellm on a background thread (first call then perceives 0.0s), and
pageindex sets LITELLM_LOCAL_MODEL_COST_MAP=True via setdefault —
LiteLLM's import otherwise blocks on a network fetch of its price map
(fresh venv: 5.6s -> 1.3s; offline it hangs to the timeout).
Also restores prompt_cache_key delivery, found dead during the flip's
gating verification: openai-agents 0.20 no longer derives it from
RunConfig.group_id, so both lanes sent nothing. ModelSettings.extra_body
is the one channel all three model classes put on the wire (the bare
kwarg is dropped by LiteLLM; extra_args[extra_body] collides with the
responses model's own parameter — both wire-verified), and it is scoped
to OpenAI destinations: LiteLLM plants extra_body as a literal field in
other providers' bodies, and Anthropic rejects unknown fields — the
anthropic wire test now pins the absence.
Verified before landing: mock-server matrix (OPENAI_BASE_URL + bare
name works through LiteLLM; gpt-named self-hosted models are NOT
bridged off a custom base_url; prompt_cache_key on the wire in every
OpenAI lane with distinct per-conversation keys; anthropic body clean)
and live (sol answers through chat(), gpt-5.4 unchanged, responses()
bare unchanged with the key on its wire).
* refactor: no prefix-triggered direct lane — chat model names are LiteLLM's, verbatim
Ray's ruling on the flip's remaining carve-out: a routing decision must
never hide in a model-name prefix. openai/ now means what LiteLLM says
it means (its openai provider), like every other name on the chat lane —
the grammar is LiteLLM's with zero exceptions.
The two defenses for keeping a direct carve-out had no concrete victim:
debugging isolation (litellm is unavoidable in indexing anyway, and
responses() IS the OpenAI-SDK-native door), and endpoint determinism
(litellm sends chatcmpl for openai-provider models except the gpt-5
bridge, which never fires against a custom base_url — wire-verified).
If a direct escape is ever needed, it will be a declared parameter,
never name grammar.
openai/-prefixed names keep the build-time OPENAI_API_KEY check for
parity with bare names; responses() is untouched (bare and openai/
still drive the OpenAI SDK — LiteLLM cannot speak that protocol).
* fix: third-audit findings — extra_body naming drift, litellm floor hint
The prompt_cache_key delivery channel went through three iterations and
settled on ModelSettings.extra_body; the _conversation_cache_key
docstring still named extra_args from the middle iteration.
The litellm install hint said >=1.30, below both our own pyproject floor
(>=1.84.0) and the floor openai-agents' litellm extra declares (>=1.83).
A user in a broken environment following it would land on a version the
package itself rules out. The hint now matches the declared floor.
* test: pin the bundle door's Agents-SDK model grammar; rename the cache-key test to extra_body
The config bundle hands its model string to the Agents SDK's own
MultiProvider grammar, which refuses unknown prefixes (probe on 0.20:
'anthropic/x' -> UserError: Unknown prefix). _normalize_retrieve_model's
litellm/ spelling is what keeps that door working — a link the existing
self-referential assert (config["model"] == client.retrieve_model)
could not catch. Pinned with a provider-slashed name.
Also renames the cache-key delivery test to its real channel,
extra_body — the extra_args name survived from the superseded delivery
attempt.
* refactor: name the retrieve_model helper for its reason — the Agents SDK's grammar
_normalize_retrieve_model said what it does, not why. The litellm/
spelling exists because the Agents SDK resolves raw model strings with
its own prefix grammar and refuses unknown prefixes — the name now
points at that constraint.
* test: the bundle-grammar test skips without openai-agents, like its file's siblings
Every agents-dependent test in this file importorskips; without the
guard this one errors where the others skip.
* feat: index_model + chat_model — two-knob model surface with full legacy fallback
The documented surface becomes two role knobs: index_model builds the
index, chat_model answers on the chat surfaces. model turns into the
set-both umbrella (its 0.2.8 indexing semantics are a strict subset, so
old configs run unchanged); summary_model and retrieve_model stay
accepted as legacy role names.
Resolution lives in ConfigLoader.load(), the one seam every consumer
already passes through (client, CLI standard/md paths, flash's
summary fallback, tree_optimize's default_model): new names win over
old, specific over general, model sets every role, and code constants
close each chain. The packaged yaml no longer ships model keys — key
presence is what separates a user's explicit choice from a built-in
default, and _validate_keys accepts the five model names explicitly.
Consequences: with no config at all, classic-mode structure extraction
now uses DEFAULT_INDEX_MODEL (gpt-5.6-luna) instead of the yaml's old
gpt-4o-2024-11-20 line (ratified; flash-default users see no change).
client.retrieve_model becomes a read-only alias for client.chat_model.
The resolution matrix test pins one row per released generation:
0.2.8 (model), 0.3.0.dev (model+retrieve_model), 0.2.10.dev (all three
legacy names), the new pair, umbrella-only, and mixed.
* feat: --index-model on the CLI; README flag docs follow
The CLI leads with --index-model; --model stays as its legacy synonym
(the CLI only indexes, so the umbrella and the index role coincide).
The flash branch's summary fallback gains the index position, and the
standard branch now forwards --summary-model, which it had silently
ignored — the flag's help always claimed it worked there. The md
branch's unfiltered model=None no longer clobbers the default: the
resolver treats None as unset.
* feat: default chat model becomes gpt-5.6-sol
Ray's pick for the out-of-box QA default; indexing stays on luna. sol
runs tools on the Responses lane — current litellm bridges chat()
there automatically; older litellm gets the guided 400 naming both
exits.
* chore: trim the model-keys comment to the constraint
The per-generation history lives in 7244ee4's message.
* feat: per-door reasoning passthrough — reasoning_effort / reasoning / thinking
Each chat door gains its own protocol's native thinking control,
forwarded verbatim with no invented vocabulary and no default of ours:
chat_completions(reasoning_effort=...), responses(reasoning={...}),
messages(thinking={...}). Unset sends nothing, so backend defaults
(sol: medium, adaptive) are untouched. chat() stays answer-only.
Delivery channels, each verified: the chat door rides
extra_args["reasoning_effort"] — LiteLLM's own top-level kwarg on
every supported openai-agents version, admitting non-enum values
("none"); newer openai-agents promotes it to the top-level argument
and pops the duplicate. Wire-captured on a mock backend
(/v1/chat/completions body carries it) and coexists with the Claude
cache marker in one dict. The responses door rides
ModelSettings.reasoning — coerced to the typed openai Reasoning object
and forwarded verbatim by the Responses model; the envelope echoes the
caller's dict. The messages door joins the existing anthropic
passthrough dict, asserted through the real tool runner.
LiteLLM semantics observed and accepted as-is: unknown models refuse
the param loudly with LiteLLM's own remedies, and gpt-5.4+ names with
an explicit effort route to /v1/responses even against a custom
api_base (its documented pre-existing arm). The sol-class 400 guidance
now names the third exit — an explicit effort routes on older litellm
releases too.
Cloud chat_completions rejects the new parameter like model/max_turns;
responses()/messages() are local-only already.
* feat: extra_body escape hatch on the three protocol doors
The industry-standard per-request extension channel (openai/anthropic
SDK trio): a dict merged verbatim into the backend request, last, so
caller keys win over SDK-set ones. Routing per door: OpenAI-compatible
destinations get a true body merge (ModelSettings.extra_body / the
anthropic SDK's native extra_body); LiteLLM-routed providers take the
keys as LiteLLM's own top-level kwargs instead, since LiteLLM plants
extra_body as literal fields other providers reject. Cloud mode rejects
it like the other local-only knobs. chat() stays answer-only.
* chore: litellm floor 1.84 -> 1.97.0
1.97.0 is where the unset-effort chatcmpl->responses bridge landed
(responses_api_bridge_check's on_constraint_enforcing_endpoint arm,
A/B-verified against 1.96.2), so sol-class models work through the
chat lane out of the box instead of 400ing until a manual upgrade.
Three spots move together: the pyproject floor, the requirements.txt
CI pin, and the install hint.
* feat: named top_p/max_tokens on the chat door, max_output_tokens on responses
extra_body could not carry these: openai-agents' LitellmModel passes
every ModelSettings sampling field as an explicit keyword and unpacks
extra_args into the same call, so the common knobs collided with a bare
TypeError on the LiteLLM lane (reproduced against a stub — Python call
semantics, callee-independent). Named params ride ModelSettings fields,
the one channel clean on every lane; responses() uses the protocol's
own name (openai_responses maps ModelSettings.max_tokens to
max_output_tokens on the wire) and the envelope now echoes the real
value instead of a constant None. Both caps bound each backend call in
the agent loop, not the whole run — documented. Cloud rejects them like
the other local-only knobs; the long tail (frequency_penalty etc.)
stays extra_body-blocked-loudly on that lane by choice.
* chore: the missing-key error names chat_model as the other exit
The default chat_model is what put keyless users on the OpenAI lane,
so the error now points at the knob that picks a different backend.
|
||
|
|
08a912c69c |
feat: openai-agents becomes a base dependency (#406)
feat: openai-agents becomes a base dependency — the chat engine ships with the SDK chat() is the SDK's front door, and its engine lived behind a vendor-named extra: pip install pageindex could index a document but failed on the first chat call, and chatting with Claude required installing '[openai]'. Measured before moving: the base tree already carries litellm (75 MB) + openai (13 MB), openai-agents adds ~15 MB (agents 8.1 + mcp 1.7 + griffe 1.4 + small pure-python deps), and current litellm's openai range (>=2.20,<3) intersects cleanly with openai-agents' (>=2.45,<3). The [openai] extra stays declared but empty, so existing pip install 'pageindex[openai]' commands keep resolving. Error messages and docstrings drop the extra; requirements.txt gains the dependency, so CI now runs the openai-agents test lane instead of skipping it. Extras now mean exactly one thing: a vendor's own SDK surface ([anthropic] for messages()/tool runner, [claude] for the Claude Agent SDK). |
||
|
|
ec4851059f |
feat: chat() front door and Claude cache marking on LiteLLM routes (#405)
* feat: chat() — the answer-out front door over chat_completions * feat: cache-mark the managed prefix on anthropic-routed LiteLLM models * refactor: cache predicate asks litellm's own provider resolution * feat: extend cache marking to Claude on Bedrock and Vertex — both live-verified * fix: point anthropic-extra users at messages(); guard the two silent vendor chains * docs: the max-tokens table is a closed set — litellm's map prunes EOL'd entries, so it cannot replace it * chore: trim the max-tokens docstring to the contract * docs: state the local text-only history contract — cloud forwards tool turns, local rejects; extra fields drop |
||
|
|
7c6e3c1c8f |
feat: Flash with full optimization becomes the default local indexing mode (#404)
* fix: break the phantom exception chain in _run_sync
Move asyncio.run(coro) out of the except RuntimeError block so real
errors no longer carry a bogus "no running event loop" context in
their traceback.
* feat: Flash with full optimization becomes the default local indexing mode
Every entrance now defaults to Flash with the full optimize pass
(deterministic merge, then LLM expand), replacing the standard LLM-built
tree as the default:
- submit_document(): mode=None now means "flash"; pass mode="standard"
for the LLM-built tree. _index_flash runs optimize="full" with the
expand model = summary_model, and fails fast with the missing key
name(s) via litellm.validate_environment before any work.
- page_index_flash(): optimize takes "full" (default) / "merge" / False;
True is accepted as "full" for compatibility, unknown values raise
instead of silently degrading to merge-only. optimize_expand stays
honored for legacy callers.
- CLI: --mode {flash,standard} replaces --flash (kept as a hidden
compatibility alias that forces flash). --optimize defaults to full in
flash mode with an `off` choice; explicitly passing it outside flash
still errors. Standard-only tuning flags (--toc-check-pages,
--max-*-per-node, --if-add-*) now error in flash mode instead of being
silently ignored, mirroring the existing flash-only flag errors. The
key pre-check runs only when an LLM will actually be called, so
--no-summary --optimize off|merge works keyless. Output drops the
_structure_flash suffix — always <name>_structure.json.
On the Disney earnings PDF the optimized default is also faster than
unoptimized flash (fewer nodes to summarize) and fixes hierarchy
mistakes; both modes emit identical schemas end to end.
Docs updated to match (mode flag, defaults, LLM usage honesty); tests
pin the new defaults: stored mode == "flash", optimize passthrough, and
the unknown-optimize rejection.
|
||
|
|
ec342c4e28 |
fix: post-merge review fixes for v0.2.10
- Conformant responses() envelope: official output (model items only) + items (full transcript for round-trip), usage details aggregated across turns; verified against real OpenAI API - Python floor corrected to >=3.10 (litellm stable requires it) - Three stale anthropic>=0.84.0 hints updated to 0.108.0 - Image-stub behavior disclosed on tool-builder docstrings - Non-essential comments trimmed |
||
|
|
4e41acdc68 |
feat: agent tools and local chat for the PageIndex SDK (v0.2.10) (#396)
* feat: agent tools — the cloud MCP tool contract on the client Four new client methods make PageIndex documents available to agent frameworks, in both modes, with the mode decided solely by the client constructor: - agent_tools(): plain functions (browse_documents, get_document, get_document_structure, get_page_content) matching the PageIndex cloud MCP server's tools/list — same names, schemas, descriptions, and JSON response envelopes — so agent prompts port unchanged between the cloud MCP connection and these in-process tools. Tools never raise; errors come back in the same envelope. remove_document ships behind include_management=False. - as_openai_tools(): the same tools wrapped for the OpenAI Agents SDK. - as_claude_mcp(): one mcp_servers entry for the Claude Agent SDK — cloud clients get the remote MCP config (the framework connects to api.pageindex.ai/mcp and discovers the full cloud tool set), local clients get an in-process SDK MCP server. - agent_instructions(doc_id=None): orchestration guidance for the agent's system prompt; doc_id (same shape as chat_completions) appends the target documents. submit_document() gains wait=True: poll get_document status until completed, raise on failed or after 30 minutes — the manual polling loop every cloud caller writes today spins forever on a failed document. Neither framework becomes a dependency: imports happen at call time with actionable errors, and pageindex[openai] / pageindex[claude] extras are floor-only pins. tests/data/cloud_mcp_contract.json freezes the tool contract; a parity test guards against drift. 36 new tests (95 total), plus a live OpenAI Agents SDK run over a seeded local store verifying the structure-first navigation flow end to end. * fix: agent tools review — next_steps order, resolve caching, error semantics - Large-doc next_steps now says structure-first, consistent with tool descriptions and agent instructions - _remove_document fetches document list once instead of per-name - call_tool returns error envelope for unknown names instead of raising - _not_ready_error timed_out flag reflects actual wait outcome - openai_agents.py docstring corrected to match default (FunctionTools) - Removed unused ModelSettings import from demo * fix: agent tools review 2 — bridge thread safety, browse paging, metadata merge - McpBridge reads session/protocol headers under the lock (now RLock: _ensure_initialized posts while holding it). openai-agents runs sync tools on threads and executes parallel tool calls concurrently, so bridge functions genuinely race; a torn read sent a new session id with a stale protocol header. Measured: one session expiry under 8 threads cost 4 initializations before, minimal 2 after. - Session-expiry retry also resets the negotiated protocol version, so the re-handshake carries no stale MCP-Protocol-Version header. - browse_documents time sort pages list_documents natively instead of fetching the whole library to slice one window (relevance still needs the full list for scoring). - _await_completion: a status refetch that nulls out metadata no longer clobbers the listing's copy (setdefault was a no-op on existing None). - Structure tool reads the raw stored tree via a named LocalAPI raw_tree() seam instead of reaching into _api._store internals; drop the redundant deepcopy before _format_structure (store re-reads from disk, formatting builds fresh containers). - Shared pageindex/_version.py replaces _sdk_version duplicated in mcp_bridge and the Claude integration. Left as-is after source verification against the cloud MCP: first-page budget bypass, pageNum falsy-zero, and the page-gap fallback text are letter-for-letter cloud behavior — parity wins over local repair. * fix: agent tools review 3 — page-span cap, duplicate names, wait resilience, contract drift - _parse_page_spec bounds the requested span arithmetically (10k pages) before materializing it; pages="1-1000000000" previously expanded to a billion integers inside the caller's process. - Local submit_document uniquifies document names the way the cloud upload does (taken name -> _1.._99, then reject with the cloud's own message). Same-name duplicates broke name-addressed tools: resolution always picks the newest, so older duplicates were unreachable. - agent_instructions(doc_id=...) now fails loud when the pinned doc's name is shadowed by a newer same-name document (legacy stores predate the rename) — it previews resolution with the same _resolve_document the tools use, so the check cannot drift from actual behavior. - submit_document(wait=True) tolerates transient network errors, not just API errors; a dropped connection at minute 25 of a 30-minute wait no longer kills it. Third strike wraps into PageIndexAPIError per the documented contract. - The live contract-parity test compares full per-param schemas, not just names and descriptions. It immediately caught real drift the shallow check had been passing: the server now emits nullables as anyOf unions and stamps MAX_SAFE_INTEGER maxima on offset/part. Contract and snapshot updated to the served wire form; _annotation_for learned anyOf so bridge signatures stay Optional[str] instead of degrading to Any. Adjudicated, not changed: the allowed_tools wildcard example stays (docstring advice covers scoping; Ray's call), and raw-length response accounting stays (letter-for-letter cloud behavior, parity wins). * feat: surface the stored document name from submit_document Compute PR #558 makes /doc/ return {"doc_id", "name"} carrying the post-dedup-rename name. Mirror it end to end: local submit returns the stored name, the client warns when it differs from the uploaded file name (read via .get so older cloud servers stay compatible), the local name-exhaustion check runs before indexing instead of after the LLM spend, and the demo caches doc_id in a file instead of name-matching — a renamed document made the name lookup re-index on every run. * fix: add missing page_list kwarg in duplicate-name test mock * revert: keep README.md unchanged from main — SDK section deferred * feat: serve cloud agent instructions live from the MCP server The cloud MCP server publishes its agent instructions in the initialize result, adapted to each key's tool set. agent_instructions() previously returned the SDK's local-subset text in both modes — a silently forked copy that lacks the guidance for cloud-only tools (search_documents escalation, folders, images) and drifts as the server's prompt evolves. Cloud clients now serve the server's live instructions, captured from the initialize handshake on a per-client bridge shared with agent_tools() (one session, no extra request). An empty server response raises instead of silently substituting the subset text — same posture as the annotation-regression guard. The local constant stays as the honest subset for the in-process tools, with its provenance noted and a consistency test that every tool it names exists in the local registry. * fix: local relevance sort answers honestly instead of imitating sort="relevance" is cloud-side semantic ranking; the local substring imitation could satisfy the letter of the interface while silently missing semantically relevant documents. Per the honest-subset rule (same treatment as folders), local now returns the "not available here" envelope for sort="relevance" or a stray query, and the local instructions steer discovery through name/description matching plus full-library paging instead of prescribing a capability that does not exist here. The tool schema keeps the cloud contract verbatim, like folder_id: honesty lives in the runtime answer, not a forked contract. * docs: note the cloud+Claude instructions duplication trade-off in as_claude_mcp * fix: unsupported-capability envelopes say local-mode-yet, point to cloud "Not available here" read as a broken feature; the honest framing is that folders and semantic ranking exist on PageIndex cloud and are not in local mode yet. Both envelopes now say so and name the cloud client in next_steps, so agents relay an accurate story to the user. * fix: local tool descriptions pre-announce cloud-only capabilities The cloud-verbatim browse_documents description invites sort="relevance" and folder drilling, so a local agent's first semantic search attempt was a guaranteed dead end discovered only from the runtime error envelope. Local registration now appends a LOCAL MODE note to the description — the agent learns what is cloud-only before calling; the runtime envelope stays as the backstop for prompts that ignore descriptions. The cloud-facing contract stays byte-verbatim. * refactor: localized tool guidance replaces the appended LOCAL MODE note Appending a retraction to the cloud-verbatim description left the model parsing an instruction and its negation — and kept the cloud text recommending search_documents and get_folder_structure, tools that are not registered locally (get_page_content likewise pointed at get_document_image). Guidance now adapts to the local surface the way AGENT_INSTRUCTIONS already does: schema structure stays byte-identical to the contract (mechanically asserted by a strip-descriptions test), while local description strings teach only what works here and point to PageIndex cloud for the rest. A dead-reference test forbids local guidance from naming tools outside the local registry, so a contract refresh that reintroduces a cloud-only reference fails loudly. * feat: hide cloud-only parameters from the local tool surface folder_id, sort, query, and recursive were exposed locally with localized "cloud-only" descriptions, leaving the dead-end calls expressible and discovered at runtime. Schema constraints beat guidance: the local surface now serves the contract minus these parameters, so strict-schema frameworks make the calls inexpressible and a prompt that insists on sort="relevance" degrades to the bare call (the correct local behavior) instead of an error round-trip. The implementations still accept the hidden parameters and answer with the guided "works on PageIndex cloud" envelope — the backstop for direct call_tool callers and hosts without schema enforcement. wait_for_completion stays: seeded or torn stores can hold documents that are genuinely not completed. The structural guard now asserts the local schema equals the contract minus the documented hidden set, descriptions aside. * fix: incremental-review findings — bridge cache, guards, envelope drift Three independent review passes over the agent-instructions increment surfaced six fixes: - The per-client bridge moved off the instance into a weak-keyed, lock-guarded module cache: cloud clients stay picklable (threading.RLock no longer rides on the client) and concurrent first calls can no longer construct duplicate bridges/sessions. - Blank or non-string initialize.instructions now hit the same honest error as a missing one — a whitespace-only or structured value could previously become the system prompt (or crash the doc_id append with a raw TypeError). - The invalid-sort envelope no longer prescribes sort="relevance" — the one error text that still taught the cloud-only value it would then reject. - "Page through the rest of the library" is emitted only when has_more is true; a fully-listed library no longer instructs a pointless call. - The mandatory full-library paging step now says limit: 50 — 6 calls instead of 30 on a 300-document library. - Docstrings and comments rescoped to what is actually true: the never-raise contract covers invocations the signatures accept (unknown params fail at the Python boundary; call_tool answers them with the guided envelope), recursive is accepted as the identity rather than errored, lenient framework arg models drop hidden params pre-call, and the module header no longer claims full schema parity. The capability-phrase guard now covers every local docstring, not just browse_documents. * chore: keep the demo's doc_id cache file out of the repo * test: live envelope field-parity guard against cloud response drift The frozen contract guards tools/list, but the response envelopes the local tools emit were hand-built to mirror the cloud's and had no drift detector. A key-gated live test now asserts every field local emits exists in the live cloud response for the analogous call (top-level keys, next_steps, document entries, structure nodes, content entries). Guidance wording is deliberately localized and not compared. Verified green against the live server: local and cloud field structures currently match exactly. * feat: local chat — three protocol surfaces over the agent tools (v0.2.10) Local mode gains managed document QA: an agent over the #393 local tool set, reachable through three wire protocols, each 1:1 with the backend and with no translation layer. - chat_completions(): standard chat.completions semantics on any OpenAI-compatible backend (openai-agents engine). Final answer only, cross-turn aggregated usage, streaming as text pieces or chunk dicts (the existing cloud signature, now implemented locally; model and max_turns are local-only additions). - responses(): the agentic surface — OpenAI Responses format, the tool process is standard output items, streaming forwards native events (tool outputs emitted as response.output_item.done, the way the platform streams its own server-side tools). Round-tripping output into the next input keeps provider prompt-cache prefix continuity and the agent's memory — live-verified: the follow-up call answered from round-tripped tool output with zero new tool calls. - messages(): Anthropic-native via the SDK's own tool runner (new pageindex[anthropic] extra, floor 0.68.0 verified for tool_runner/beta_tool(input_schema)). tool_use/tool_result round-trip is the format's native behavior; the envelope is the final message with aggregated usage plus the full new-turn sequence; the managed system blocks carry cache_control breakpoints. Shared skeleton: thin chat header + the local AGENT_INSTRUCTIONS (caller system content is appended, not rejected), the doc_id targeting block as a leading context item (factored out of build_agent_instructions), read-only toolset, structural-only validation (no arbitrary caps — backend limits govern), sampling params passed through, per-run tracing disabled, enable_citations rejected as cloud-only. Design basis is industry-standard formats rather than the cloud chat endpoint; responses()/messages() raise on cloud clients until the cloud converges. Tests run the real engines against scripted backends (a Model fake for openai-agents, a mock HTTP transport under the real anthropic SDK) with real tool execution against a seeded store, including the round-trip prefix-extension assertions on both engines. * fix: local-chat review findings — truncation, serialization, streams Three independent review passes (bug scan, claims-vs-code, adversarial runtime probes) over the local-chat increment; every fix below was reproduced before being fixed. messages(): - A max_turns cut no longer duplicates the final assistant turn: the runner has already appended it when iterations exhaust, so the round-trip history carried a duplicate tool_use id and ended on an unanswered tool_use — a guaranteed 400 on continuation. The append now keys on stop_reason, and truncation reads natively as stop_reason: "tool_use" with a continuable history. - The envelope is JSON-serializable end to end: runner-stored turns carry pydantic content blocks; everything is dumped to plain dicts, excluding SDK-internal __api_exclude__ fields (parsed_output) that the API rejects on round-trip. - Bounded by default (max_iterations 10, like the OpenAI surfaces); usage aggregation now preserves the final turn's native fields and sums the token counters None-safely; empty caller system strings are skipped; non-dict message entries and bad doc_id types raise PageIndexAPIError; anthropic < 0.68 gets an actionable version error; the doc block no longer spends a cache_control breakpoint. chat_completions()/responses(): - MaxTurnsExceeded wraps into PageIndexAPIError on all four run paths. - responses(stream=True) is one logical response: per-turn backend lifecycle events are collapsed (a canonical consumer previously stopped at turn 1's response.completed and never saw the answer), sequence numbers are reassigned monotonically, and the synthesized tool-output event carries output_index/sequence_number. - The responses envelope carries the real request surface (instructions, the actual function tool definitions, tool_choice, parallel_tool_calls, error/incomplete_details). - RunConfig(group_id) pins a stable prompt_cache_key: openai-agents otherwise stamps each run with a fresh key, tagging round-tripped prefixes as different cache groups and defeating the feature the round-trip exists for. - Abandoning a stream now cancels the run: a watchdog task lets the cancellation land even while the pump awaits the backend, and the per-call AsyncOpenAI client is closed before its loop ends (fixes "Task exception was never retrieved" noise). The opening role chunk is emitted even for empty outputs; empty responses() input and enable_citations-before-extra ordering fixed. Docs rescoped to what is true: finish_reason/status reflect loop completion on the OpenAI surfaces (the engine does not surface per-turn backend reasons); chat streaming yields visible narration including pre-tool text; messages(stream=True) forwards the Anthropic SDK's native event objects (not wire-verbatim); the doc block is a leading conversation item on OpenAI surfaces and a system block on messages(). Tests: 25 in the file (11 new), with per-extra skip sections so a machine with only one framework still covers the other surface; without-frameworks matrix re-verified; live smoke re-run green with a clean exit. * feat: as_anthropic_tools — Anthropic tool-runner export, both modes Fills the last cell of the agent-connection matrix: users driving their own anthropic tool_runner loop get runnable tools directly. Cloud wraps the live MCP tool set with input schemas passing through verbatim (MCP inputSchema is the Messages API schema shape); local exposes the same set messages() runs internally. The beta_tool wrapping moves from local_chat into integrations/anthropic_sdk.py, parallel to openai_agents.py, and messages() now consumes the shared builder. agent_tools grows _bridge_invoker/_read_only_tools so the plain-function and beta_tool cloud paths share invocation containment and the read-only gate. * fix: as_anthropic_tools review findings — async flavor, schema isolation Adversarial + best-practice review of |
||
|
|
d375c00a5a |
feat: local mode for the PageIndex SDK (v0.2.9) (#389)
* feat: add local mode to the PageIndex SDK client
One PageIndexClient, two backends. With api_key: the 0.2.x cloud SDK,
request for request (with the reviewed fixes: bounded timeouts on JSON
endpoints, none on uploads, URL-encoded ids, 401 key hint, empty DELETE
body tolerated). Without api_key: the same methods run locally —
page_index builds the tree in submit_document (mode="flash" uses
PageIndex Flash), documents are stored as plain JSON per doc under
storage_path, submit_query is LLM tree search with retrieve_model, and
chat_completions answers over the retrieved nodes with OpenAI-style
responses and streaming.
Local responses mirror the cloud wire shapes verified against the server
source: tree nodes rename start_index to page_index and drop end_index,
a non-leaf summary becomes prefix_summary, and the metadata/list/delete/
retrieval envelopes match key for key. Cloud-only features (folders,
beta_headers, enable_citations) raise instead of pretending.
Replaces the demo-only workspace client (index/get_document_structure/
get_page_content had no real users) and its retrieve.py helpers.
page_index_main gains an optional logger param so the SDK can keep
./logs out of the caller's working directory; pymupdf import is now lazy
(only the optional PyMuPDF parser path needs it).
* chore: package pageindex 0.3.0.dev4 for PyPI
Poetry packaging for the combined SDK + local pipeline: every production
import is a declared dependency (openai and requests join requirements.txt
for the same reason), config.yaml and the flash data tables ship in the
wheel, the benchmark PNG does not. pymupdf drops to an optional note now
that its import is lazy. dev4 follows the already-published 0.3.0.dev1-3;
pip still resolves plain 'pip install pageindex' to 0.2.8 until a final
0.3.0 — install with --pre.
* docs: add SDK section to README; move the agentic demo onto the SDK
The demo keeps its flow and the post-cutoff demo paper, swapping the
removed workspace client for PageIndexClient local mode (list_documents
for the doc-id cache, get_tree/get_ocr behind the agent tools). The old
examples/workspace JSONs demoed the removed format and go with it.
* refactor: rebuild cloud_api on the 0.2.8 client text
Ray's rule for the cloud half: his 0.2.8 code is the base; a Kylin-lineage
change survives only when strictly better — invisible on healthy traffic
while fixing a real failure mode. Kept under that bar: request timeouts
(dead connections hung forever; uploads still pass none, exactly like
0.2.8), the upload handle closed via with (leaked on request errors),
URL-encoded path ids (a crafted id could reroute the URL), the
empty-DELETE-body guard, and stream hardening (choices guard,
response.close in finally). Reverted as not strictly better: the
lowercase summary param (the server accepts both spellings), the 401
message hint (visible text change; AUTH_HINT dropped from errors.py), and
the _request/requests.request reorganization — every method body,
docstring, and section comment is 0.2.8's text again.
diff -w against ../pageindex_sdk/pageindex/client.py now reads as that
surgical patch plus plumbing: CloudAPI reads BASE_URL/api_key through the
owning client, PageIndexAPIError comes from errors.py, and
is_retrieval_ready lives verbatim on PageIndexClient shared by both
modes. A mocked-requests harness driving 0.2.8 and this file through 19
identical calls shows the only remaining request-level difference is
timeout.
* refactor: drop the local retrieval endpoints — cloud-only, deprecated
Ray: the cloud already marks POST /retrieval/ and GET /retrieval/{id}/
deprecated in favor of chat completions, so local mode should not grow a
fresh implementation of a retiring surface. submit_query/get_retrieval
now raise in local mode with a pointer to chat_completions; cloud mode is
untouched (the endpoint still works there and 0.2.8 code keeps running).
The tree search that backed them stays as chat_completions' retrieval
engine; the retrievals/ storage goes away. This also closes the one real
cross-mode parity gap — the retrieved_nodes inner shape — by removing
its local half.
* feat: manifest.json — one-file document listings for the local store
Ray wanted a metafile that shows every document in one place instead of
per-directory reads. It is a cache, never a second source of truth:
writers update it best-effort after save/delete (atomic replace, no
locks), and list_metas trusts it only while its id set matches the docs/
directory names — documents are immutable, so matching names imply valid
content. Any mismatch (lost concurrent update, crash, corrupt or deleted
manifest) rebuilds it from the doc.json files, reading only the missing
entries. Incomplete dirs (no doc.json) stay invisible and are never
recorded, so a save that completes later is still picked up.
1000-doc listing: 37ms of per-dir reads -> 2.3ms warm (scandir names +
one manifest read); one-time rebuild 160ms.
* fix: align local doc_id prefix and createdAt format with the cloud
Local doc ids now carry the cloud's pi- namespace prefix (random token
stays uuid4 hex — nothing parses cuid internals, the prefix is the
contract; chat ids already mirrored chatcmpl-). createdAt now matches
the server byte for byte: the cloud emits the DB datetime's bare
isoformat — naive UTC, second precision — while we emitted microseconds
plus +00:00.
* feat: optional metadata tags on submit_document (both modes)
Ray's call after the alignment review: the metadata key in tree/OCR
envelopes and list entries should carry real data, and the only honest
way is exposing the field the server already accepts. Cloud mode
forwards it as the existing metadata form field; local mode validates it
early (a JSON-serializable dict, checked before any LLM spend), stores
it in doc.json, and returns it from the same three places. get_document
still omits it, mirroring the server, whose metadata-endpoint SQL never
selects that column. Scope stays deliberately narrow: set at submit and
read back — no metadata_filter, no update API.
* docs: state that createdAt is UTC and show how to localize it
The value is naive UTC in both modes (the cloud column is timestamp
DEFAULT CURRENT_TIMESTAMP on a UTC server, emitted via bare isoformat).
Wall-clock display is the consumer's layer: emitting local time under
the same format would silently change meaning per machine, and adding an
offset marker would break both the byte-format parity and string-order
sorting.
* fix: createdAt carries milliseconds, matching the cloud's datetime(3)
The earlier second-precision alignment was reasoned from the postgres
schema file, but production is MySQL (DATABASE_BACKEND defaults to
mysql) and its FilePageIndex.createdAt is datetime(3) DEFAULT
CURRENT_TIMESTAMP(3) — so the server isoformat()s a millisecond-
precision naive-UTC datetime, emitting .XXX000 fractions (bare seconds
only when the millisecond happens to be zero). Local now generates
through the same mechanism. The docs' 2024-01-15T10:30:00.000Z sample is
a JS-style string Python isoformat cannot produce — not evidence.
* fix: createdAt at millisecond precision, matching the timestamp(3) column
The cloud column is timestamp(3); its datetimes render through bare
isoformat() as six fractional digits ending in 000 (or no fraction when
the millisecond is exactly zero). Truncate to the millisecond and render
the same way, replacing the second-precision guess from
|
||
|
|
b723c9f0a7 |
ci: run the test suite on every push and pull request
Publishing was the only automation touching code: a tag builds and ships to PyPI without ever running a test, and pull requests get no checks at all. This runs pytest on a small matrix — Python 3.10 (the floor) and 3.13, each with and without the agent frameworks installed, so the lazy-import contract (the package must work with neither framework present) is enforced rather than assumed. |
||
|
|
d5c4e62c20 |
Exclude common words from roman numeral conversion (#387)
Update PageIndex Flash |
||
|
|
933df35b2f |
Add a Flash summary flag (#386)
* Update PageIndex Flash * Add a summary flag |
||
|
|
502b763876 | Update README (#385) | ||
|
|
3c4c7e8353 |
Use embedded PDF bookmarks in Flash trees (#384)
* Update PageIndex Flash * Add embedded ToC CLI flag and README note |
||
|
|
fb3f6e441a | Fail fast on rejected keys and unknown models (#381) | ||
|
|
03cf0fa0ef | Sync PageIndex Flash from private branch (#380) | ||
|
|
d409566464 | Sync PageIndex Flash from private branch (#379) | ||
|
|
1b2fdadf51 | Keep merge inside --optimize (#377) | ||
|
|
2a29ac5aa0 |
Merge pull request #375 from VectifyAI/fix/default-merge
Refine the Flash tree with merge |
||
|
|
9e789babf3 | Add benchmark to Flash README (#376) | ||
|
|
e17dac2053 | Strip internal marker from optimize output | ||
|
|
4a116daa2f | Replace thinning with merge | ||
|
|
3f33a53b50 | Add tree optimization and node summaries for Flash | ||
|
|
afb3100232 |
Merge pull request #374 from VectifyAI/fix/flash-requirements
Add Flash dependencies to requirements |
||
|
|
6e248f8f05 | Add Flash dependencies to requirements | ||
|
|
d178bc557d |
Merge pull request #372 from VectifyAI/openai-sdk-fast-path
Use OpenAI SDK directly for OpenAI models |
||
|
|
548d3201b7 | Use OpenAI SDK directly for OpenAI models | ||
|
|
66d5b9ac9b |
Merge pull request #371 from VectifyAI/fix-flash-readme
Add --flash CLI flag and update README |
||
|
|
faf3c57fc8 | Add --flash flag to CLI | ||
|
|
a77301e5c2 |
Merge pull request #370 from VectifyAI/fix-flash-readme
Update Flash README links |
||
|
|
165fcea3db | Update Flash links to main branch | ||
|
|
c9a01fad13 |
Merge pull request #369 from VectifyAI/pageindex-flash
Add PageIndex Flash |
||
|
|
fffc2c5512 | Add PageIndex Flash | ||
|
|
0e4b68c92d |
Merge pull request #367 from VectifyAI/add-summary-model-config
Add summary_model config |
||
|
|
acfcde2c58 | Remove summary_model from page_index() signature | ||
|
|
820e8bf103 | Use summary_model for doc description, expose in page_index() |