mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-10-02 07:44:37 +08:00
feat: agent tools and local chat for the PageIndex SDK (v0.2.10) (#396)
* feat: agent tools — the cloud MCP tool contract on the client Four new client methods make PageIndex documents available to agent frameworks, in both modes, with the mode decided solely by the client constructor: - agent_tools(): plain functions (browse_documents, get_document, get_document_structure, get_page_content) matching the PageIndex cloud MCP server's tools/list — same names, schemas, descriptions, and JSON response envelopes — so agent prompts port unchanged between the cloud MCP connection and these in-process tools. Tools never raise; errors come back in the same envelope. remove_document ships behind include_management=False. - as_openai_tools(): the same tools wrapped for the OpenAI Agents SDK. - as_claude_mcp(): one mcp_servers entry for the Claude Agent SDK — cloud clients get the remote MCP config (the framework connects to api.pageindex.ai/mcp and discovers the full cloud tool set), local clients get an in-process SDK MCP server. - agent_instructions(doc_id=None): orchestration guidance for the agent's system prompt; doc_id (same shape as chat_completions) appends the target documents. submit_document() gains wait=True: poll get_document status until completed, raise on failed or after 30 minutes — the manual polling loop every cloud caller writes today spins forever on a failed document. Neither framework becomes a dependency: imports happen at call time with actionable errors, and pageindex[openai] / pageindex[claude] extras are floor-only pins. tests/data/cloud_mcp_contract.json freezes the tool contract; a parity test guards against drift. 36 new tests (95 total), plus a live OpenAI Agents SDK run over a seeded local store verifying the structure-first navigation flow end to end. * fix: agent tools review — next_steps order, resolve caching, error semantics - Large-doc next_steps now says structure-first, consistent with tool descriptions and agent instructions - _remove_document fetches document list once instead of per-name - call_tool returns error envelope for unknown names instead of raising - _not_ready_error timed_out flag reflects actual wait outcome - openai_agents.py docstring corrected to match default (FunctionTools) - Removed unused ModelSettings import from demo * fix: agent tools review 2 — bridge thread safety, browse paging, metadata merge - McpBridge reads session/protocol headers under the lock (now RLock: _ensure_initialized posts while holding it). openai-agents runs sync tools on threads and executes parallel tool calls concurrently, so bridge functions genuinely race; a torn read sent a new session id with a stale protocol header. Measured: one session expiry under 8 threads cost 4 initializations before, minimal 2 after. - Session-expiry retry also resets the negotiated protocol version, so the re-handshake carries no stale MCP-Protocol-Version header. - browse_documents time sort pages list_documents natively instead of fetching the whole library to slice one window (relevance still needs the full list for scoring). - _await_completion: a status refetch that nulls out metadata no longer clobbers the listing's copy (setdefault was a no-op on existing None). - Structure tool reads the raw stored tree via a named LocalAPI raw_tree() seam instead of reaching into _api._store internals; drop the redundant deepcopy before _format_structure (store re-reads from disk, formatting builds fresh containers). - Shared pageindex/_version.py replaces _sdk_version duplicated in mcp_bridge and the Claude integration. Left as-is after source verification against the cloud MCP: first-page budget bypass, pageNum falsy-zero, and the page-gap fallback text are letter-for-letter cloud behavior — parity wins over local repair. * fix: agent tools review 3 — page-span cap, duplicate names, wait resilience, contract drift - _parse_page_spec bounds the requested span arithmetically (10k pages) before materializing it; pages="1-1000000000" previously expanded to a billion integers inside the caller's process. - Local submit_document uniquifies document names the way the cloud upload does (taken name -> _1.._99, then reject with the cloud's own message). Same-name duplicates broke name-addressed tools: resolution always picks the newest, so older duplicates were unreachable. - agent_instructions(doc_id=...) now fails loud when the pinned doc's name is shadowed by a newer same-name document (legacy stores predate the rename) — it previews resolution with the same _resolve_document the tools use, so the check cannot drift from actual behavior. - submit_document(wait=True) tolerates transient network errors, not just API errors; a dropped connection at minute 25 of a 30-minute wait no longer kills it. Third strike wraps into PageIndexAPIError per the documented contract. - The live contract-parity test compares full per-param schemas, not just names and descriptions. It immediately caught real drift the shallow check had been passing: the server now emits nullables as anyOf unions and stamps MAX_SAFE_INTEGER maxima on offset/part. Contract and snapshot updated to the served wire form; _annotation_for learned anyOf so bridge signatures stay Optional[str] instead of degrading to Any. Adjudicated, not changed: the allowed_tools wildcard example stays (docstring advice covers scoping; Ray's call), and raw-length response accounting stays (letter-for-letter cloud behavior, parity wins). * feat: surface the stored document name from submit_document Compute PR #558 makes /doc/ return {"doc_id", "name"} carrying the post-dedup-rename name. Mirror it end to end: local submit returns the stored name, the client warns when it differs from the uploaded file name (read via .get so older cloud servers stay compatible), the local name-exhaustion check runs before indexing instead of after the LLM spend, and the demo caches doc_id in a file instead of name-matching — a renamed document made the name lookup re-index on every run. * fix: add missing page_list kwarg in duplicate-name test mock * revert: keep README.md unchanged from main — SDK section deferred * feat: serve cloud agent instructions live from the MCP server The cloud MCP server publishes its agent instructions in the initialize result, adapted to each key's tool set. agent_instructions() previously returned the SDK's local-subset text in both modes — a silently forked copy that lacks the guidance for cloud-only tools (search_documents escalation, folders, images) and drifts as the server's prompt evolves. Cloud clients now serve the server's live instructions, captured from the initialize handshake on a per-client bridge shared with agent_tools() (one session, no extra request). An empty server response raises instead of silently substituting the subset text — same posture as the annotation-regression guard. The local constant stays as the honest subset for the in-process tools, with its provenance noted and a consistency test that every tool it names exists in the local registry. * fix: local relevance sort answers honestly instead of imitating sort="relevance" is cloud-side semantic ranking; the local substring imitation could satisfy the letter of the interface while silently missing semantically relevant documents. Per the honest-subset rule (same treatment as folders), local now returns the "not available here" envelope for sort="relevance" or a stray query, and the local instructions steer discovery through name/description matching plus full-library paging instead of prescribing a capability that does not exist here. The tool schema keeps the cloud contract verbatim, like folder_id: honesty lives in the runtime answer, not a forked contract. * docs: note the cloud+Claude instructions duplication trade-off in as_claude_mcp * fix: unsupported-capability envelopes say local-mode-yet, point to cloud "Not available here" read as a broken feature; the honest framing is that folders and semantic ranking exist on PageIndex cloud and are not in local mode yet. Both envelopes now say so and name the cloud client in next_steps, so agents relay an accurate story to the user. * fix: local tool descriptions pre-announce cloud-only capabilities The cloud-verbatim browse_documents description invites sort="relevance" and folder drilling, so a local agent's first semantic search attempt was a guaranteed dead end discovered only from the runtime error envelope. Local registration now appends a LOCAL MODE note to the description — the agent learns what is cloud-only before calling; the runtime envelope stays as the backstop for prompts that ignore descriptions. The cloud-facing contract stays byte-verbatim. * refactor: localized tool guidance replaces the appended LOCAL MODE note Appending a retraction to the cloud-verbatim description left the model parsing an instruction and its negation — and kept the cloud text recommending search_documents and get_folder_structure, tools that are not registered locally (get_page_content likewise pointed at get_document_image). Guidance now adapts to the local surface the way AGENT_INSTRUCTIONS already does: schema structure stays byte-identical to the contract (mechanically asserted by a strip-descriptions test), while local description strings teach only what works here and point to PageIndex cloud for the rest. A dead-reference test forbids local guidance from naming tools outside the local registry, so a contract refresh that reintroduces a cloud-only reference fails loudly. * feat: hide cloud-only parameters from the local tool surface folder_id, sort, query, and recursive were exposed locally with localized "cloud-only" descriptions, leaving the dead-end calls expressible and discovered at runtime. Schema constraints beat guidance: the local surface now serves the contract minus these parameters, so strict-schema frameworks make the calls inexpressible and a prompt that insists on sort="relevance" degrades to the bare call (the correct local behavior) instead of an error round-trip. The implementations still accept the hidden parameters and answer with the guided "works on PageIndex cloud" envelope — the backstop for direct call_tool callers and hosts without schema enforcement. wait_for_completion stays: seeded or torn stores can hold documents that are genuinely not completed. The structural guard now asserts the local schema equals the contract minus the documented hidden set, descriptions aside. * fix: incremental-review findings — bridge cache, guards, envelope drift Three independent review passes over the agent-instructions increment surfaced six fixes: - The per-client bridge moved off the instance into a weak-keyed, lock-guarded module cache: cloud clients stay picklable (threading.RLock no longer rides on the client) and concurrent first calls can no longer construct duplicate bridges/sessions. - Blank or non-string initialize.instructions now hit the same honest error as a missing one — a whitespace-only or structured value could previously become the system prompt (or crash the doc_id append with a raw TypeError). - The invalid-sort envelope no longer prescribes sort="relevance" — the one error text that still taught the cloud-only value it would then reject. - "Page through the rest of the library" is emitted only when has_more is true; a fully-listed library no longer instructs a pointless call. - The mandatory full-library paging step now says limit: 50 — 6 calls instead of 30 on a 300-document library. - Docstrings and comments rescoped to what is actually true: the never-raise contract covers invocations the signatures accept (unknown params fail at the Python boundary; call_tool answers them with the guided envelope), recursive is accepted as the identity rather than errored, lenient framework arg models drop hidden params pre-call, and the module header no longer claims full schema parity. The capability-phrase guard now covers every local docstring, not just browse_documents. * chore: keep the demo's doc_id cache file out of the repo * test: live envelope field-parity guard against cloud response drift The frozen contract guards tools/list, but the response envelopes the local tools emit were hand-built to mirror the cloud's and had no drift detector. A key-gated live test now asserts every field local emits exists in the live cloud response for the analogous call (top-level keys, next_steps, document entries, structure nodes, content entries). Guidance wording is deliberately localized and not compared. Verified green against the live server: local and cloud field structures currently match exactly. * feat: local chat — three protocol surfaces over the agent tools (v0.2.10) Local mode gains managed document QA: an agent over the #393 local tool set, reachable through three wire protocols, each 1:1 with the backend and with no translation layer. - chat_completions(): standard chat.completions semantics on any OpenAI-compatible backend (openai-agents engine). Final answer only, cross-turn aggregated usage, streaming as text pieces or chunk dicts (the existing cloud signature, now implemented locally; model and max_turns are local-only additions). - responses(): the agentic surface — OpenAI Responses format, the tool process is standard output items, streaming forwards native events (tool outputs emitted as response.output_item.done, the way the platform streams its own server-side tools). Round-tripping output into the next input keeps provider prompt-cache prefix continuity and the agent's memory — live-verified: the follow-up call answered from round-tripped tool output with zero new tool calls. - messages(): Anthropic-native via the SDK's own tool runner (new pageindex[anthropic] extra, floor 0.68.0 verified for tool_runner/beta_tool(input_schema)). tool_use/tool_result round-trip is the format's native behavior; the envelope is the final message with aggregated usage plus the full new-turn sequence; the managed system blocks carry cache_control breakpoints. Shared skeleton: thin chat header + the local AGENT_INSTRUCTIONS (caller system content is appended, not rejected), the doc_id targeting block as a leading context item (factored out of build_agent_instructions), read-only toolset, structural-only validation (no arbitrary caps — backend limits govern), sampling params passed through, per-run tracing disabled, enable_citations rejected as cloud-only. Design basis is industry-standard formats rather than the cloud chat endpoint; responses()/messages() raise on cloud clients until the cloud converges. Tests run the real engines against scripted backends (a Model fake for openai-agents, a mock HTTP transport under the real anthropic SDK) with real tool execution against a seeded store, including the round-trip prefix-extension assertions on both engines. * fix: local-chat review findings — truncation, serialization, streams Three independent review passes (bug scan, claims-vs-code, adversarial runtime probes) over the local-chat increment; every fix below was reproduced before being fixed. messages(): - A max_turns cut no longer duplicates the final assistant turn: the runner has already appended it when iterations exhaust, so the round-trip history carried a duplicate tool_use id and ended on an unanswered tool_use — a guaranteed 400 on continuation. The append now keys on stop_reason, and truncation reads natively as stop_reason: "tool_use" with a continuable history. - The envelope is JSON-serializable end to end: runner-stored turns carry pydantic content blocks; everything is dumped to plain dicts, excluding SDK-internal __api_exclude__ fields (parsed_output) that the API rejects on round-trip. - Bounded by default (max_iterations 10, like the OpenAI surfaces); usage aggregation now preserves the final turn's native fields and sums the token counters None-safely; empty caller system strings are skipped; non-dict message entries and bad doc_id types raise PageIndexAPIError; anthropic < 0.68 gets an actionable version error; the doc block no longer spends a cache_control breakpoint. chat_completions()/responses(): - MaxTurnsExceeded wraps into PageIndexAPIError on all four run paths. - responses(stream=True) is one logical response: per-turn backend lifecycle events are collapsed (a canonical consumer previously stopped at turn 1's response.completed and never saw the answer), sequence numbers are reassigned monotonically, and the synthesized tool-output event carries output_index/sequence_number. - The responses envelope carries the real request surface (instructions, the actual function tool definitions, tool_choice, parallel_tool_calls, error/incomplete_details). - RunConfig(group_id) pins a stable prompt_cache_key: openai-agents otherwise stamps each run with a fresh key, tagging round-tripped prefixes as different cache groups and defeating the feature the round-trip exists for. - Abandoning a stream now cancels the run: a watchdog task lets the cancellation land even while the pump awaits the backend, and the per-call AsyncOpenAI client is closed before its loop ends (fixes "Task exception was never retrieved" noise). The opening role chunk is emitted even for empty outputs; empty responses() input and enable_citations-before-extra ordering fixed. Docs rescoped to what is true: finish_reason/status reflect loop completion on the OpenAI surfaces (the engine does not surface per-turn backend reasons); chat streaming yields visible narration including pre-tool text; messages(stream=True) forwards the Anthropic SDK's native event objects (not wire-verbatim); the doc block is a leading conversation item on OpenAI surfaces and a system block on messages(). Tests: 25 in the file (11 new), with per-extra skip sections so a machine with only one framework still covers the other surface; without-frameworks matrix re-verified; live smoke re-run green with a clean exit. * feat: as_anthropic_tools — Anthropic tool-runner export, both modes Fills the last cell of the agent-connection matrix: users driving their own anthropic tool_runner loop get runnable tools directly. Cloud wraps the live MCP tool set with input schemas passing through verbatim (MCP inputSchema is the Messages API schema shape); local exposes the same set messages() runs internally. The beta_tool wrapping moves from local_chat into integrations/anthropic_sdk.py, parallel to openai_agents.py, and messages() now consumes the shared builder. agent_tools grows _bridge_invoker/_read_only_tools so the plain-function and beta_tool cloud paths share invocation containment and the read-only gate. * fix: as_anthropic_tools review findings — async flavor, schema isolation Adversarial + best-practice review of4590dd8(three independent passes) surfaced two holes. The export was sync-only: AsyncAnthropic's runner accepts only BetaAsyncFunctionTool and splices anything else into the request body unserialized, so the first call died with an opaque TypeError — asynchronous=True now builds beta_async_tool runnables (present since the 0.68.0 floor) that run the blocking bridge/store call in a worker thread, keeping I/O off the caller's event loop. And beta_tool stores input_schema by reference, so cloud tools aliased the bridge's cached metas while the local path deep-copied — the builder now copies, and the passthrough test asserts equal-but-not-aliased so it can no longer compare an object with itself. Docstring fixes from the same round: the MCP-connector pointer now carries the full live-verified shape (authorization_token was missing — following it literally gave a 401), and the manual messages.create loop's to_dict() serialization is documented. Tests pin the runnable flavor both ways (isinstance), which existing tests could not distinguish. * docs: doc_id is per-call table-setting — keep it identical across a conversation The targeting block doc_id adds is re-set on every call and sits in the cached prompt prefix, so a round-trip that drops (or changes) doc_id silently diverges the prefix and loses the cache continuation. State the rule on all three chat surfaces' doc_id docs, and pin it with a prefix test that passes the same doc_id on both calls. * feat: every chat surface takes a bare query string query + doc_id is the minimal PageIndex contract, so it now works uniformly: chat_completions and messages accept a plain string (one user message), as responses always did per its wire format. The wrap is input sugar at the SDK surface, not a translation layer — the outgoing wire is unchanged, and managed agent surfaces taking strings is the ecosystem convention (Runner.run, claude_agent_sdk.query). Cloud chat_completions gains the same acceptance; blank strings raise on every path. * feat: messages() defaults max_tokens to 4096 The Messages API requires a per-turn output budget on the wire, but that is table-setting, not a PageIndex-layer user obligation — the simple call is now a question + model + doc_id. The knob stays overridable (passthrough intact); model stays required because no cross-vendor default is honest to guess. * fix: raise messages() max_tokens default to 8192 max_tokens is a cap, not consumption, so the default should be the highest universally safe value: 4096 could truncate long-form answers (whole-document summaries), while 8192 is the output ceiling every non-EOL Claude model accepts and stays under the SDK's non-streaming long-request threshold. * fix: restore per-extra skip markers the string-input tests displaced Inserting tests above decorated ones absorbed their @needs_agents markers, so two tests ran (and failed) in the without-frameworks CI job. Both simulated-bare and full runs are green again. * fix: close 17 findings from the v0.2.10 max review Tool layer: - anthropic adapter: failed tool calls raise ToolError so the runner emits tool_result is_error:true; McpBridge.call_tool returns (text, is_error) and surfaces the server's MCP isError marking - as_openai_tools builds FunctionTool with the contract/server schema verbatim (strict off) — function_tool() regenerated schemas from signatures, dropping items/enum/pattern/bounds and aborting the whole list on object-typed params; shared _tool_specs() feeds both adapters - remove_document validates every name before deleting anything; call_tool classifies only bind-time TypeErrors as INVALID_INPUT - unknown-tool envelope formatted with _dumps like every other envelope Local chat: - doc_id is enforced at the tool layer (allowlist threaded through call_tool and the adapters), not just prompted; the shadow check runs inside the scope - _openai_model routes litellm/ and provider/ paths via LitellmModel and strips openai/ — the normalized retrieve_model 404'd as a raw wire name - responses() reports the backend's real terminal status (recorded at the transport client; the framework discards Response.status) and wraps framework exceptions in PageIndexAPIError - chat_completions streaming yields its opening chunk inside try, so an abandoned iterator still cancels the run and closes the backend - prompt-cache group_id is per-conversation (model+instructions+first item) instead of one global constant pooling every user - messages() max_tokens default resolves per model (claude-3 caps at 4096) Packaging / surface: - __init__ registers the 0.2.10 modules in _SUBMODULES; unknown names raise AttributeError instead of eagerly importing page_index_classic - anthropic floor 0.84.0: first release with ToolError whose runner also executes the final turn's tools on a max_iterations cut - client docstrings caught up with local chat landing Claude Agent SDK gate: - claude_allowed_tools(mcp_servers) derives mcp__<key>__<tool> entries from the caller's own registration map (live server annotations on cloud, the contract locally) — no name is ever spelled twice - claude_agent_config() bundles the three slots as one-call sugar over the explicit form Examples: - demo runs against cloud again (getattr for local-only attrs) and finds an existing indexed copy by name before re-indexing Tests: monkeypatches replace the consuming module's binding instead of mutating the shared time/requests modules; 185 -> 211. * feat: one-call config bundles for every bring-your-own-framework surface claude_agent_config() gets two symmetric siblings, so each framework's front door is a single splat over the same explicit primitives: - openai_agent_config(): Agent(**...) kwargs — instructions, tools, and the local retrieve_model (cloud omits model for the framework default) - anthropic_runner_config(): tool_runner(**...) kwargs — system, tools, and the messages() defaults (per-model max_tokens, 10-iteration bound); only the user's messages remain Bundles stay pure sugar: doc_id rides agent_instructions, no extra semantics over the explicit form, docstrings point both ways. The demo agent shrinks to Agent(**client.openai_agent_config(doc_id=...)). Construction is pinned against the real frameworks in tests (Agent and tool_runner both built offline), so an upstream kwargs rename fails loudly; 211 -> 215 tests. * fix: three more review findings — partial-read reporting, reply correlation, output_index axis - get_page_content: the summary is additive, not either/or — a call that both truncates for size and has out-of-range pages reported only the latter, telling the agent every in-range page was returned (#2) - McpBridge._extract_result: strict request-id correlation only; the eager fallback could hand back a stale or mis-correlated JSON-RPC message as this call's reply (#16) - responses() streaming: output_index now addresses the logical response.output — backend per-turn indexes are re-based past prior turns' items and the SDK-injected tool outputs take the next slot on that axis, instead of reusing the event-sequence counter (#15) 215 -> 217 tests. * feat: gate the config-handoff surfaces by the read-only MCP endpoint pageindex-chat#448 adds /mcp?tools=read — the server registers only readOnlyHint-annotated tools — so the URL itself becomes the gate for every surface that hands a config to a third party: - as_claude_mcp: include_management now picks the endpoint on cloud; the parameter is real in both modes - as_openai_tools(hosted=True): OpenAI connects to the read-only endpoint by default and require_approval simplifies to "never" — the approval-flow middle ground becomes hard absence, matching every other surface's default - claude_allowed_tools() retired before ever shipping: with the server gated, allowed_tools degenerates to whole-server pre-approval, which claude_agent_config emits as the constant ["mcp__<name>"] — no setup-time bridge round-trip remains - in-process surfaces (agent_tools, as_openai_tools, as_anthropic_tools over the bridge) keep bare /mcp + client-side annotation filtering: they materialize tools locally and hand no URL to anyone Release ordering: 0.2.10 must ship after pageindex-chat#448 deploys — an older server ignores unknown query params and would silently serve the full set behind a URL that promises read-only. * fix: chat_completions wraps framework exceptions like responses() The AgentsException -> PageIndexAPIError wrap from the responses() fix covered only that surface; a backend stream dying without a terminal event (or any engine failure) still escaped chat_completions as a raw openai-agents exception type on both its paths. * fix: same-name documents in different folders no longer refuse doc_id targeting The shadow check in doc_targeting_block compared names across the whole library, so agent_instructions(doc_id=...) hard-raised for a legal cloud layout — one file name in two folders — with advice (rename/remove) that contradicts the contract, whose folder_id parameter exists precisely to disambiguate this case. Shadowing is now judged per folder: only a newer same-name document in the SAME folder makes the name unreachable and raises. A same-name document in another folder serves the call, and the targeting block adds a directive to pass folder_id on every tool call — dropping the raise alone would have traded a loud refusal for the agent silently reading the newer document. Local mode (folderId always None) and the scoped chat path (allowlist resolution, fixed with the doc_id enforcement) are behaviorally unchanged. * Revert "fix: same-name documents in different folders no longer refuse doc_id targeting" This reverts commit 9d16dbf6, whose premise collapsed on verification against the cloud upload paths. Review finding #10 inferred from the folder_id tool description that one file name in two folders is a legal cloud layout; both upload paths actually dedup names per USER SPACE with no folder dimension — chat's getSignedUploadUrl queries fileName + sourceName + mode + owner (file-access.service.ts), and compute's get_upload_url probes the S3 key (user, source, file_name) — so own documents cannot share a name across any folders. The only legitimate same-name source is shared mounts (shared-with-me/following), which the api-proxy surface this SDK talks to never carries. Cloud and local therefore share one invariant — names unique per space, server-enforced — and the original global shadow check was the right shape: a duplicate is an anomaly worth refusing loudly, not a layout to accommodate with per-folder adjudication and conditional prompt notes. The invariant is now stated in doc_targeting_block's docstring so the finding does not get re-raised. * docs: state the name-uniqueness invariant in library terms * docs: trim doc_targeting_block docstring to the contract * fix: compress out-of-range page lists in get_page_content Two message strings enumerated every out-of-range page number one by one while the payload fields beside them already used _format_page_spec. Against a 2-page document, pages="3-10000" produced a 59,310-character error whose own requested_pages field expressed the identical set as "3-10000"; the mixed case pages="1-10000" produced 59,479. Both now render through the helper: 431 and 600 characters. This was inherited behaviour, not a local slip — the cloud MCP server enumerated at the same two sites, so local reproduced it verbatim. The cloud fixed it first (pageindex-chat #449), and this follows to keep the strings byte-identical; the error message now matches the served one character for character. A differential run of the two compressors over 94 inputs (empty, single, unsorted, duplicated, 10k spans, 80 random) agrees on every one, separator included. The new test pins all three shapes, including the non-contiguous case ("5,9" must not collapse into a range) that the compressor had no direct coverage for. * docs: messages() marks only the managed prefix with cache_control The method docstring claimed the doc targeting block carries a cache_control breakpoint too, and the doc_id note called that block part of the cached prompt prefix. _anthropic_system deliberately marks only the stable managed prefix — the API allows four breakpoints and the varying doc block must not consume one — and the block is appended after the sole breakpoint, so it is never cached.a45b554added the block with cache_control, making the claim true when written;daac9d2removed it without touching the docstring, andadb2f1fthen added the "cached prompt prefix" sentence after the fact. The same phrase at the chat_completions and responses docstrings is correct — there the block is a leading conversation item inside the auto-cached prefix — so only the Messages surface is reworded. The doc_id advice itself stands: the block is per-call table-setting and should stay identical across a conversation. Only the caching rationale was wrong. * docs: _stream_sync cancels on close, not on abandonment The docstring promised that closing "or abandoning" the iterator cancels the run. Abandoning only works when refcounting collects the generator: a caller that breaks out of the loop while keeping the reference never runs the finally that sets the cancel event, so the pump thread stays parked on the full queue and the backend client is never released. Closing is correct and is what the dedicated test exercises. Narrowing the promise to the behaviour the code actually provides is the honest fix; a watchdog or finalizer would be machinery bought for a shape the sync surface is not meant to serve, and the async client planned for 0.2.11 gets native task cancellation instead. * fix: raise the openai-agents floor to 0.14.0 _conversation_group_id feeds RunConfig.group_id into OpenAI's prompt_cache_key so a round-tripped prefix stays in one cache group. That wiring first appears in openai-agents 0.14.0: 0.8.0 through 0.13.x have no prompt_cache_key at all, and group_id there is a tracing group id only — inert, since tracing is disabled on the line above. An install resolving to the declared floor lost the cache continuity that the responses() docstring sells, silently and with no test able to catch it. The old floor's rationale (0.8.0 offloads sync tools to a thread) is subsumed by the new one. Every symbol the package imports predates 0.14.0, so nothing else constrains the bound. * fix: enforce doc_id at the tool layer in the framework config helpers openai_agent_config / anthropic_runner_config / claude_agent_config accepted doc_id but built unscoped tools, so the parameter that is a structural allowlist on chat_completions() was prompt-only advice here — the agent could read every document in the store regardless. - as_openai_tools / as_anthropic_tools / as_claude_mcp take a doc_id tail parameter and thread it to the existing _allowed_ids channel; the config helpers pass it through in local mode - cloud config helpers keep prompt-level targeting (tool scoping is server-side there, documented); explicit as_*(doc_id=...) raises on cloud instead of silently dropping the allowlist — including the hosted branch, which returned before _tool_specs' existing guard - _require_local_scope consolidates the cloud rejection that was inlined in _tool_specs - doc_id=[] is an empty allowlist, not "unscoped": dropped the `or None` at the three local chat surfaces * fix: two chat findings — final-turn append and cache-key seeding run_messages keyed its re-append guard on stop_reason, but the anthropic runner executes tools whenever the turn's content carries tool_use blocks (refusal excepted) — a max_tokens turn with complete tool_use blocks was already appended by the runner, so the guard re-appended it, duplicating tool_use ids and 400ing the documented verbatim continuation. The guard now checks whether final's tool_use ids already sit in the appended history; unexecuted tool_use blocks (refusal turns) are stripped from the appendable history, as the SDK itself does when rebuilding params around an unresulted turn. _conversation_group_id seeded on items[0], which is the doc-targeting block whenever doc_id is set — byte-identical across every conversation about a document, so all of them pooled under one prompt_cache_key and evicted each other's prefixes. Seed on the conversation's own first item instead: continuations keep their key, unrelated conversations never share one. Also drop the dead pytestmark_openai assignment (pytest's magic name is pytestmark; the section gate it implied never existed). * fix: six review findings — pagination, compat, and containment - _all_documents advances by what actually arrived and treats `total` as an optimization: absent/null totals and short pages silently truncated the library behind every name resolution - _make_bridge_function survives description: null (the parallel _tool_specs path already did) - as_openai_tools answers a malformed argument string with the guided error envelope instead of raising through the caller's whole run - the pre-0.2.10 package attributes (ConfigLoader, count_tokens, ...) resolve again: main's underscore-guarded fallthrough is restored — dunder probes stay lazy, a non-underscore typo pays one classic import before its AttributeError - _split_structure chunks are always lists: the structure field no longer changes JSON type between parts of one paginated response - the bridge replays only session-carrying 404s (the spec's expiry status); 400 raises instead of re-running side effects, and the reset double-checks under the lock so concurrent retries cannot clobber a freshly re-initialized session * fix: five secondary review findings — containment and guards - the bridge maps content blocks individually: base64 payloads (image/audio) become metadata stubs instead of handing the model the raw blob, text blocks pass verbatim, anything else keeps the JSON dump (revisit if tool results become real multimodal input) - cloud proxy annotations keep array item types (list[str], not bare list) so strict function calling accepts the round-trip; a type-array in items degrades to bare list instead of crashing the build - run_messages raises when set_messages_params stops delivering params instead of silently dropping every tool turn from the envelope - call_tool drops None-valued arguments (None ≡ omitted, the contract's semantics) — adapters that forward the model's nulls verbatim no longer trip parameter validation - client._parse_pages bounds the span arithmetically before materializing it, like the tool layer: "1-999999999" raises instead of allocating a billion integers * fix: three review findings — protocol honesty, model echo, containment - responses() promised the Responses protocol ("no translation layer") but _openai_model ignored protocol on the LiteLLM branch: provider- prefixed models silently ran chat.completions under a responses-shaped envelope, and with no transport hook to record status (LitellmModel has no _client.responses) a turn truncated at the output cap reported status "completed". The branch now raises for protocol == "responses" — at agent-build time, before any backend call — naming the routes out: chat_completions(), messages() for Anthropic models, or OPENAI_BASE_URL + a bare/openai/-prefixed name for backends that genuinely speak /responses. Refusal, not emulation: most providers have no /responses endpoint to drive. - chat_completions envelopes echoed retrieve_model verbatim, which carries the SDK's litellm/ routing marker after normalization — a name no provider catalog contains, and a different string than the same model passed per-call. The envelope and every streaming chunk now report the name the provider actually serves; routing and the prompt-cache group key keep the prefixed form. responses() needs no change (post-refusal the prefix cannot reach its envelope), and the user-typed openai/ prefix stays echoed as typed. - _remove_document caught only PageIndexAPIError around the per-doc delete, so a bare OSError (local_store re-raises them) or a transport error (cloud delete_document wraps nothing) escaped mid-batch, discarded the entries for documents already irreversibly deleted, and surfaced as a generic INTERNAL_ERROR envelope inviting a retry — which then reports the destroyed document as not_found. The loop now catches Exception, keeping the per-document results the contract promises. * fix: config bundles use the scoped shadow check their tools earned9f67fddmade the three config helpers enforce doc_id at the tool layer but left their instructions on doc_targeting_block's unscoped default, so a bundle refused any doc_id whose name a newer library-wide duplicate shadows — a raise whose message ("the tools address documents by name and would read the newer one") had just become false: the bundle's own tools resolve names inside the allowlist and read the targeted document correctly. chat_completions() accepted the same doc_id via _doc_block's scoped=True. Each helper now computes scope = _local_doc_scope(doc_id) once and derives both slots from it — scoped=scope is not None for the instructions, doc_id=scope for the tools — so the check mode and the tool allowlist come from one fact and cannot drift apart again. build_agent_instructions grows a scoped passthrough; cloud stays on the whole-library check (scope is None there and the tools are genuinely unscoped), and the public agent_instructions() keeps its unscoped default for the same reason. An in-set duplicate still raises — and in that case the message is true on every surface that emits it. * docs: as_openai_tools' remote-MCP note moves to the Cloud paragraph The MCPServerStreamableHttp alternative sat in the Local: paragraph pointing at bare {BASE_URL}/mcp — a cloud-only route (BASE_URL is the hosted API; local has no HTTP MCP server) that as written would connect unauthenticated to the full tool set. Now stated where it applies, in the as_anthropic_tools connector-note form: Cloud paragraph, Bearer auth spelled out, ?tools=read default with the drop-it escape. * fix: nine review findings — argument coercion, scope, and honest envelopes - call_tool coerces string booleans per the TOOL_CONTRACT schema ("false"/"no"/"0" read as False, not a truthy 3-minute wait) and survives arguments: null (json.loads("null") reaches the seam as None) - _local_doc_scope raises on an explicitly empty doc_id on cloud: with no tool-layer allowlist there, dropping it silently widened an empty scope to the whole library - both page-spec caps count distinct pages instead of summing parts, so overlapping ranges (a parent section plus its children) within the 10k union pass again as they did in 0.2.9; the per-part arithmetic bound still rejects billion-page specs before materializing anything - _remove_document deduplicates doc_names: a repeated name is one deletion, not a second "failed" row with an internal error string - doc_targeting_block merges the user's metadata tags from the listing (local get_document keeps the 7-key cloud detail wire shape, which carries none) so the block delivers the metadata it promises - _wait_until_ready folds its two raise branches into one that carries the doc_id: a poll that dies no longer discards the handle to an uploaded, billed document - _reported_model strips both routing prefixes (litellm/ and openai/) and responses() now reports it too, instead of echoing a model id the provider never served - _openai_model wraps AsyncOpenAI() construction so a missing backend credential surfaces as PageIndexAPIError like every other gate on the chat surfaces (and builds the client once for both protocols) - _browse_documents advances its cursor by the rows that actually arrived and guards a null/absent total — the same hazards _all_documents already guards — and an empty window ends pagination instead of freezing the cursor * fix: two chat findings — protocol terminal states, provider error types - responses(stream=True) raised PageIndexAPIError when the backend ended the response with response.failed / response.incomplete: openai-agents yields the terminal lifecycle event, then re-raises it as ModelBehaviorError, so the generic AgentsException wrap short-circuited the emit the agen's tail was built for — its failed/incomplete terminal mapping was dead code against the real engine, and the caller lost both the partial output and the real status. The wrap now steps aside when the recorded terminal state is failed/incomplete, and the stream ends with the honest terminal event (committed output, real status, error/incomplete_details) — the backend's terminal state is a protocol event, not an engine failure. Non-stream was already honest for incomplete via the transport recorder; a failed response arrives there as an HTTP error, covered below. Known ceiling: the truncated final turn's partial text was already streamed as deltas but is not reconstructed into the terminal event's output (the engine commits items only on turn completion). - Provider exceptions (network, auth, rate limit) leaked as raw openai/anthropic types through every chat surface, against the layer's own "never raw engine types" contract. Every engine boundary now wraps its vendor's base exception into PageIndexAPIError (chained): the four OpenAI-engine sites catch openai.OpenAIError — LiteLLM's exception types subclass openai's, so one handler covers both routing paths — and messages() catches anthropic.AnthropicError around the batch drive and the stream generator. * fix: guided failure for unknown LiteLLM providers, non-object call_tool args - _openai_model pre-checks the first path segment against litellm.provider_list (fail-open if the attribute ever disappears): a HuggingFace repo id like Qwen/Qwen2.5-7B-Instruct on an OpenAI-compatible server now fails at build time with the escape spelled out — 'openai/<id>' plus OPENAI_BASE_URL — instead of at request time inside LiteLLM with "LLM Provider NOT provided". The slash-means-provider routing convention itself is unchanged; the retrieve_model and chat_completions docstrings now document it where they promise "any OpenAI-compatible server works" - call_tool answers a non-dict arguments value (a JSON array or scalar from a misbehaving caller) with the guided INVALID_INPUT envelope instead of raising AttributeError through the agent loop, matching the openai adapter's own non-object guard * fix: wrap litellm import in PageIndexAPIError when not installed * fix: silence CodeQL findings — merge implicit string concat, drop unused vars * fix: two external review findings — init-notification race, SDK floor notifications/initialized moves inside the bridge lock: a concurrent first use could send tools/list between the handshake and the notification, which strict MCP servers reject with a 400 the bridge never replays. Regression test races two threads through a stalled notification window. claude-agent-sdk floor rises to 0.1.53 — below it, string prompts with SDK MCP servers (the documented local-mode flow) hit invisible registration (#597) and a deadlock (#780). * refactor: drop the unused exc parameter from _wrap_max_turns The parameter was dead from the moment it was introduced (daac9d2): the body reads only max_turns, and every call site already carries the cause via `raise ... from exc`. The signature implied the helper inspected the engine exception, which it never did. No behavior change — message text and __cause__ chaining verified identical across all four call sites (chat_completions and responses, stream and non-stream). * fix: raise the anthropic and openai-agents floors past broken releases Both declared floors named a version that cannot work, and CI never caught either because it installs the latest. anthropic >=0.84.0 -> >=0.108.0. Probed against a mock transport: on a turn with stop_reason="refusal" carrying a tool_use block, 0.84.0, 0.92.0 and 0.100.0 all execute the tool and post the tool_result back; 0.108.0 and later stop at the refusal. test_messages_refusal_with_ tool_use_stays_appendable asserts the latter, so that test was false at the floor. messages() is unaffected in practice (it never passes include_management, so remove_document is not registered), but as_anthropic_tools(include_management=True) hands it to a caller's own runner. openai-agents >=0.14.0 -> >=0.18.1. 0.14.0 and 0.16.0 raise pydantic ValidationError on InputTokensDetails.cache_write_tokens before any request reaches the transport when paired with openai 2.54.0 — and they declare openai <3,>=2.26.0, so pip resolves exactly that pair. 0.18.1 is clean. The 0.14.0 rationale (RunConfig.group_id -> prompt_cache_key) still holds above the new floor. The three extras' floor comments are cut to the binding constraint; the reasoning lives here. * test: cover max_turns wrapping on every chat surface test_chat_completions_max_turns_wrapped only drove chat_completions, so the two responses() call sites had no coverage, and no test asserted that the engine exception survives as __cause__. Parametrized over both surfaces and both stream modes; the non-positive max_turns rejection splits out, since it is input validation rather than wrapping. * fix: four review findings — envelope size honesty, contained tool errors - _dumps drops indent=2: emission now matches _serialized_size's compact accounting, so the pagination budget bounds what is actually sent (indented parts measured under 95k but emitted ~1.8x the 100k cap) - call_tool builds the _allowed_ids frozenset inside the guarded block: a non-iterable doc_id returns the INVALID_INPUT envelope instead of raising into the agent loop; same move for _bridge_invoker's arguments normalization - next_steps strings qualify submit_document() as PageIndexClient.submit_document() (three sites), matching the one already-qualified site — it is a client method, not a registered tool - tests: import httpx at module scope (guaranteed via the hard openai dependency) so agents-gated tests survive an install without the anthropic extra; formatting assertion follows the compact envelope
This commit is contained in:
@@ -31,5 +31,5 @@ jobs:
|
||||
cache: pip
|
||||
- run: pip install -r requirements.txt pytest
|
||||
- if: matrix.agent-frameworks == 'with'
|
||||
run: pip install openai-agents claude-agent-sdk
|
||||
run: pip install openai-agents claude-agent-sdk anthropic
|
||||
- run: python -m pytest -q
|
||||
|
||||
@@ -6,3 +6,4 @@ __pycache__
|
||||
logs/
|
||||
.pageindex/
|
||||
dist/
|
||||
*.doc_id
|
||||
|
||||
@@ -6,20 +6,21 @@ local mode and the OpenAI Agents SDK. Instead of vector similarity search and
|
||||
chunking, PageIndex builds a hierarchical tree index and uses agentic LLM
|
||||
reasoning for human-like, context-aware retrieval.
|
||||
|
||||
Agent tools:
|
||||
- get_document() — document metadata (status, page count, etc.)
|
||||
- get_document_structure() — tree structure index of a document
|
||||
- get_page_content() — retrieve text content of specific pages
|
||||
The agent tools come straight from the SDK — ``client.as_openai_tools()``
|
||||
exposes the PageIndex tool contract (browse_documents, get_document,
|
||||
get_document_structure, get_page_content) and ``client.agent_instructions()``
|
||||
provides the retrieval playbook, so the whole agent is a few lines. Swap
|
||||
``PageIndexLocalClient()`` for ``PageIndexCloudClient(api_key=...)`` and the
|
||||
same code runs against the cloud.
|
||||
|
||||
Steps:
|
||||
1 — Index a PDF locally and view its tree structure index
|
||||
2 — View document metadata
|
||||
3 — Ask a question (agent reasons over the index and auto-calls tools)
|
||||
|
||||
Requirements: pip install openai-agents; OPENAI_API_KEY in the environment.
|
||||
Requirements: pip install "pageindex[openai]"; OPENAI_API_KEY in the environment.
|
||||
"""
|
||||
import sys
|
||||
import json
|
||||
import asyncio
|
||||
import concurrent.futures
|
||||
from pathlib import Path
|
||||
@@ -27,62 +28,30 @@ import requests
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).parent.parent))
|
||||
|
||||
from agents import Agent, Runner, function_tool, set_tracing_disabled
|
||||
from agents.model_settings import ModelSettings
|
||||
from agents import Agent, Runner, set_tracing_disabled
|
||||
from agents.stream_events import RawResponsesStreamEvent, RunItemStreamEvent
|
||||
from openai.types.responses import ResponseTextDeltaEvent, ResponseReasoningSummaryTextDeltaEvent
|
||||
|
||||
from pageindex import PageIndexClient
|
||||
from pageindex import PageIndexAPIError, PageIndexLocalClient
|
||||
import pageindex.utils as utils
|
||||
|
||||
PDF_URL = "https://arxiv.org/pdf/2603.15031"
|
||||
|
||||
_EXAMPLES_DIR = Path(__file__).parent
|
||||
PDF_PATH = _EXAMPLES_DIR / "documents" / "attention-residuals.pdf"
|
||||
DOC_ID_PATH = _EXAMPLES_DIR / "documents" / "attention-residuals.doc_id"
|
||||
STORAGE_PATH = _EXAMPLES_DIR / ".pageindex"
|
||||
|
||||
AGENT_SYSTEM_PROMPT = """
|
||||
You are PageIndex, a document QA assistant.
|
||||
TOOL USE:
|
||||
- Call get_document() first to confirm status and page count.
|
||||
- Call get_document_structure() to identify relevant page ranges.
|
||||
- Call get_page_content(pages="5-7") with tight ranges; never fetch the whole document.
|
||||
- Before each tool call, output one short sentence explaining the reason.
|
||||
Answer based only on tool output. Be concise.
|
||||
"""
|
||||
|
||||
|
||||
def query_agent(client: PageIndexClient, doc_id: str, prompt: str, verbose: bool = False) -> str:
|
||||
def query_agent(client: PageIndexLocalClient, doc_id: str, prompt: str, verbose: bool = False) -> str:
|
||||
"""Run a document QA agent using the OpenAI Agents SDK.
|
||||
|
||||
Streams text output token-by-token and returns the full answer string.
|
||||
Tool calls are always printed; verbose=True also prints arguments and output previews.
|
||||
"""
|
||||
|
||||
@function_tool
|
||||
def get_document() -> str:
|
||||
"""Get document metadata: status, page count, name, and description."""
|
||||
return json.dumps(client.get_document(doc_id))
|
||||
|
||||
@function_tool
|
||||
def get_document_structure() -> str:
|
||||
"""Get the document's full tree structure (without text) to find relevant sections."""
|
||||
return json.dumps(client.get_document_structure(doc_id), ensure_ascii=False)
|
||||
|
||||
@function_tool
|
||||
def get_page_content(pages: str) -> str:
|
||||
"""
|
||||
Get the text content of specific pages.
|
||||
Use tight ranges: e.g. '5-7' for pages 5 to 7, '3,8' for pages 3 and 8, '12' for page 12.
|
||||
"""
|
||||
return json.dumps(client.get_page_content(doc_id, pages), ensure_ascii=False)
|
||||
|
||||
agent = Agent(
|
||||
name="PageIndex",
|
||||
instructions=AGENT_SYSTEM_PROMPT,
|
||||
tools=[get_document, get_document_structure, get_page_content],
|
||||
model=getattr(client, "retrieve_model", None),
|
||||
# model_settings=ModelSettings(reasoning={"effort": "low", "summary": "auto"}), # Uncomment to enable reasoning
|
||||
**client.openai_agent_config(doc_id=doc_id),
|
||||
# model_settings=ModelSettings(reasoning={"effort": "low", "summary": "auto"}), # from agents.model_settings import ModelSettings
|
||||
)
|
||||
|
||||
async def _run():
|
||||
@@ -152,21 +121,33 @@ if __name__ == "__main__":
|
||||
print("Download complete.\n")
|
||||
|
||||
# Setup: local mode — no PageIndex API key needed, your LLM key does the work
|
||||
client = PageIndexClient(storage_path=str(STORAGE_PATH))
|
||||
client = PageIndexLocalClient(storage_path=str(STORAGE_PATH))
|
||||
|
||||
# Step 1: Index PDF and view tree structure
|
||||
print("=" * 60)
|
||||
print("Step 1: Index PDF and view tree structure")
|
||||
print("=" * 60)
|
||||
doc_id = next(
|
||||
(doc["id"] for doc in client.list_documents(limit=100)["documents"]
|
||||
if doc["name"] == PDF_PATH.name),
|
||||
None,
|
||||
)
|
||||
doc_id = None
|
||||
if DOC_ID_PATH.exists():
|
||||
cached = DOC_ID_PATH.read_text().strip()
|
||||
try:
|
||||
client.get_document(cached)
|
||||
doc_id = cached
|
||||
except PageIndexAPIError:
|
||||
DOC_ID_PATH.unlink()
|
||||
if doc_id is None:
|
||||
# The .doc_id cache is gitignored — on a fresh clone with an
|
||||
# existing store, find the already-indexed copy by name instead of
|
||||
# re-indexing it.
|
||||
doc_id = next(
|
||||
(doc["id"] for doc in client.list_documents(limit=100)["documents"]
|
||||
if doc["name"] == PDF_PATH.name), None)
|
||||
if doc_id:
|
||||
DOC_ID_PATH.write_text(doc_id)
|
||||
print(f"\nLoaded cached doc_id: {doc_id}")
|
||||
else:
|
||||
doc_id = client.submit_document(str(PDF_PATH))["doc_id"]
|
||||
doc_id = client.submit_document(str(PDF_PATH), wait=True)["doc_id"]
|
||||
DOC_ID_PATH.write_text(doc_id)
|
||||
print(f"\nIndexed. doc_id: {doc_id}")
|
||||
print("\nTree Structure (top-level sections):")
|
||||
structure = client.get_tree(doc_id, node_summary=True)["result"]
|
||||
|
||||
+16
-6
@@ -18,26 +18,36 @@ __all__ = [
|
||||
]
|
||||
|
||||
_LAZY = {
|
||||
"page_index": ".page_index_classic",
|
||||
"page_index_main": ".page_index_classic",
|
||||
"page_index_flash": ".flash",
|
||||
"optimize_tree": ".tree_optimize",
|
||||
"md_to_tree": ".page_index_md",
|
||||
}
|
||||
_SUBMODULES = {"client", "cloud_api", "errors", "flash", "local_api",
|
||||
"local_store", "page_index_classic", "page_index_md", "tree_optimize",
|
||||
"utils"}
|
||||
|
||||
_SUBMODULES = {"agent_tools", "client", "cloud_api", "errors", "flash",
|
||||
"integrations", "local_api", "local_chat", "local_store",
|
||||
"mcp_bridge", "page_index_classic", "page_index_md",
|
||||
"tree_optimize", "utils"}
|
||||
|
||||
def __getattr__(name):
|
||||
if name.startswith("_"):
|
||||
# Dunder probes (copy, pickle, inspect) are the frequent unknown
|
||||
# names — they must not trigger the classic import below.
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
import importlib
|
||||
if name in _SUBMODULES:
|
||||
return importlib.import_module(f".{name}", __name__)
|
||||
module = importlib.import_module(_LAZY.get(name, ".page_index_classic"), __name__)
|
||||
# Pre-0.2.10 compat: unknown names fall through to the classic module,
|
||||
# whose public surface (ConfigLoader, count_tokens, ...) resolved as
|
||||
# package attributes. A non-underscore typo pays one classic import
|
||||
# before its AttributeError — not worth an allowlist.
|
||||
module = importlib.import_module(_LAZY.get(name, ".page_index_classic"),
|
||||
__name__)
|
||||
try:
|
||||
value = getattr(module, name)
|
||||
except AttributeError:
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
|
||||
raise AttributeError(
|
||||
f"module {__name__!r} has no attribute {name!r}") from None
|
||||
globals()[name] = value
|
||||
return value
|
||||
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
"""Installed-package version, shared by every surface that reports it upstream."""
|
||||
from __future__ import annotations
|
||||
|
||||
|
||||
def sdk_version() -> str:
|
||||
try:
|
||||
from importlib.metadata import version
|
||||
return version("pageindex")
|
||||
except Exception:
|
||||
return "0.0.0"
|
||||
File diff suppressed because it is too large
Load Diff
+588
-35
@@ -1,23 +1,36 @@
|
||||
"""PageIndex SDK client: the 0.2.x cloud surface, now with a local mode."""
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, Iterator, Optional, Union
|
||||
import os
|
||||
import time
|
||||
import warnings
|
||||
from typing import Any, Callable, Iterator, Optional, Union
|
||||
|
||||
from .errors import PageIndexAPIError
|
||||
|
||||
|
||||
def _parse_pages(pages: str) -> list[int]:
|
||||
result = []
|
||||
result: set[int] = set()
|
||||
too_many = (f"Page specification '{pages}' spans more than "
|
||||
"10000 pages; request a narrower range")
|
||||
for part in pages.split(","):
|
||||
part = part.strip()
|
||||
if "-" in part:
|
||||
start, end = (int(x) for x in part.split("-", 1))
|
||||
if start > end:
|
||||
raise ValueError(f"Invalid range '{part}': start must be <= end")
|
||||
result.extend(range(start, end + 1))
|
||||
else:
|
||||
result.append(int(part))
|
||||
return sorted(set(result))
|
||||
start = end = int(part)
|
||||
# Bound each part arithmetically before materializing it — a spec
|
||||
# like "1-999999999" would otherwise expand to a billion integers.
|
||||
# The cap is on distinct pages, so overlapping parts (a parent
|
||||
# section plus its children) don't double-count.
|
||||
if end - start + 1 > 10_000:
|
||||
raise ValueError(too_many)
|
||||
result.update(range(start, end + 1))
|
||||
if len(result) > 10_000:
|
||||
raise ValueError(too_many)
|
||||
return sorted(result)
|
||||
|
||||
|
||||
def _normalize_retrieve_model(model: str) -> str:
|
||||
@@ -47,10 +60,12 @@ class PageIndexClient:
|
||||
trees. Defaults to the packaged config (see pageindex/config.yaml).
|
||||
summary_model (str, optional): Local mode only — LLM used for node
|
||||
summaries and document descriptions.
|
||||
retrieve_model (str, optional): Local mode only — exposed as
|
||||
``client.retrieve_model`` (the agent demo reads it); the SDK
|
||||
itself consumes it once agent-based local chat lands in a
|
||||
later release.
|
||||
retrieve_model (str, optional): Local mode only — the model the
|
||||
local chat surfaces (``chat_completions``, ``responses``)
|
||||
default to, exposed as ``client.retrieve_model``.
|
||||
``provider/model`` names route through LiteLLM; for an
|
||||
OpenAI-compatible server that itself serves slashed model ids
|
||||
(vLLM, TGI), prefix ``openai/`` (e.g. ``openai/Qwen/...``).
|
||||
storage_path (str, optional): Local mode only — directory where
|
||||
indexed documents are stored. Defaults to ``./.pageindex``.
|
||||
|
||||
@@ -62,10 +77,9 @@ class PageIndexClient:
|
||||
instead of inferring it from api_key.
|
||||
|
||||
Local mode differences (all documented per method): indexing is
|
||||
synchronous, only PDFs are supported, and ``chat_completions`` (until
|
||||
agent-based local chat lands in a later release) / folders /
|
||||
``beta_headers`` / the deprecated retrieval API (``submit_query``,
|
||||
``get_retrieval``) are cloud-only.
|
||||
synchronous, only PDFs are supported, and folders / ``beta_headers`` /
|
||||
the deprecated retrieval API (``submit_query``, ``get_retrieval``) are
|
||||
cloud-only.
|
||||
"""
|
||||
|
||||
BASE_URL = "https://api.pageindex.ai"
|
||||
@@ -126,12 +140,14 @@ class PageIndexClient:
|
||||
beta_headers: Optional[list[str]] = None,
|
||||
folder_id: Optional[str] = None,
|
||||
metadata: Optional[dict] = None,
|
||||
wait: bool = False,
|
||||
) -> dict[str, Any]:
|
||||
"""
|
||||
Submit a PDF document for processing. Returns {'doc_id': ...}.
|
||||
Submit a PDF document for processing. Returns {'doc_id': ..., 'name': ...}.
|
||||
|
||||
Cloud: uploads the file; processing is asynchronous — poll
|
||||
``is_retrieval_ready(doc_id)`` before retrieving.
|
||||
Cloud: uploads the file; processing is asynchronous. Pass
|
||||
``wait=True`` to block until the document is ready, or poll
|
||||
``get_document(doc_id)['status']`` yourself.
|
||||
|
||||
Local: indexes the document in this call (it blocks while your LLM
|
||||
builds the tree — minutes for a standard index of a long document),
|
||||
@@ -151,14 +167,67 @@ class PageIndexClient:
|
||||
metadata (dict, optional): Your own JSON-serializable tags for the
|
||||
document; returned in get_tree/get_ocr responses and
|
||||
list_documents entries (both modes).
|
||||
wait (bool): Return only once the document is ready for use.
|
||||
Cloud: polls status until "completed" (raises on "failed" or
|
||||
after 30 minutes). Local: indexing is synchronous already, so
|
||||
this changes nothing. Leave False to submit many documents
|
||||
concurrently and poll afterwards.
|
||||
|
||||
Returns:
|
||||
dict: {'doc_id': ...}
|
||||
dict: {'doc_id': ..., 'name': ...}. 'name' is the stored document
|
||||
name: a taken name gains a numeric suffix (name_1..name_99)
|
||||
and a UserWarning is emitted. Older cloud servers omit 'name'.
|
||||
"""
|
||||
return self._api.submit_document(
|
||||
result = self._api.submit_document(
|
||||
file_path=file_path, mode=mode,
|
||||
beta_headers=beta_headers, folder_id=folder_id, metadata=metadata,
|
||||
)
|
||||
stored = result.get("name")
|
||||
if stored and stored != os.path.basename(file_path):
|
||||
warnings.warn(
|
||||
f'Document "{os.path.basename(file_path)}" was stored as '
|
||||
f'"{stored}".',
|
||||
stacklevel=2,
|
||||
)
|
||||
if wait:
|
||||
self._wait_until_ready(result["doc_id"])
|
||||
return result
|
||||
|
||||
def _wait_until_ready(self, doc_id: str, timeout: float = 1800.0) -> None:
|
||||
import requests
|
||||
interval = 2.0
|
||||
deadline = time.monotonic() + timeout
|
||||
poll_failures = 0
|
||||
while True:
|
||||
try:
|
||||
status = self.get_document(doc_id).get("status")
|
||||
poll_failures = 0
|
||||
except (PageIndexAPIError, requests.RequestException) as exc:
|
||||
# Tolerate transient poll failures; a 30-minute wait should
|
||||
# not die on one 502 or dropped connection.
|
||||
poll_failures += 1
|
||||
if poll_failures >= 3:
|
||||
raise PageIndexAPIError(
|
||||
f"Could not poll document status (doc_id: {doc_id}): "
|
||||
f"{exc}. Processing continues in the cloud — poll "
|
||||
"get_document(doc_id) for status."
|
||||
) from exc
|
||||
status = None
|
||||
if status == "completed":
|
||||
return
|
||||
if status == "failed":
|
||||
raise PageIndexAPIError(
|
||||
f"Document processing failed (doc_id: {doc_id})."
|
||||
)
|
||||
if time.monotonic() >= deadline:
|
||||
raise PageIndexAPIError(
|
||||
f"Timed out after {int(timeout)}s waiting for document "
|
||||
f"processing (doc_id: {doc_id}, last status: {status}). "
|
||||
"Processing continues in the cloud — poll "
|
||||
"get_document(doc_id) for status."
|
||||
)
|
||||
time.sleep(interval)
|
||||
interval = min(interval * 1.5, 15.0)
|
||||
|
||||
# ---------- OCR FUNCTIONALITY ----------
|
||||
|
||||
@@ -256,11 +325,11 @@ class PageIndexClient:
|
||||
|
||||
Cloud-only: the cloud API marks this endpoint deprecated in favor of
|
||||
chat completions, so local mode does not implement it — raises
|
||||
PageIndexAPIError. Use ``chat_completions`` (cloud) instead.
|
||||
PageIndexAPIError. Use ``chat_completions`` instead.
|
||||
"""
|
||||
return self._require_cloud(
|
||||
"submit_query is cloud-only — the retrieval API is deprecated in "
|
||||
"favor of chat completions; use chat_completions in cloud mode."
|
||||
"favor of chat completions; use chat_completions instead."
|
||||
).submit_query(doc_id=doc_id, query=query, thinking=thinking)
|
||||
|
||||
def get_retrieval(self, retrieval_id: str) -> dict[str, Any]:
|
||||
@@ -269,55 +338,217 @@ class PageIndexClient:
|
||||
|
||||
Cloud-only: the cloud API marks this endpoint deprecated in favor of
|
||||
chat completions, so local mode does not implement it — raises
|
||||
PageIndexAPIError. Use ``chat_completions`` (cloud) instead.
|
||||
PageIndexAPIError. Use ``chat_completions`` instead.
|
||||
"""
|
||||
return self._require_cloud(
|
||||
"get_retrieval is cloud-only — the retrieval API is deprecated in "
|
||||
"favor of chat completions; use chat_completions in cloud mode."
|
||||
"favor of chat completions; use chat_completions instead."
|
||||
).get_retrieval(retrieval_id=retrieval_id)
|
||||
|
||||
# ---------- CHAT COMPLETIONS ----------
|
||||
|
||||
def chat_completions(
|
||||
self,
|
||||
messages: list[dict[str, str]],
|
||||
messages: Union[str, list[dict[str, str]]],
|
||||
stream: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
temperature: Optional[float] = None,
|
||||
stream_metadata: bool = False,
|
||||
enable_citations: bool = False,
|
||||
model: Optional[str] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict[str, Any], Iterator[str], Iterator[dict[str, Any]]]:
|
||||
"""
|
||||
PageIndex Chat Completions, scoped to specific PageIndex documents.
|
||||
PageIndex Chat Completions: document QA in one call.
|
||||
|
||||
Cloud: the hosted chat endpoint. Local: a managed document-QA agent
|
||||
run over the local tools against your own LLM backend's
|
||||
/chat/completions (requires ``pageindex[openai]``; the OpenAI SDK's
|
||||
usual env config — OPENAI_API_KEY, OPENAI_BASE_URL — selects the
|
||||
backend, so any OpenAI-compatible server works; a ``/`` in the
|
||||
model name means LiteLLM provider routing, so prefix ``openai/``
|
||||
when the backend itself serves slashed ids, e.g.
|
||||
``openai/Qwen/...`` on vLLM). The non-stream
|
||||
response carries the final answer only; streaming yields the
|
||||
agent's visible text as it is produced, including narration before
|
||||
tool calls. ``finish_reason`` reports loop completion ("stop") —
|
||||
the engine does not surface per-turn backend finish reasons. For
|
||||
the tool-use process and prompt-cache round-trip use
|
||||
``responses()`` or ``messages()``.
|
||||
|
||||
Args:
|
||||
messages: Conversation messages with 'role' and 'content' keys.
|
||||
messages: Conversation messages with 'role' and 'content' keys,
|
||||
or a bare query string (it becomes a single user message).
|
||||
Local also accepts system/developer messages — their content
|
||||
is appended to the managed system prompt.
|
||||
stream: Enable streaming responses.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
temperature: Sampling temperature (0.0-1.0).
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call and is part
|
||||
of the cached prompt prefix.
|
||||
temperature: Sampling temperature, passed through to the model.
|
||||
stream_metadata: With stream=True, yield chunk dicts instead of
|
||||
text pieces.
|
||||
enable_citations: Enable citation instructions in responses.
|
||||
enable_citations: Cloud-only — local mode raises (citations need
|
||||
block-level OCR data local mode does not store).
|
||||
model: Local only — backend model name (defaults to
|
||||
``retrieve_model``). The cloud endpoint selects its own.
|
||||
max_turns: Local only — cap on agent turns per call.
|
||||
|
||||
Returns:
|
||||
- stream=False: complete response dict ({'id', 'object', 'created',
|
||||
'choices', 'usage'})
|
||||
- stream=True, stream_metadata=False: iterator of text chunks
|
||||
- stream=True, stream_metadata=True: iterator of chunk dicts
|
||||
|
||||
Local: not yet supported — raises PageIndexAPIError. Agent-based
|
||||
local chat arrives in a later release.
|
||||
"""
|
||||
return self._require_cloud(
|
||||
"chat_completions is not yet supported in local mode — it arrives "
|
||||
"in a later release. Create the client with an api_key to use "
|
||||
"cloud chat."
|
||||
).chat_completions(
|
||||
if isinstance(messages, str):
|
||||
if not messages.strip():
|
||||
raise PageIndexAPIError(
|
||||
"messages must be a non-empty string or a list of "
|
||||
"message dicts.")
|
||||
messages = [{"role": "user", "content": messages}]
|
||||
from .cloud_api import CloudAPI
|
||||
if not isinstance(self._api, CloudAPI):
|
||||
from .local_chat import run_chat_completions
|
||||
return run_chat_completions(
|
||||
self, messages, stream=stream, doc_id=doc_id,
|
||||
temperature=temperature, stream_metadata=stream_metadata,
|
||||
enable_citations=enable_citations, model=model,
|
||||
max_turns=max_turns,
|
||||
)
|
||||
if model is not None or max_turns is not None:
|
||||
raise PageIndexAPIError(
|
||||
"model and max_turns are local-mode parameters — the cloud "
|
||||
"chat endpoint selects its own model."
|
||||
)
|
||||
return self._api.chat_completions(
|
||||
messages=messages, stream=stream, doc_id=doc_id,
|
||||
temperature=temperature, stream_metadata=stream_metadata,
|
||||
enable_citations=enable_citations,
|
||||
)
|
||||
|
||||
def responses(
|
||||
self,
|
||||
input: Union[str, list[dict[str, Any]]],
|
||||
model: Optional[str] = None,
|
||||
stream: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
instructions: Optional[str] = None,
|
||||
temperature: Optional[float] = None,
|
||||
top_p: Optional[float] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict[str, Any], Iterator[dict[str, Any]]]:
|
||||
"""
|
||||
Document QA over the OpenAI Responses protocol — the agentic surface.
|
||||
|
||||
Local only for now. Drives your backend's /responses end to end (no
|
||||
translation layer), so the ``output`` carries the whole process as
|
||||
standard items — messages, function calls, and function outputs
|
||||
(the SDK executes the tools). Append the returned ``output`` to your
|
||||
next call's ``input`` verbatim to keep provider prompt-cache prefix
|
||||
continuity and the agent's memory of what it already read.
|
||||
|
||||
Requires ``pageindex[openai]`` and a backend that supports the
|
||||
Responses API; backends that only speak chat.completions should use
|
||||
``chat_completions()``. Provider-prefixed models (``anthropic/…``)
|
||||
route through LiteLLM's chat.completions adapter and are therefore
|
||||
refused here — use ``chat_completions()`` or ``messages()`` for
|
||||
those.
|
||||
|
||||
Args:
|
||||
input: A user message string, or a list of Responses input items
|
||||
(round-trip prior ``output`` items here).
|
||||
model: Backend model name (defaults to ``retrieve_model``).
|
||||
stream: Yield Responses stream events as dicts — one logical
|
||||
response per call: per-turn backend lifecycle events are
|
||||
collapsed, sequence numbers are reassigned monotonically,
|
||||
and ``output_index`` is re-based onto the single logical
|
||||
``output``; tool outputs are emitted as
|
||||
``response.output_item.done`` events and the single final
|
||||
event is the terminal ``response.*`` for the run's status.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call and is part
|
||||
of the cached prompt prefix.
|
||||
instructions: Appended to the managed system prompt.
|
||||
temperature / top_p: Passed through to the model.
|
||||
max_turns: Cap on agent turns per call.
|
||||
"""
|
||||
from .cloud_api import CloudAPI
|
||||
if isinstance(self._api, CloudAPI):
|
||||
raise PageIndexAPIError(
|
||||
"responses is not available on PageIndex cloud yet — it is "
|
||||
"a local-mode surface for now."
|
||||
)
|
||||
from .local_chat import run_responses
|
||||
return run_responses(
|
||||
self, input, model=model, stream=stream, doc_id=doc_id,
|
||||
instructions=instructions, temperature=temperature, top_p=top_p,
|
||||
max_turns=max_turns,
|
||||
)
|
||||
|
||||
def messages(
|
||||
self,
|
||||
messages: Union[str, list[dict[str, Any]]],
|
||||
model: str,
|
||||
max_tokens: Optional[int] = None,
|
||||
stream: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
system: Optional[Union[str, list[dict[str, Any]]]] = None,
|
||||
temperature: Optional[float] = None,
|
||||
top_p: Optional[float] = None,
|
||||
top_k: Optional[int] = None,
|
||||
stop_sequences: Optional[list[str]] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict[str, Any], Iterator[Any]]:
|
||||
"""
|
||||
Document QA over the Anthropic Messages protocol — Claude-native.
|
||||
|
||||
Local only for now. Drives Anthropic's /v1/messages via the
|
||||
Anthropic SDK's own tool runner (requires ``pageindex[anthropic]``;
|
||||
ANTHROPIC_API_KEY selects the backend). ``tool_use``/``tool_result``
|
||||
round-trip is the format's native behavior: the response is the
|
||||
final message envelope with cross-turn aggregated ``usage`` plus a
|
||||
``messages`` field — the full new turn sequence, valid for verbatim
|
||||
append to your history. The managed system prompt carries a
|
||||
``cache_control`` breakpoint.
|
||||
|
||||
Args:
|
||||
messages: Native Messages-format history (including prior
|
||||
tool_use/tool_result blocks on round-trip), or a bare query
|
||||
string (it becomes a single user message).
|
||||
model: Required — there is no cross-vendor default to guess.
|
||||
max_tokens: Per-turn output budget the Messages API requires on
|
||||
the wire; the default is resolved per model (8192, or 4096
|
||||
for the claude-3 generation whose ceiling is lower) so the
|
||||
simple call needs only a question. Passed through.
|
||||
stream: Yield the Anthropic SDK's event stream across turns
|
||||
(its native event objects, including SDK-synthesized
|
||||
convenience events), one message sequence per turn.
|
||||
doc_id: Document ID or list of IDs to scope the conversation.
|
||||
Keep it identical across a conversation's calls — the
|
||||
targeting block it adds is re-set each call.
|
||||
system: Appended after the managed system blocks.
|
||||
temperature / top_p / top_k / stop_sequences: Passed through.
|
||||
max_turns: Cap on agent turns per call (default 10, like the
|
||||
OpenAI surfaces). A truncated run reports
|
||||
``stop_reason: "tool_use"`` and its ``messages`` remain
|
||||
valid for continuation.
|
||||
"""
|
||||
from .cloud_api import CloudAPI
|
||||
if isinstance(self._api, CloudAPI):
|
||||
raise PageIndexAPIError(
|
||||
"messages is not available on PageIndex cloud yet — it is "
|
||||
"a local-mode surface for now."
|
||||
)
|
||||
from .local_chat import run_messages
|
||||
return run_messages(
|
||||
self, messages, model=model, max_tokens=max_tokens,
|
||||
stream=stream, doc_id=doc_id, system=system,
|
||||
temperature=temperature, top_p=top_p, top_k=top_k,
|
||||
stop_sequences=stop_sequences, max_turns=max_turns,
|
||||
)
|
||||
|
||||
# ---------- DOCUMENT MANAGEMENT ----------
|
||||
|
||||
def get_document(self, doc_id: str) -> dict[str, Any]:
|
||||
@@ -365,6 +596,328 @@ class PageIndexClient:
|
||||
"""
|
||||
return self._api.list_documents(limit=limit, offset=offset, folder_id=folder_id)
|
||||
|
||||
# ---------- AGENT INTEGRATION ----------
|
||||
|
||||
def agent_tools(self, include_management: bool = False) -> list[Callable[..., str]]:
|
||||
"""
|
||||
Plain functions for any agent framework (LangChain, PydanticAI, ...).
|
||||
For the OpenAI / Claude Agent SDKs, prefer ``as_openai_tools()`` /
|
||||
``as_claude_mcp()``.
|
||||
|
||||
Cloud: the full cloud tool set, discovered live from the PageIndex
|
||||
MCP server when this method is called — one function per tool,
|
||||
signature and docstring synthesized from the server's schemas, calls
|
||||
executed from your process over MCP. Raises PageIndexAPIError if the
|
||||
server cannot be reached. Local: the built-in tools over the local
|
||||
store (``browse_documents``, ``get_document``,
|
||||
``get_document_structure``, ``get_page_content``).
|
||||
|
||||
Each function takes JSON-serializable arguments, returns a JSON
|
||||
string, and reports failures inside that JSON instead of raising.
|
||||
|
||||
Args:
|
||||
include_management (bool): Also expose tools that modify the
|
||||
library. Local: adds ``remove_document``. Cloud: by default
|
||||
only tools the server marks read-only are exposed; True
|
||||
exposes the server's complete list (upload, delete, ...).
|
||||
"""
|
||||
from .agent_tools import build_agent_tools
|
||||
return build_agent_tools(self, include_management)
|
||||
|
||||
def as_openai_tools(self, include_management: bool = False,
|
||||
hosted: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None) -> list:
|
||||
"""
|
||||
Tools for the OpenAI Agents SDK — pass to ``Agent(tools=...)``
|
||||
(or ``openai_agent_config()`` for all the Agent slots in one
|
||||
call).
|
||||
|
||||
Cloud (default): the full live read tool set (search, folders,
|
||||
images — as enabled for your key) as plain function tools,
|
||||
discovered from the PageIndex MCP server and executed from your
|
||||
process — works with any model backend. Pass ``hosted=True`` to
|
||||
hand the connection to OpenAI instead: one hosted MCP tool, tool
|
||||
calls executed server-side (lowest latency; requires an
|
||||
OpenAI-hosted model on the Responses API). The framework's own
|
||||
``MCPServerStreamableHttp`` — ``params={"url":
|
||||
f"{BASE_URL}/mcp?tools=read", "headers": {"Authorization":
|
||||
"Bearer <your PageIndex API key>"}}`` (drop ``?tools=read`` for
|
||||
the full tool set) — is the async-native alternative for its
|
||||
``mcp_servers=`` slot.
|
||||
|
||||
Local: the in-process tools, any model backend; ``hosted`` does
|
||||
not apply.
|
||||
|
||||
Requires ``openai-agents`` (``pip install 'pageindex[openai]'``),
|
||||
imported only when this method is called.
|
||||
|
||||
Args:
|
||||
include_management (bool): Also expose tools that modify the
|
||||
library (delete, upload). Default off: the in-process
|
||||
cloud default serves only server-annotated read-only
|
||||
tools, and ``hosted=True`` connects OpenAI to the
|
||||
read-only endpoint (``/mcp?tools=read``) instead.
|
||||
hosted (bool): Cloud only — hand the MCP connection to OpenAI
|
||||
for server-side tool execution (OpenAI models only).
|
||||
doc_id: Local only — restrict the tools to this document ID
|
||||
(or list of IDs), enforced at the tool layer: out-of-scope
|
||||
lookups return NOT_FOUND. Raises on cloud, where scoping
|
||||
is server-side.
|
||||
"""
|
||||
from .integrations.openai_agents import build_openai_tools
|
||||
return build_openai_tools(self, include_management, hosted,
|
||||
doc_ids=doc_id)
|
||||
|
||||
def _local_doc_scope(self, doc_id):
|
||||
"""doc_id for the tool layer: passed through locally (structural
|
||||
allowlist), dropped on cloud where scoping is server-side and the
|
||||
config helpers keep prompt-level targeting."""
|
||||
if not getattr(self, "api_key", None):
|
||||
return doc_id
|
||||
if doc_id is not None and not doc_id:
|
||||
# Cloud has no tool-layer allowlist to make an empty scope mean
|
||||
# "nothing"; dropping it would silently mean "everything".
|
||||
raise PageIndexAPIError(
|
||||
"doc_id is empty. Pass one or more document IDs, or omit "
|
||||
"doc_id to give the agent the whole library.")
|
||||
return None
|
||||
|
||||
def openai_agent_config(
|
||||
self,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
include_management: bool = False,
|
||||
model: Optional[str] = None,
|
||||
) -> dict[str, Any]:
|
||||
"""
|
||||
Document QA ``Agent`` kwargs for the OpenAI Agents SDK in one
|
||||
call::
|
||||
|
||||
agent = Agent(**client.openai_agent_config())
|
||||
|
||||
Sugar over the explicit form — ``agent_instructions`` (with
|
||||
``doc_id`` targeting) as the instructions and
|
||||
``as_openai_tools`` as the tools; local clients also carry their
|
||||
configured ``retrieve_model`` (cloud omits ``model`` so the
|
||||
framework default applies). To customize further, switch to
|
||||
those methods directly.
|
||||
|
||||
Args:
|
||||
doc_id: Document ID or list of IDs to target, as in
|
||||
``agent_instructions``. Local: also enforced at the tool
|
||||
layer, not just prompted. Cloud: prompt-level targeting
|
||||
(tool scoping is server-side).
|
||||
include_management (bool): Also expose tools that modify the
|
||||
library.
|
||||
model: Backend model name; overrides the local default.
|
||||
"""
|
||||
from .agent_tools import build_agent_instructions
|
||||
scope = self._local_doc_scope(doc_id)
|
||||
config: dict[str, Any] = {
|
||||
"name": "PageIndex",
|
||||
"instructions": build_agent_instructions(self, doc_id,
|
||||
scoped=scope is not None),
|
||||
"tools": self.as_openai_tools(include_management, doc_id=scope),
|
||||
}
|
||||
model = model or getattr(self, "retrieve_model", None)
|
||||
if model:
|
||||
config["model"] = model
|
||||
return config
|
||||
|
||||
def as_anthropic_tools(self, include_management: bool = False,
|
||||
asynchronous: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
) -> list:
|
||||
"""
|
||||
Runnable tools for the Anthropic SDK's tool runner — pass to
|
||||
``client.beta.messages.tool_runner(tools=...)`` (or
|
||||
``anthropic_runner_config()`` for the whole setup in one call).
|
||||
The default flavor is for the sync ``Anthropic`` client; pass
|
||||
``asynchronous=True`` for ``AsyncAnthropic``. For a manual
|
||||
``messages.create`` loop, serialize with
|
||||
``[tool.to_dict() for tool in ...]``.
|
||||
|
||||
Cloud: the full live read tool set (search, folders, images — as
|
||||
enabled for your key), discovered from the PageIndex MCP server
|
||||
and executed from your process; the server's input schemas pass
|
||||
through verbatim (MCP and the Messages API share the schema
|
||||
shape). The server-side alternative is the Messages API's beta
|
||||
MCP connector — ``mcp_servers=[{"type": "url", "name":
|
||||
"pageindex", "url": f"{BASE_URL}/mcp?tools=read",
|
||||
"authorization_token": <your PageIndex API key>}]`` (drop
|
||||
``?tools=read`` for the full tool set) — with no client-side
|
||||
tools involved. Local: the in-process tools — the same set
|
||||
``messages()`` runs internally.
|
||||
|
||||
Requires ``anthropic>=0.84.0``
|
||||
(``pip install 'pageindex[anthropic]'``), imported only when this
|
||||
method is called.
|
||||
|
||||
Args:
|
||||
include_management (bool): Also expose tools that modify the
|
||||
library. Local: adds ``remove_document``. Cloud: by default
|
||||
only tools the server marks read-only are exposed; True
|
||||
exposes the server's complete list (upload, delete, ...).
|
||||
asynchronous (bool): Build ``beta_async_tool`` runnables for
|
||||
``AsyncAnthropic`` (each tool call runs in a worker
|
||||
thread, keeping blocking I/O off your event loop). The
|
||||
sync and async runners each accept only their own flavor.
|
||||
doc_id: Local only — restrict the tools to this document ID
|
||||
(or list of IDs), enforced at the tool layer: out-of-scope
|
||||
lookups return NOT_FOUND. Raises on cloud, where scoping
|
||||
is server-side.
|
||||
"""
|
||||
from .integrations.anthropic_sdk import build_anthropic_tools
|
||||
return build_anthropic_tools(self, include_management, asynchronous,
|
||||
doc_ids=doc_id)
|
||||
|
||||
def anthropic_runner_config(
|
||||
self,
|
||||
model: str,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
include_management: bool = False,
|
||||
asynchronous: bool = False,
|
||||
max_tokens: Optional[int] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> dict[str, Any]:
|
||||
"""
|
||||
Document QA ``tool_runner`` kwargs for the Anthropic SDK in one
|
||||
call — only your ``messages`` remain::
|
||||
|
||||
runner = anthropic_client.beta.messages.tool_runner(
|
||||
**client.anthropic_runner_config(model="claude-sonnet-4-5"),
|
||||
messages=[{"role": "user", "content": "..."}],
|
||||
)
|
||||
|
||||
Sugar over the explicit form — ``agent_instructions`` (with
|
||||
``doc_id`` targeting) as the system prompt and
|
||||
``as_anthropic_tools`` as the tools — plus the same defaults
|
||||
``messages()`` applies: a per-model ``max_tokens`` and a
|
||||
``max_iterations`` bound of 10. To customize further, switch to
|
||||
those methods directly.
|
||||
|
||||
Args:
|
||||
model: Backend model name (also resolves the ``max_tokens``
|
||||
default).
|
||||
doc_id: Document ID or list of IDs to target, as in
|
||||
``agent_instructions``. Local: also enforced at the tool
|
||||
layer, not just prompted. Cloud: prompt-level targeting
|
||||
(tool scoping is server-side).
|
||||
include_management (bool): Also expose tools that modify the
|
||||
library.
|
||||
asynchronous (bool): Build async runnables for
|
||||
``AsyncAnthropic``.
|
||||
max_tokens: Per-turn output budget; default resolved per
|
||||
model.
|
||||
max_turns: Agent-loop bound; default 10.
|
||||
"""
|
||||
from .agent_tools import build_agent_instructions
|
||||
from .local_chat import _default_max_tokens
|
||||
scope = self._local_doc_scope(doc_id)
|
||||
return {
|
||||
"model": model,
|
||||
"max_tokens": (max_tokens if max_tokens is not None
|
||||
else _default_max_tokens(model)),
|
||||
"system": build_agent_instructions(self, doc_id,
|
||||
scoped=scope is not None),
|
||||
"tools": self.as_anthropic_tools(include_management, asynchronous,
|
||||
doc_id=scope),
|
||||
"max_iterations": max_turns if max_turns is not None else 10,
|
||||
}
|
||||
|
||||
def as_claude_mcp(self, include_management: bool = False,
|
||||
doc_id: Optional[Union[str, list[str]]] = None):
|
||||
"""
|
||||
``mcp_servers`` entry for the Claude Agent SDK.
|
||||
|
||||
Cloud: returns the remote PageIndex MCP config.
|
||||
``include_management`` picks the endpoint, so the URL itself is
|
||||
the gate — the default connects to the read-only endpoint
|
||||
(``/mcp?tools=read``: the server registers only read-only tools),
|
||||
``True`` connects to the full tool set. Local: returns an
|
||||
in-process SDK MCP server exposing the agent tools, gated the
|
||||
same way at registration (requires ``claude-agent-sdk``;
|
||||
``pip install 'pageindex[claude]'``). ``doc_id`` (local only)
|
||||
restricts those tools to that document ID (or list), enforced at
|
||||
the tool layer; it raises on cloud, where scoping is server-side.
|
||||
|
||||
Cloud hosts that surface MCP server instructions receive the same
|
||||
guidance ``agent_instructions()`` returns natively — passing both
|
||||
duplicates the text (harmless). ``system_prompt`` stays the
|
||||
recommended channel: it is guaranteed delivery, carries ``doc_id``
|
||||
targeting, and is the only channel local mode has.
|
||||
|
||||
Usage (or ``claude_agent_config()`` for all three slots in one
|
||||
call)::
|
||||
|
||||
options = ClaudeAgentOptions(
|
||||
system_prompt=client.agent_instructions(),
|
||||
mcp_servers={"pageindex": client.as_claude_mcp()},
|
||||
# Pre-approval only — the server itself is already gated.
|
||||
allowed_tools=["mcp__pageindex"],
|
||||
)
|
||||
"""
|
||||
from .integrations.claude_agent_sdk import build_claude_mcp
|
||||
return build_claude_mcp(self, include_management, doc_ids=doc_id)
|
||||
|
||||
def claude_agent_config(
|
||||
self,
|
||||
doc_id: Optional[Union[str, list[str]]] = None,
|
||||
include_management: bool = False,
|
||||
server_name: str = "pageindex",
|
||||
) -> dict[str, Any]:
|
||||
"""
|
||||
Document QA ``ClaudeAgentOptions`` kwargs in one call::
|
||||
|
||||
options = ClaudeAgentOptions(**client.claude_agent_config())
|
||||
|
||||
Sugar over the explicit form — the managed system prompt
|
||||
(``agent_instructions``) and the server entry (``as_claude_mcp``,
|
||||
itself the tool gate) with its ``allowed_tools`` pre-approval,
|
||||
one ``include_management`` and ``server_name`` applied
|
||||
everywhere. To customize (your own system prompt, extra
|
||||
servers), switch to those methods directly.
|
||||
|
||||
Args:
|
||||
doc_id: Document ID or list of IDs to target, as in
|
||||
``agent_instructions``. Local: also enforced at the tool
|
||||
layer, not just prompted. Cloud: prompt-level targeting
|
||||
(tool scoping is server-side).
|
||||
include_management (bool): Also allow tools that modify the
|
||||
library.
|
||||
server_name (str): Key the server is registered under.
|
||||
"""
|
||||
from .agent_tools import build_agent_instructions
|
||||
scope = self._local_doc_scope(doc_id)
|
||||
return {
|
||||
"system_prompt": build_agent_instructions(self, doc_id,
|
||||
scoped=scope is not None),
|
||||
"mcp_servers": {server_name: self.as_claude_mcp(
|
||||
include_management, doc_id=scope)},
|
||||
# Pre-approval only — the server itself is already gated (the
|
||||
# read-only endpoint on cloud, the registered set locally).
|
||||
"allowed_tools": [f"mcp__{server_name}"],
|
||||
}
|
||||
|
||||
def agent_instructions(self, doc_id: Optional[Union[str, list[str]]] = None) -> str:
|
||||
"""
|
||||
Orchestration guidance for document QA agents — pass as the agent's
|
||||
system prompt (or append to your own).
|
||||
|
||||
Cloud: the live instructions the PageIndex MCP server serves for
|
||||
your key's tool set, fetched over the same session as
|
||||
``agent_tools()`` — server-side guidance updates arrive without an
|
||||
SDK release. Raises PageIndexAPIError if the server cannot be
|
||||
reached. Local: the built-in guidance for the in-process tools.
|
||||
|
||||
With ``doc_id`` (str or list, same shape as ``chat_completions``),
|
||||
appends the target documents' names and metadata and directs the
|
||||
agent to work within them. Raises PageIndexAPIError if a doc_id does
|
||||
not exist, or if its name is shadowed by a newer same-name document
|
||||
(the name-addressed tools could not reach it).
|
||||
"""
|
||||
from .agent_tools import build_agent_instructions
|
||||
return build_agent_instructions(self, doc_id)
|
||||
|
||||
# ---------- FOLDER MANAGEMENT ----------
|
||||
|
||||
def create_folder(
|
||||
|
||||
@@ -55,7 +55,9 @@ class CloudAPI:
|
||||
returned in get_tree/get_ocr responses and list_documents entries. Defaults to None.
|
||||
|
||||
Returns:
|
||||
dict: {'doc_id': ...}
|
||||
dict: {'doc_id': ...} — plus 'name', the stored document name
|
||||
(a taken name gains a numeric suffix), when the server
|
||||
returns it.
|
||||
"""
|
||||
data = {'if_retrieval': True}
|
||||
if mode is not None:
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Framework adapters for the agent tools layer.
|
||||
|
||||
These modules import their target frameworks lazily, at call time — the
|
||||
frameworks are never required to install or import pageindex.
|
||||
"""
|
||||
@@ -0,0 +1,59 @@
|
||||
"""Anthropic SDK adapter for the tool runner's tools=... slot.
|
||||
|
||||
Cloud clients get one runnable tool per live cloud MCP tool — the server's
|
||||
input schemas pass through verbatim (MCP inputSchema and Messages API
|
||||
input_schema are the same shape), calls proxied over MCP. Local clients get
|
||||
the in-process tools — the same set messages() runs internally. Failed
|
||||
calls raise ToolError so the runner emits the tool_result with
|
||||
``is_error: true`` and the envelope as its content.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from typing import Any
|
||||
|
||||
from ..errors import PageIndexAPIError
|
||||
|
||||
|
||||
def build_anthropic_tools(client, include_management: bool = False,
|
||||
asynchronous: bool = False, doc_ids=None) -> list:
|
||||
try:
|
||||
from anthropic import beta_async_tool, beta_tool
|
||||
from anthropic.lib.tools import ToolError
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
"as_anthropic_tools requires the Anthropic SDK tool runner "
|
||||
"(anthropic>=0.84.0) — pip install -U anthropic (or pip install "
|
||||
"'pageindex[anthropic]')."
|
||||
) from exc
|
||||
from ..agent_tools import _tool_specs
|
||||
|
||||
def wrap(name, description, schema, invoke):
|
||||
"""One runnable tool in the caller's flavor: the sync runner and the
|
||||
async runner each accept only their own kind, and the async variant
|
||||
moves the blocking bridge/store call into a worker thread so it
|
||||
never blocks the caller's event loop."""
|
||||
def run(kwargs: dict) -> str:
|
||||
text, is_error = invoke(kwargs)
|
||||
if is_error:
|
||||
raise ToolError(text)
|
||||
return text
|
||||
|
||||
if asynchronous:
|
||||
async def _afn(**kwargs: Any) -> str:
|
||||
return await asyncio.to_thread(run, kwargs)
|
||||
|
||||
_afn.__name__ = name
|
||||
return beta_async_tool(_afn, name=name, description=description,
|
||||
input_schema=schema)
|
||||
|
||||
def _fn(**kwargs: Any) -> str:
|
||||
return run(kwargs)
|
||||
|
||||
_fn.__name__ = name
|
||||
return beta_tool(_fn, name=name, description=description,
|
||||
input_schema=schema)
|
||||
|
||||
return [wrap(*spec)
|
||||
for spec in _tool_specs(client, include_management,
|
||||
doc_ids=doc_ids)]
|
||||
@@ -0,0 +1,69 @@
|
||||
"""Claude Agent SDK adapter: one value for the mcp_servers slot.
|
||||
|
||||
Cloud clients get the remote PageIndex MCP config — the framework connects
|
||||
directly, and include_management picks the endpoint (the read-only
|
||||
``?tools=read`` URL by default); local clients get an in-process SDK MCP
|
||||
server over the same tool contract, gated the same way at registration.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from typing import Any
|
||||
|
||||
from .._version import sdk_version
|
||||
from ..errors import PageIndexAPIError
|
||||
|
||||
|
||||
def build_claude_mcp(client, include_management: bool = False, doc_ids=None):
|
||||
from ..agent_tools import _require_local_scope
|
||||
# The cloud branch returns a URL config — reject cloud doc_ids so they
|
||||
# are never silently dropped.
|
||||
_require_local_scope(client, doc_ids)
|
||||
if getattr(client, "api_key", None):
|
||||
# include_management picks the endpoint — the URL itself is the
|
||||
# gate (?tools=read serves only readOnlyHint-annotated tools).
|
||||
suffix = "" if include_management else "?tools=read"
|
||||
return {
|
||||
"type": "http",
|
||||
"url": f"{client.BASE_URL}/mcp{suffix}",
|
||||
"headers": {"Authorization": f"Bearer {client.api_key}"},
|
||||
}
|
||||
|
||||
try:
|
||||
from claude_agent_sdk import create_sdk_mcp_server, tool
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
"as_claude_mcp in local mode requires the Claude Agent SDK — "
|
||||
"pip install claude-agent-sdk (or pip install 'pageindex[claude]')."
|
||||
) from exc
|
||||
from ..agent_tools import (TOOL_CONTRACT, _local_description,
|
||||
_local_schema, call_tool, tool_names)
|
||||
|
||||
def make_handler(name: str):
|
||||
async def handler(arguments: dict[str, Any]) -> dict[str, Any]:
|
||||
text, is_error = await asyncio.to_thread(
|
||||
call_tool, client, name, arguments or {}, doc_ids
|
||||
)
|
||||
result: dict[str, Any] = {"content": [{"type": "text", "text": text}]}
|
||||
if is_error:
|
||||
result["is_error"] = True
|
||||
return result
|
||||
return handler
|
||||
|
||||
def tool_kwargs(name: str) -> dict:
|
||||
annotations = TOOL_CONTRACT[name].get("annotations")
|
||||
if not annotations:
|
||||
return {}
|
||||
try:
|
||||
from claude_agent_sdk import ToolAnnotations
|
||||
except ImportError:
|
||||
return {}
|
||||
return {"annotations": ToolAnnotations(**annotations)}
|
||||
|
||||
tools = [
|
||||
tool(name, _local_description(name),
|
||||
_local_schema(name), **tool_kwargs(name))(make_handler(name))
|
||||
for name in tool_names(include_management)
|
||||
]
|
||||
return create_sdk_mcp_server(name="pageindex", version=sdk_version(),
|
||||
tools=tools)
|
||||
@@ -0,0 +1,80 @@
|
||||
"""OpenAI Agents SDK adapter for the Agent(tools=...) slot.
|
||||
|
||||
Cloud clients default to the live read tool set as plain FunctionTools via
|
||||
the MCP bridge; pass hosted=True to use a single HostedMCPTool instead
|
||||
(the model connects to the PageIndex cloud MCP server from OpenAI's side —
|
||||
the read-only ``?tools=read`` endpoint by default). Local clients get the
|
||||
in-process tools wrapped as FunctionTools. Tools are built as FunctionTool
|
||||
directly so the contract/server JSON schema goes to the model verbatim —
|
||||
function_tool() would regenerate it from a Python signature, dropping
|
||||
items/enum/pattern/bounds and rejecting object-typed parameters.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
from typing import Any
|
||||
|
||||
from ..errors import PageIndexAPIError
|
||||
|
||||
|
||||
def build_openai_tools(client, include_management: bool = False,
|
||||
hosted: bool = False, doc_ids=None) -> list:
|
||||
try:
|
||||
from agents import FunctionTool, HostedMCPTool
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
"as_openai_tools requires the OpenAI Agents SDK — "
|
||||
"pip install openai-agents (or pip install 'pageindex[openai]')."
|
||||
) from exc
|
||||
from ..agent_tools import (_dumps, _failure, _require_local_scope,
|
||||
_tool_specs)
|
||||
# The hosted branch returns before _tool_specs — reject cloud doc_ids
|
||||
# here so they are never silently dropped.
|
||||
_require_local_scope(client, doc_ids)
|
||||
if getattr(client, "api_key", None) and hosted:
|
||||
# include_management picks the endpoint — the URL itself is the
|
||||
# gate (?tools=read serves only readOnlyHint-annotated tools), so
|
||||
# nothing needs the Responses API approval flow.
|
||||
suffix = "" if include_management else "?tools=read"
|
||||
return [HostedMCPTool(tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "pageindex",
|
||||
"server_url": f"{client.BASE_URL}/mcp{suffix}",
|
||||
"headers": {"Authorization": f"Bearer {client.api_key}"},
|
||||
"require_approval": "never",
|
||||
})]
|
||||
|
||||
def wrap(name, description, schema, invoke):
|
||||
async def on_invoke_tool(ctx: Any, args_json: str) -> str:
|
||||
# strict_json_schema is off, so the provider never validates the
|
||||
# payload; a malformed or non-object argument string must come
|
||||
# back as the guided error envelope — raising here aborts the
|
||||
# caller's whole run (hand-built FunctionTools have no
|
||||
# failure_error_function to hand the error back to the model).
|
||||
try:
|
||||
parsed = json.loads(args_json) if args_json else {}
|
||||
except ValueError:
|
||||
parsed = None
|
||||
if not isinstance(parsed, dict):
|
||||
payload, _ = _failure(
|
||||
f"Invalid arguments for {name}: expected a JSON object, "
|
||||
f"got: {(args_json or '')[:200]!r}", None,
|
||||
{"summary": "Malformed tool arguments",
|
||||
"options": [f"Re-send the {name} call with a JSON "
|
||||
"object of its parameters"]},
|
||||
"INVALID_INPUT")
|
||||
return _dumps(payload)
|
||||
arguments = {key: value for key, value in parsed.items()
|
||||
if value is not None}
|
||||
text, _ = await asyncio.to_thread(invoke, arguments)
|
||||
return text
|
||||
|
||||
return FunctionTool(name=name, description=description,
|
||||
params_json_schema=schema,
|
||||
on_invoke_tool=on_invoke_tool,
|
||||
strict_json_schema=False)
|
||||
|
||||
return [wrap(*spec)
|
||||
for spec in _tool_specs(client, include_management,
|
||||
doc_ids=doc_ids)]
|
||||
+26
-2
@@ -97,6 +97,9 @@ class LocalAPI:
|
||||
raise PageIndexAPIError(
|
||||
"Failed to submit document: PDF has no content. All pages are blank."
|
||||
)
|
||||
# Fail before paying for indexing when _1.._99 are all taken; the
|
||||
# binding name resolution happens again at save.
|
||||
self._unique_doc_name(os.path.basename(file_path))
|
||||
|
||||
try:
|
||||
if mode == "flash":
|
||||
@@ -115,7 +118,7 @@ class LocalAPI:
|
||||
doc_id = "pi-" + uuid.uuid4().hex
|
||||
meta = {
|
||||
"id": doc_id,
|
||||
"name": os.path.basename(file_path),
|
||||
"name": self._unique_doc_name(os.path.basename(file_path)),
|
||||
"description": description,
|
||||
"status": "completed",
|
||||
"createdAt": _now_iso(),
|
||||
@@ -129,7 +132,23 @@ class LocalAPI:
|
||||
from .utils import remove_fields
|
||||
self._store.save_document(
|
||||
doc_id, meta, remove_fields(structure, fields=["text"]), pages)
|
||||
return {"doc_id": doc_id}
|
||||
return {"doc_id": doc_id, "name": meta["name"]}
|
||||
|
||||
def _unique_doc_name(self, name: str) -> str:
|
||||
"""Mirror the cloud upload: a taken name gets _1.._99 appended,
|
||||
beyond that the submit is rejected."""
|
||||
taken = {meta.get("name") for meta in self._store.list_metas()}
|
||||
if name not in taken:
|
||||
return name
|
||||
base, ext = os.path.splitext(name)
|
||||
for num in range(1, 100):
|
||||
candidate = f"{base}_{num}{ext}"
|
||||
if candidate not in taken:
|
||||
return candidate
|
||||
raise PageIndexAPIError(
|
||||
"Failed to submit document: Too many files with similar names. "
|
||||
"Please use a different file name."
|
||||
)
|
||||
|
||||
@staticmethod
|
||||
def _extract_page_texts(file_path: str) -> list[str]:
|
||||
@@ -190,6 +209,11 @@ class LocalAPI:
|
||||
add_node_text(structure, pdf_pages)
|
||||
return structure
|
||||
|
||||
def raw_tree(self, doc_id: str) -> list | None:
|
||||
"""Stored tree verbatim — keeps start_index/end_index, which
|
||||
get_tree's cloud wire shape renames and drops."""
|
||||
return self._store.get_tree(doc_id)
|
||||
|
||||
def get_tree(self, doc_id: str, node_summary: bool = False,
|
||||
include_text: bool = True) -> dict[str, Any]:
|
||||
meta = self._require_doc(doc_id, "Failed to get tree result")
|
||||
|
||||
@@ -0,0 +1,840 @@
|
||||
"""Managed local chat: document-QA agents over the local tools.
|
||||
|
||||
Three methods, three backend protocols, routed 1:1: ``chat_completions``
|
||||
drives the backend's /chat/completions (any OpenAI-compatible backend,
|
||||
final answer only), ``responses`` drives /responses (process items are
|
||||
standard output; round-trip them for provider prompt-cache continuation and
|
||||
agent memory), ``messages`` drives Anthropic's /v1/messages via the SDK's
|
||||
own tool runner (tool_use/tool_result round-trip is the format's native
|
||||
behavior).
|
||||
|
||||
Content passes through untouched — the caller's messages, the model's
|
||||
answers, tool outputs. Native stop reasons pass through on ``messages``;
|
||||
the OpenAI engine's abstraction does not surface per-turn finish reasons,
|
||||
so ``chat_completions`` reports loop completion as ``"stop"``, while
|
||||
``responses`` reports the backend's terminal ``status`` where the wire
|
||||
surfaces one (recorded at the transport layer — the framework discards
|
||||
it). The SDK owns gatekeeping (structural validation), table-setting
|
||||
(managed instructions, tools, doc targeting), tool execution, and billing
|
||||
(usage aggregation, envelope ids).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import concurrent.futures
|
||||
import hashlib
|
||||
import json
|
||||
import queue
|
||||
import threading
|
||||
import time
|
||||
import uuid
|
||||
from typing import Any, Iterator, Optional, Union
|
||||
|
||||
from .agent_tools import AGENT_INSTRUCTIONS, doc_targeting_block
|
||||
from .errors import PageIndexAPIError
|
||||
|
||||
CHAT_HEADER = (
|
||||
"You are PageIndex by Vectify AI, a document-focused assistant. "
|
||||
"Be concise, never use emojis, and do not expose tool names."
|
||||
)
|
||||
|
||||
|
||||
# ── shared: prompt, doc targeting, validation, sync bridges ──
|
||||
|
||||
def _managed_instructions(extra_system: list[str]) -> str:
|
||||
return "\n\n".join([CHAT_HEADER, AGENT_INSTRUCTIONS, *extra_system])
|
||||
|
||||
|
||||
def _doc_block(client, doc_id) -> Optional[str]:
|
||||
if doc_id is None:
|
||||
return None
|
||||
if not isinstance(doc_id, (str, list)):
|
||||
raise PageIndexAPIError("doc_id must be a string or a list of "
|
||||
"strings.")
|
||||
doc_ids = [doc_id] if isinstance(doc_id, str) else list(doc_id)
|
||||
missing = []
|
||||
for one_id in doc_ids:
|
||||
try:
|
||||
client.get_document(one_id)
|
||||
except PageIndexAPIError:
|
||||
missing.append(str(one_id))
|
||||
if missing:
|
||||
raise PageIndexAPIError(
|
||||
"Documents not found or access denied: " + ", ".join(missing)
|
||||
)
|
||||
# scoped: the chat surfaces also pass doc_id into the tool layer, so
|
||||
# name resolution happens inside the allowlist — only a duplicate name
|
||||
# within the targeted set shadows.
|
||||
return doc_targeting_block(client, doc_id, scoped=True)
|
||||
|
||||
|
||||
def _system_text(content: Any) -> str:
|
||||
"""Text of a system/developer message: a string, or text parts joined."""
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
if isinstance(content, list):
|
||||
texts = [part.get("text") for part in content
|
||||
if isinstance(part, dict) and isinstance(part.get("text"), str)]
|
||||
if texts:
|
||||
return "\n".join(texts)
|
||||
raise PageIndexAPIError(
|
||||
"system message content must be a string or a list of text parts."
|
||||
)
|
||||
|
||||
|
||||
def _split_chat_messages(messages) -> "tuple[list[str], list[dict]]":
|
||||
"""Validate the chat_completions surface's messages: system/developer
|
||||
content joins the managed instructions; user/assistant history passes
|
||||
through. Tool-history round-trips belong to responses()/messages()."""
|
||||
if not isinstance(messages, list) or not messages:
|
||||
raise PageIndexAPIError("messages must be a non-empty list.")
|
||||
system_texts: list[str] = []
|
||||
history: list[dict] = []
|
||||
for message in messages:
|
||||
if not isinstance(message, dict) or "role" not in message:
|
||||
raise PageIndexAPIError(
|
||||
"Each message must be a dict with 'role' and 'content'.")
|
||||
role = message["role"]
|
||||
if role in ("system", "developer"):
|
||||
system_texts.append(_system_text(message.get("content")))
|
||||
elif role in ("user", "assistant"):
|
||||
content = message.get("content")
|
||||
if not isinstance(content, str):
|
||||
raise PageIndexAPIError(
|
||||
"chat_completions content must be a string; for "
|
||||
"structured items use responses() or messages()."
|
||||
)
|
||||
history.append({"role": role, "content": content})
|
||||
else:
|
||||
raise PageIndexAPIError(
|
||||
f"Unsupported role for chat_completions: {role!r}. Tool "
|
||||
"history round-trips belong to responses() or messages()."
|
||||
)
|
||||
if not history:
|
||||
raise PageIndexAPIError("messages must contain a user or assistant "
|
||||
"message.")
|
||||
return system_texts, history
|
||||
|
||||
|
||||
def _run_sync(coro):
|
||||
try:
|
||||
asyncio.get_running_loop()
|
||||
except RuntimeError:
|
||||
return asyncio.run(coro)
|
||||
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as pool:
|
||||
return pool.submit(asyncio.run, coro).result()
|
||||
|
||||
|
||||
_SENTINEL = object()
|
||||
|
||||
|
||||
def _stream_sync(agen_factory) -> Iterator[Any]:
|
||||
"""Drive an async generator from a background thread; yield synchronously.
|
||||
|
||||
Closing the iterator cancels the run between items: the pump stops, and
|
||||
the async generator's cleanup cancels the underlying agent task, so no
|
||||
further model turns or tool executions start. An in-flight backend
|
||||
request cannot be aborted mid-turn.
|
||||
"""
|
||||
items: "queue.Queue[Any]" = queue.Queue(maxsize=32)
|
||||
cancelled = threading.Event()
|
||||
|
||||
def deliver(item) -> bool:
|
||||
while not cancelled.is_set():
|
||||
try:
|
||||
items.put(item, timeout=0.1)
|
||||
return True
|
||||
except queue.Full:
|
||||
continue
|
||||
return False
|
||||
|
||||
def pump():
|
||||
async def consume():
|
||||
agen = agen_factory()
|
||||
|
||||
async def drain():
|
||||
async for item in agen:
|
||||
if not deliver(item):
|
||||
break
|
||||
|
||||
# The watchdog lets cancellation land even while drain() is
|
||||
# awaiting the backend — a plain async-for would only notice
|
||||
# between items.
|
||||
task = asyncio.ensure_future(drain())
|
||||
try:
|
||||
while not task.done():
|
||||
if cancelled.is_set():
|
||||
task.cancel()
|
||||
break
|
||||
await asyncio.sleep(0.05)
|
||||
try:
|
||||
await task
|
||||
except asyncio.CancelledError:
|
||||
pass
|
||||
finally:
|
||||
await agen.aclose()
|
||||
|
||||
try:
|
||||
asyncio.run(consume())
|
||||
except BaseException as exc: # re-raised on the consumer thread
|
||||
deliver(exc)
|
||||
return
|
||||
deliver(_SENTINEL)
|
||||
|
||||
threading.Thread(target=pump, daemon=True).start()
|
||||
try:
|
||||
while True:
|
||||
item = items.get()
|
||||
if item is _SENTINEL:
|
||||
return
|
||||
if isinstance(item, BaseException):
|
||||
raise item
|
||||
yield item
|
||||
finally:
|
||||
cancelled.set()
|
||||
|
||||
|
||||
# ── OpenAI engine (chat_completions / responses) ──
|
||||
|
||||
def _require_openai_agents(method: str) -> None:
|
||||
try:
|
||||
import agents # noqa: F401
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"{method} in local mode requires the OpenAI Agents SDK — "
|
||||
"pip install openai-agents (or pip install 'pageindex[openai]')."
|
||||
) from exc
|
||||
|
||||
|
||||
def _openai_model(protocol: str, model_name: str):
|
||||
"""The backend protocol driver — the seam tests replace with a fake.
|
||||
|
||||
``litellm/<provider>/<model>`` (the client's normalized retrieve_model
|
||||
form) and bare ``<provider>/<model>`` paths drive the provider through
|
||||
LiteLLM — chat.completions only, so the responses protocol refuses them
|
||||
instead of silently downgrading; a first segment LiteLLM does not know
|
||||
(a HuggingFace repo id like ``Qwen/...``) is refused with the
|
||||
``openai/`` escape instead of failing inside LiteLLM at request time;
|
||||
an ``openai/`` prefix strips to the OpenAI SDK; bare names go to the
|
||||
OpenAI SDK as-is."""
|
||||
if "/" in model_name and not model_name.startswith("openai/"):
|
||||
if protocol == "responses":
|
||||
raise PageIndexAPIError(
|
||||
f"responses() cannot drive "
|
||||
f"'{model_name.removeprefix('litellm/')}': provider-prefixed "
|
||||
"models route through LiteLLM, which speaks chat.completions, "
|
||||
"not the Responses API. Use chat_completions() (or messages() "
|
||||
"for Anthropic models), or point OPENAI_BASE_URL at a "
|
||||
"Responses-capable backend and use a bare or "
|
||||
"'openai/'-prefixed model name."
|
||||
)
|
||||
try:
|
||||
from agents.extensions.models.litellm_model import LitellmModel
|
||||
import litellm
|
||||
except ImportError:
|
||||
raise PageIndexAPIError(
|
||||
f"'{model_name}' routes through LiteLLM, but litellm is not "
|
||||
"installed. Run: pip install 'litellm>=1.30'"
|
||||
)
|
||||
wire = model_name.removeprefix("litellm/")
|
||||
providers = getattr(litellm, "provider_list", None)
|
||||
if providers and wire.split("/", 1)[0] not in providers:
|
||||
raise PageIndexAPIError(
|
||||
f"'{wire}' routes through LiteLLM, but "
|
||||
f"'{wire.split('/', 1)[0]}' is not a LiteLLM provider. For an "
|
||||
"OpenAI-compatible server (vLLM, TGI, Ollama) serving this "
|
||||
f"model id, use 'openai/{wire}' and point OPENAI_BASE_URL "
|
||||
"at the server."
|
||||
)
|
||||
return LitellmModel(wire)
|
||||
import openai
|
||||
model_name = model_name.removeprefix("openai/")
|
||||
try:
|
||||
backend = openai.AsyncOpenAI()
|
||||
except openai.OpenAIError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The OpenAI backend is not configured: {exc}") from exc
|
||||
if protocol == "chat":
|
||||
from agents.models.openai_chatcompletions import (
|
||||
OpenAIChatCompletionsModel)
|
||||
return OpenAIChatCompletionsModel(model_name, backend)
|
||||
from agents.models.openai_responses import OpenAIResponsesModel
|
||||
return OpenAIResponsesModel(model_name, openai_client=backend)
|
||||
|
||||
|
||||
def _reported_model(model_name: str) -> str:
|
||||
"""The name the provider actually serves — routing prefixes stripped."""
|
||||
return model_name.removeprefix("litellm/").removeprefix("openai/")
|
||||
|
||||
|
||||
def _openai_agent(client, protocol: str, model_name: str, instructions: str,
|
||||
temperature, top_p, doc_ids=None):
|
||||
from agents import Agent, ModelSettings
|
||||
from .integrations.openai_agents import build_openai_tools
|
||||
return Agent(
|
||||
name="PageIndex",
|
||||
instructions=instructions,
|
||||
tools=build_openai_tools(client, doc_ids=doc_ids),
|
||||
model=_openai_model(protocol, model_name),
|
||||
model_settings=ModelSettings(temperature=temperature, top_p=top_p),
|
||||
)
|
||||
|
||||
|
||||
def _validate_max_turns(max_turns) -> None:
|
||||
if max_turns is not None and (not isinstance(max_turns, int)
|
||||
or max_turns < 1):
|
||||
raise PageIndexAPIError("max_turns must be a positive integer.")
|
||||
|
||||
|
||||
def _conversation_group_id(model_name: str, instructions: str, items) -> str:
|
||||
"""Stable per-conversation cache-routing key: openai-agents hashes
|
||||
RunConfig.group_id into the OpenAI prompt_cache_key, and without one it
|
||||
stamps every run with a fresh key, tagging a round-tripped prefix as a
|
||||
different cache group. Keyed on the prefix identity — model,
|
||||
instructions, first conversation item — so a conversation's
|
||||
continuations share one route without pooling unrelated conversations.
|
||||
Callers pass the conversation's own items, never the SDK-prepended
|
||||
doc-targeting block: that block is byte-identical for every
|
||||
conversation about a document and would pool them all under one key."""
|
||||
seed = json.dumps([model_name, instructions,
|
||||
items[0] if items else None],
|
||||
sort_keys=True, default=str)
|
||||
return "pageindex-" + hashlib.sha256(seed.encode()).hexdigest()[:16]
|
||||
|
||||
|
||||
def _run_kwargs(max_turns, group_id: str) -> dict:
|
||||
# Managed runs never export traces — the caller opted into document QA,
|
||||
# not telemetry.
|
||||
from agents import RunConfig
|
||||
kwargs: dict = {"run_config": RunConfig(tracing_disabled=True,
|
||||
group_id=group_id)}
|
||||
if max_turns is not None:
|
||||
kwargs["max_turns"] = max_turns
|
||||
return kwargs
|
||||
|
||||
|
||||
def _record_response_status(agent, recorded: dict) -> None:
|
||||
"""Capture each turn's terminal Response status at the transport client:
|
||||
openai-agents' non-streaming path discards Response.status, so a final
|
||||
turn truncated at the output cap would otherwise report as a clean
|
||||
completion. No-op for backends without an OpenAI responses resource
|
||||
(the streaming path records from lifecycle events instead)."""
|
||||
responses = getattr(getattr(getattr(agent, "model", None), "_client", None),
|
||||
"responses", None)
|
||||
create = getattr(responses, "create", None)
|
||||
if create is None:
|
||||
return
|
||||
|
||||
async def recording_create(*args, **kwargs):
|
||||
response = await create(*args, **kwargs)
|
||||
if getattr(response, "status", None):
|
||||
recorded["status"] = response.status
|
||||
for field in ("incomplete_details", "error"):
|
||||
value = getattr(response, field, None)
|
||||
recorded[field] = (value.model_dump(mode="json")
|
||||
if hasattr(value, "model_dump") else value)
|
||||
return response
|
||||
|
||||
responses.create = recording_create
|
||||
|
||||
|
||||
async def _aclose_backend(agent) -> None:
|
||||
"""Close the per-call AsyncOpenAI client before its event loop ends —
|
||||
otherwise httpx tears down pooled connections on a closed loop and
|
||||
emits 'Task exception was never retrieved' noise."""
|
||||
backend = getattr(getattr(agent, "model", None), "_client", None)
|
||||
close = getattr(backend, "close", None)
|
||||
if close is not None:
|
||||
try:
|
||||
await close()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
async def _run_closing(agent, coro):
|
||||
try:
|
||||
return await coro
|
||||
finally:
|
||||
await _aclose_backend(agent)
|
||||
|
||||
|
||||
def _wrap_max_turns(max_turns) -> PageIndexAPIError:
|
||||
limit = max_turns if max_turns is not None else "the default limit"
|
||||
return PageIndexAPIError(
|
||||
f"The agent did not finish within max_turns ({limit}). Raise "
|
||||
"max_turns, or narrow the question."
|
||||
)
|
||||
|
||||
|
||||
def _openai_usage(raw_responses) -> dict:
|
||||
prompt = sum(r.usage.input_tokens for r in raw_responses)
|
||||
completion = sum(r.usage.output_tokens for r in raw_responses)
|
||||
return {"prompt_tokens": prompt, "completion_tokens": completion,
|
||||
"total_tokens": prompt + completion}
|
||||
|
||||
|
||||
def run_chat_completions(client, messages, stream: bool = False,
|
||||
doc_id=None, temperature: Optional[float] = None,
|
||||
stream_metadata: bool = False,
|
||||
enable_citations: bool = False,
|
||||
model: Optional[str] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict, Iterator[str], Iterator[dict]]:
|
||||
if enable_citations:
|
||||
raise PageIndexAPIError(
|
||||
"enable_citations is cloud-only — citations need block-level OCR "
|
||||
"data that local mode does not store."
|
||||
)
|
||||
_require_openai_agents("chat_completions")
|
||||
_validate_max_turns(max_turns)
|
||||
system_texts, history = _split_chat_messages(messages)
|
||||
block = _doc_block(client, doc_id)
|
||||
items = ([{"role": "user", "content": block}] if block else []) + history
|
||||
model_name = model or client.retrieve_model
|
||||
# litellm/ and openai/ are the SDK's routing markers, not model names —
|
||||
# report the name the provider actually serves.
|
||||
reported_model = _reported_model(model_name)
|
||||
managed = _managed_instructions(system_texts)
|
||||
agent = _openai_agent(client, "chat", model_name, managed,
|
||||
temperature, None, doc_ids=doc_id)
|
||||
run_kwargs = _run_kwargs(max_turns,
|
||||
_conversation_group_id(model_name, managed,
|
||||
history))
|
||||
import openai
|
||||
from agents import Runner
|
||||
from agents.exceptions import AgentsException, MaxTurnsExceeded
|
||||
if not stream:
|
||||
try:
|
||||
result = _run_sync(_run_closing(agent,
|
||||
Runner.run(agent, input=items, **run_kwargs)))
|
||||
except MaxTurnsExceeded as exc:
|
||||
raise _wrap_max_turns(max_turns) from exc
|
||||
except AgentsException as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The agent backend failed: {exc}") from exc
|
||||
except openai.OpenAIError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
return {
|
||||
"id": f"chatcmpl-{uuid.uuid4().hex}",
|
||||
"object": "chat.completion",
|
||||
"created": int(time.time()),
|
||||
"model": reported_model,
|
||||
"choices": [{
|
||||
"index": 0,
|
||||
"message": {"role": "assistant",
|
||||
"content": result.final_output or ""},
|
||||
"finish_reason": "stop",
|
||||
}],
|
||||
"usage": _openai_usage(result.raw_responses),
|
||||
}
|
||||
|
||||
chat_id = f"chatcmpl-{uuid.uuid4().hex}"
|
||||
created = int(time.time())
|
||||
|
||||
def chunk(delta: dict, finish=None) -> dict:
|
||||
return {
|
||||
"id": chat_id, "object": "chat.completion.chunk",
|
||||
"created": created, "model": reported_model,
|
||||
"choices": [{"index": 0, "delta": delta,
|
||||
"finish_reason": finish}],
|
||||
}
|
||||
|
||||
async def agen():
|
||||
from openai.types.responses import ResponseTextDeltaEvent
|
||||
streamed = Runner.run_streamed(agent, input=items, **run_kwargs)
|
||||
completed = False
|
||||
# First yield inside the try: a consumer that stops on the opening
|
||||
# chunk must still tear the run down via the finally below.
|
||||
try:
|
||||
yield chunk({"role": "assistant", "content": ""})
|
||||
async for event in streamed.stream_events():
|
||||
if (event.type == "raw_response_event"
|
||||
and isinstance(event.data, ResponseTextDeltaEvent)):
|
||||
yield chunk({"content": event.data.delta})
|
||||
completed = True
|
||||
except MaxTurnsExceeded as exc:
|
||||
raise _wrap_max_turns(max_turns) from exc
|
||||
except AgentsException as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The agent backend failed: {exc}") from exc
|
||||
except openai.OpenAIError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
finally:
|
||||
if not completed and hasattr(streamed, "cancel"):
|
||||
streamed.cancel() # abandoned/failed: stop the agent task
|
||||
await _aclose_backend(agent)
|
||||
yield chunk({}, finish="stop")
|
||||
yield {
|
||||
"id": chat_id, "object": "chat.completion.chunk",
|
||||
"created": created, "model": reported_model, "choices": [],
|
||||
"usage": _openai_usage(streamed.raw_responses),
|
||||
}
|
||||
|
||||
if stream_metadata:
|
||||
return _stream_sync(agen)
|
||||
return (piece["choices"][0]["delta"]["content"]
|
||||
for piece in _stream_sync(agen)
|
||||
if piece.get("choices")
|
||||
and "content" in piece["choices"][0]["delta"]
|
||||
and piece["choices"][0]["delta"]["content"])
|
||||
|
||||
|
||||
def run_responses(client, input, model: Optional[str] = None,
|
||||
stream: bool = False, doc_id=None,
|
||||
instructions: Optional[str] = None,
|
||||
temperature: Optional[float] = None,
|
||||
top_p: Optional[float] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict, Iterator[dict]]:
|
||||
_require_openai_agents("responses")
|
||||
_validate_max_turns(max_turns)
|
||||
if isinstance(input, str) and input.strip():
|
||||
items = [{"role": "user", "content": input}]
|
||||
elif (isinstance(input, list) and input
|
||||
and all(isinstance(item, dict) for item in input)):
|
||||
items = list(input)
|
||||
else:
|
||||
raise PageIndexAPIError("input must be a non-empty string or list "
|
||||
"of item dicts.")
|
||||
block = _doc_block(client, doc_id)
|
||||
conversation = items
|
||||
if block:
|
||||
items = [{"role": "user", "content": block}] + items
|
||||
extra = [instructions] if instructions else []
|
||||
model_name = model or client.retrieve_model
|
||||
managed = _managed_instructions(extra)
|
||||
agent = _openai_agent(client, "responses", model_name, managed,
|
||||
temperature, top_p, doc_ids=doc_id)
|
||||
run_kwargs = _run_kwargs(max_turns,
|
||||
_conversation_group_id(model_name, managed,
|
||||
conversation))
|
||||
recorded: dict = {}
|
||||
import openai
|
||||
from agents import Runner
|
||||
from agents.exceptions import AgentsException, MaxTurnsExceeded
|
||||
|
||||
def envelope(output: list, raw_responses) -> dict:
|
||||
usage = _openai_usage(raw_responses)
|
||||
return {
|
||||
"id": f"resp_{uuid.uuid4().hex}",
|
||||
"object": "response",
|
||||
"created_at": int(time.time()),
|
||||
"model": _reported_model(model_name),
|
||||
"status": recorded.get("status") or "completed",
|
||||
"output": output,
|
||||
"usage": {"input_tokens": usage["prompt_tokens"],
|
||||
"output_tokens": usage["completion_tokens"],
|
||||
"total_tokens": usage["total_tokens"]},
|
||||
"instructions": managed,
|
||||
"tools": [{"type": "function", "name": tool.name,
|
||||
"description": tool.description,
|
||||
"parameters": tool.params_json_schema,
|
||||
"strict": getattr(tool, "strict_json_schema", True)}
|
||||
for tool in agent.tools],
|
||||
"tool_choice": "auto",
|
||||
"parallel_tool_calls": True,
|
||||
"temperature": temperature,
|
||||
"top_p": top_p,
|
||||
"max_output_tokens": None,
|
||||
"error": recorded.get("error"),
|
||||
"incomplete_details": recorded.get("incomplete_details"),
|
||||
"metadata": None,
|
||||
}
|
||||
|
||||
if not stream:
|
||||
_record_response_status(agent, recorded)
|
||||
try:
|
||||
result = _run_sync(_run_closing(agent,
|
||||
Runner.run(agent, input=[dict(item) for item in items],
|
||||
**run_kwargs)))
|
||||
except MaxTurnsExceeded as exc:
|
||||
raise _wrap_max_turns(max_turns) from exc
|
||||
except AgentsException as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The agent backend failed: {exc}") from exc
|
||||
except openai.OpenAIError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
output = result.to_input_list()[len(items):]
|
||||
return envelope(output, result.raw_responses)
|
||||
|
||||
# One logical response per call: per-turn backend lifecycle events
|
||||
# (created/completed/...) are collapsed — forwarding them verbatim would
|
||||
# end a canonical consumer at the first turn — and sequence numbers are
|
||||
# reassigned monotonically across the whole run.
|
||||
lifecycle = {"response.created", "response.in_progress",
|
||||
"response.completed", "response.failed",
|
||||
"response.incomplete", "response.queued"}
|
||||
|
||||
async def agen():
|
||||
streamed = Runner.run_streamed(agent,
|
||||
input=[dict(item) for item in items],
|
||||
**run_kwargs)
|
||||
sequence = 0
|
||||
# output_index addresses an item's position in the logical
|
||||
# response.output (the final envelope's list). Backend events
|
||||
# carry per-turn indexes that restart at 0 each turn, so they are
|
||||
# re-based by the count of items already committed by prior turns
|
||||
# — and the tool outputs the SDK injects between turns take the
|
||||
# next slot on that same axis.
|
||||
output_offset = 0
|
||||
completed = False
|
||||
try:
|
||||
async for event in streamed.stream_events():
|
||||
if event.type == "raw_response_event":
|
||||
data = event.data.model_dump(exclude_unset=True)
|
||||
if data.get("type") in lifecycle:
|
||||
if data["type"] in ("response.completed",
|
||||
"response.incomplete",
|
||||
"response.failed"):
|
||||
# Per-turn terminal state; the last turn's wins
|
||||
# and feeds the final envelope below.
|
||||
state = data.get("response") or {}
|
||||
for field in ("status", "incomplete_details",
|
||||
"error"):
|
||||
recorded[field] = state.get(field)
|
||||
output_offset += len(state.get("output") or [])
|
||||
continue
|
||||
if isinstance(data.get("output_index"), int):
|
||||
data["output_index"] += output_offset
|
||||
sequence += 1
|
||||
data["sequence_number"] = sequence
|
||||
yield data
|
||||
elif (event.type == "run_item_stream_event"
|
||||
and event.item.type == "tool_call_output_item"):
|
||||
# We are the tool executor, so we emit the output item
|
||||
# the way the platform streams its own server-side tools.
|
||||
sequence += 1
|
||||
yield {"type": "response.output_item.done",
|
||||
"output_index": output_offset,
|
||||
"sequence_number": sequence,
|
||||
"item": dict(event.item.to_input_item())}
|
||||
output_offset += 1
|
||||
completed = True
|
||||
except MaxTurnsExceeded as exc:
|
||||
raise _wrap_max_turns(max_turns) from exc
|
||||
except AgentsException as exc:
|
||||
if recorded.get("status") not in ("failed", "incomplete"):
|
||||
raise PageIndexAPIError(
|
||||
f"The agent backend failed: {exc}") from exc
|
||||
# response.failed / response.incomplete: the engine re-raises
|
||||
# the backend's terminal state as an exception — it is a
|
||||
# protocol event, emitted as the terminal event below.
|
||||
completed = True
|
||||
except openai.OpenAIError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
finally:
|
||||
if not completed and hasattr(streamed, "cancel"):
|
||||
streamed.cancel() # abandoned/failed: stop the agent task
|
||||
await _aclose_backend(agent)
|
||||
output = streamed.to_input_list()[len(items):]
|
||||
sequence += 1
|
||||
status = recorded.get("status") or "completed"
|
||||
terminal = {"incomplete": "response.incomplete",
|
||||
"failed": "response.failed"}.get(status,
|
||||
"response.completed")
|
||||
yield {"type": terminal, "sequence_number": sequence,
|
||||
"response": envelope(output, streamed.raw_responses)}
|
||||
|
||||
return _stream_sync(agen)
|
||||
|
||||
|
||||
# ── Anthropic engine (messages) ──
|
||||
|
||||
def _require_anthropic() -> None:
|
||||
try:
|
||||
import anthropic # noqa: F401
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
"messages in local mode requires the Anthropic SDK — "
|
||||
"pip install anthropic (or pip install 'pageindex[anthropic]')."
|
||||
) from exc
|
||||
try:
|
||||
from anthropic import beta_tool # noqa: F401
|
||||
from anthropic.lib.tools import ToolError # noqa: F401
|
||||
except ImportError as exc:
|
||||
raise PageIndexAPIError(
|
||||
"messages in local mode requires anthropic >= 0.84.0 (the tool "
|
||||
"runner with ToolError) — pip install -U anthropic."
|
||||
) from exc
|
||||
|
||||
|
||||
def _anthropic_client():
|
||||
"""The backend client — the seam tests replace with a fake transport."""
|
||||
import anthropic
|
||||
return anthropic.Anthropic()
|
||||
|
||||
|
||||
def _anthropic_system(extra_system, block: Optional[str]) -> list[dict]:
|
||||
"""System blocks: cache_control marks the stable managed prefix only
|
||||
(the API allows 4 breakpoints total — the varying doc block and caller
|
||||
blocks must not consume the budget); the doc block and caller system
|
||||
content follow as their own blocks."""
|
||||
blocks = [{"type": "text",
|
||||
"text": CHAT_HEADER + "\n\n" + AGENT_INSTRUCTIONS,
|
||||
"cache_control": {"type": "ephemeral"}}]
|
||||
if block:
|
||||
blocks.append({"type": "text", "text": block})
|
||||
if extra_system is None:
|
||||
return blocks
|
||||
if isinstance(extra_system, str):
|
||||
if extra_system.strip():
|
||||
blocks.append({"type": "text", "text": extra_system})
|
||||
return blocks
|
||||
if isinstance(extra_system, list):
|
||||
return blocks + list(extra_system)
|
||||
raise PageIndexAPIError("system must be a string or a list of blocks.")
|
||||
|
||||
|
||||
def _dump_block(block) -> Any:
|
||||
"""A content block as a plain JSON dict, minus SDK-internal fields the
|
||||
API rejects (ParsedBetaTextBlock.__api_exclude__, e.g. parsed_output)."""
|
||||
if hasattr(block, "model_dump"):
|
||||
exclude = getattr(type(block), "__api_exclude__", None)
|
||||
return block.model_dump(mode="json",
|
||||
exclude=set(exclude) if exclude else None)
|
||||
return block
|
||||
|
||||
|
||||
def _dump_message(message) -> dict:
|
||||
message = dict(message)
|
||||
content = message.get("content")
|
||||
if isinstance(content, list):
|
||||
message["content"] = [_dump_block(item) for item in content]
|
||||
return message
|
||||
|
||||
|
||||
def _anthropic_usage(turns, final_usage: dict) -> dict:
|
||||
"""The final turn's native usage dict with the token counters replaced
|
||||
by cross-turn sums (None-safe); all other native fields survive."""
|
||||
totals = dict(final_usage)
|
||||
for field in ("input_tokens", "output_tokens",
|
||||
"cache_creation_input_tokens", "cache_read_input_tokens"):
|
||||
values = [getattr(turn.usage, field, None) for turn in turns]
|
||||
counted = [value for value in values if isinstance(value, int)]
|
||||
if counted:
|
||||
totals[field] = sum(counted)
|
||||
return totals
|
||||
|
||||
|
||||
_CLAUDE_4096_MODELS = ("claude-3-opus", "claude-3-sonnet", "claude-3-haiku",
|
||||
"claude-3-5-sonnet-20240620")
|
||||
|
||||
|
||||
def _default_max_tokens(model: str) -> int:
|
||||
"""The wire-required per-turn budget when the caller sets none: 8192,
|
||||
except the claude-3 generation whose output ceiling is 4096."""
|
||||
return 4096 if model.startswith(_CLAUDE_4096_MODELS) else 8192
|
||||
|
||||
|
||||
def run_messages(client, messages, model: str,
|
||||
max_tokens: Optional[int] = None,
|
||||
stream: bool = False, doc_id=None, system=None,
|
||||
temperature: Optional[float] = None,
|
||||
top_p: Optional[float] = None,
|
||||
top_k: Optional[int] = None,
|
||||
stop_sequences: Optional[list[str]] = None,
|
||||
max_turns: Optional[int] = None,
|
||||
) -> Union[dict, Iterator[Any]]:
|
||||
from .integrations.anthropic_sdk import build_anthropic_tools
|
||||
|
||||
_require_anthropic()
|
||||
import anthropic
|
||||
_validate_max_turns(max_turns)
|
||||
if isinstance(messages, str) and messages.strip():
|
||||
messages = [{"role": "user", "content": messages}]
|
||||
if (not isinstance(messages, list) or not messages
|
||||
or not all(isinstance(message, dict) for message in messages)):
|
||||
raise PageIndexAPIError("messages must be a non-empty string or a "
|
||||
"list of message dicts.")
|
||||
block = _doc_block(client, doc_id)
|
||||
prepared = [dict(message) for message in messages]
|
||||
passthrough = {key: value for key, value in {
|
||||
"temperature": temperature, "top_p": top_p, "top_k": top_k,
|
||||
"stop_sequences": stop_sequences,
|
||||
}.items() if value is not None}
|
||||
runner = _anthropic_client().beta.messages.tool_runner(
|
||||
max_tokens=(max_tokens if max_tokens is not None
|
||||
else _default_max_tokens(model)),
|
||||
messages=prepared,
|
||||
model=model,
|
||||
tools=build_anthropic_tools(client, doc_ids=doc_id),
|
||||
system=_anthropic_system(system, block),
|
||||
stream=stream,
|
||||
# Bounded like the OpenAI surfaces (their framework default is 10).
|
||||
max_iterations=max_turns if max_turns is not None else 10,
|
||||
**passthrough,
|
||||
)
|
||||
|
||||
if stream:
|
||||
def events() -> Iterator[Any]:
|
||||
try:
|
||||
for turn_stream in runner:
|
||||
for event in turn_stream:
|
||||
yield event
|
||||
except anthropic.AnthropicError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
return events()
|
||||
|
||||
try:
|
||||
turns = [turn for turn in runner]
|
||||
except anthropic.AnthropicError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"The model backend failed: {exc}") from exc
|
||||
if not turns:
|
||||
raise PageIndexAPIError("The model returned no response.")
|
||||
captured: dict = {}
|
||||
|
||||
def capture(params):
|
||||
captured.update(params)
|
||||
return params
|
||||
|
||||
runner.set_messages_params(capture)
|
||||
if not captured.get("messages"):
|
||||
# The conversation is read back through a mutator; if a vendor
|
||||
# change stops it delivering params, the envelope would silently
|
||||
# lose the tool turns — fail loudly instead.
|
||||
raise PageIndexAPIError(
|
||||
"Could not read the conversation back from the anthropic tool "
|
||||
"runner — the installed anthropic version is incompatible with "
|
||||
"this pageindex release."
|
||||
)
|
||||
conversation = list(captured["messages"])
|
||||
final = turns[-1]
|
||||
envelope = final.model_dump(mode="json")
|
||||
envelope["content"] = [_dump_block(item) for item in final.content]
|
||||
envelope["usage"] = _anthropic_usage(turns, envelope.get("usage") or {})
|
||||
# The full turn sequence (assistant tool_use + user tool_result + final),
|
||||
# valid for verbatim append to the caller's history. The runner appends
|
||||
# a turn to its params only when it executed tools from it — content
|
||||
# carried tool_use blocks and the turn was not a refusal. stop_reason
|
||||
# alone cannot tell: a max_tokens turn with complete tool_use blocks
|
||||
# still executes. Whether final's tool_use ids already sit in the
|
||||
# history is the ground truth for "already appended".
|
||||
new_messages = [_dump_message(message)
|
||||
for message in conversation[len(prepared):]]
|
||||
final_blocks = [_dump_block(item) for item in final.content]
|
||||
final_ids = {block["id"] for block in final_blocks
|
||||
if block.get("type") == "tool_use"}
|
||||
history_ids = {block.get("id")
|
||||
for message in new_messages
|
||||
if (message.get("role") == "assistant"
|
||||
and isinstance(message.get("content"), list))
|
||||
for block in message["content"]
|
||||
if (isinstance(block, dict)
|
||||
and block.get("type") == "tool_use")}
|
||||
if not final_ids or not final_ids <= history_ids:
|
||||
# Unexecuted tool_use blocks (refusal turns) have no tool_result,
|
||||
# so they cannot enter an appendable history — strip them, as the
|
||||
# SDK itself does when it rebuilds params around such a turn.
|
||||
appendable = [block for block in final_blocks
|
||||
if block.get("type") != "tool_use"]
|
||||
if appendable:
|
||||
new_messages = new_messages + [
|
||||
{"role": "assistant", "content": appendable}]
|
||||
envelope["messages"] = new_messages
|
||||
return envelope
|
||||
@@ -0,0 +1,210 @@
|
||||
"""Minimal MCP client (streamable HTTP) for the PageIndex cloud MCP server.
|
||||
|
||||
Backs the cloud branches of ``client.agent_tools()`` and
|
||||
``client.agent_instructions()``: ``tools/list`` discovers the live tool set,
|
||||
``tools/call`` executes a tool, and the ``initialize`` handshake carries the
|
||||
server's agent instructions. Synchronous, requests-only.
|
||||
Works against both stateful and stateless servers: a session id returned by
|
||||
``initialize`` is echoed back, and a session-carrying request rejected with
|
||||
HTTP 404 (the spec's expired-session status) re-initializes once and
|
||||
retries; a 400 is an ordinary bad request and is never replayed.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import threading
|
||||
from typing import Any, Optional
|
||||
|
||||
import requests
|
||||
|
||||
from ._version import sdk_version
|
||||
from .errors import PageIndexAPIError
|
||||
|
||||
_PROTOCOL_VERSION = "2025-06-18"
|
||||
_TIMEOUT = (10, 240) # tools may wait server-side (wait_for_completion: 3 min)
|
||||
|
||||
|
||||
def _parse_sse(text: str) -> list[dict]:
|
||||
"""JSON-RPC messages out of a text/event-stream body."""
|
||||
messages = []
|
||||
text = text.replace("\r\n", "\n").replace("\r", "\n")
|
||||
for block in text.split("\n\n"):
|
||||
data_lines = [line[5:].removeprefix(" ") for line in block.splitlines()
|
||||
if line.startswith("data:")]
|
||||
if not data_lines:
|
||||
continue
|
||||
try:
|
||||
messages.append(json.loads("\n".join(data_lines)))
|
||||
except ValueError:
|
||||
continue
|
||||
return messages
|
||||
|
||||
|
||||
class McpBridge:
|
||||
def __init__(self, url: str, headers: dict[str, str]):
|
||||
self._url = url
|
||||
self._auth_headers = dict(headers)
|
||||
self._session_id: Optional[str] = None
|
||||
self._protocol_version: Optional[str] = None
|
||||
self._instructions: Optional[str] = None
|
||||
self._initialized = False
|
||||
self._lock = threading.RLock()
|
||||
self._next_id = 0
|
||||
|
||||
# ── JSON-RPC over streamable HTTP ──
|
||||
|
||||
def _post(self, payload: dict, session_id: Optional[str] = None,
|
||||
protocol_version: Optional[str] = None) -> requests.Response:
|
||||
headers = {
|
||||
"Content-Type": "application/json",
|
||||
"Accept": "application/json, text/event-stream",
|
||||
**self._auth_headers,
|
||||
}
|
||||
if session_id:
|
||||
headers["Mcp-Session-Id"] = session_id
|
||||
if protocol_version:
|
||||
headers["MCP-Protocol-Version"] = protocol_version
|
||||
try:
|
||||
return requests.post(self._url, json=payload, headers=headers,
|
||||
timeout=_TIMEOUT)
|
||||
except requests.RequestException as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"Could not reach the PageIndex MCP server: {exc}"
|
||||
) from exc
|
||||
|
||||
def _extract_result(self, response: requests.Response, request_id: int) -> Any:
|
||||
content_type = response.headers.get("Content-Type", "")
|
||||
if "text/event-stream" in content_type:
|
||||
# SSE is UTF-8 by spec; requests guesses latin-1 for charset-less
|
||||
# text/* and would mojibake every non-ASCII character.
|
||||
messages = _parse_sse(response.content.decode("utf-8",
|
||||
errors="replace"))
|
||||
else:
|
||||
try:
|
||||
messages = [response.json()]
|
||||
except ValueError as exc:
|
||||
raise PageIndexAPIError(
|
||||
f"MCP server returned a non-JSON response "
|
||||
f"(HTTP {response.status_code})."
|
||||
) from exc
|
||||
# Strict id correlation only — accepting any result-bearing message
|
||||
# would return a stale or mis-correlated reply as this call's.
|
||||
reply = next((m for m in messages if m.get("id") == request_id), None)
|
||||
if reply is None:
|
||||
raise PageIndexAPIError(
|
||||
"MCP server response contained no reply matching the request."
|
||||
)
|
||||
if "error" in reply:
|
||||
error = reply["error"] or {}
|
||||
raise PageIndexAPIError(
|
||||
f"MCP error {error.get('code')}: {error.get('message')}"
|
||||
)
|
||||
return reply.get("result")
|
||||
|
||||
def _request(self, method: str, params: Optional[dict] = None,
|
||||
_retry: bool = True) -> Any:
|
||||
self._ensure_initialized()
|
||||
with self._lock:
|
||||
self._next_id += 1
|
||||
request_id = self._next_id
|
||||
session_id = self._session_id
|
||||
protocol_version = self._protocol_version
|
||||
payload: dict[str, Any] = {"jsonrpc": "2.0", "id": request_id,
|
||||
"method": method}
|
||||
if params is not None:
|
||||
payload["params"] = params
|
||||
response = self._post(payload, session_id, protocol_version)
|
||||
if response.status_code == 404 and session_id and _retry:
|
||||
# Session expired (stateful servers; the spec's 404): the server
|
||||
# refused the request at session validation, so replaying it is
|
||||
# safe. 400 is an ordinary bad request — replaying one would
|
||||
# re-run side effects. Reset only if no other thread has already
|
||||
# re-initialized, then retry once on the fresh session.
|
||||
with self._lock:
|
||||
if self._session_id == session_id:
|
||||
self._initialized = False
|
||||
self._session_id = None
|
||||
self._protocol_version = None
|
||||
return self._request(method, params, _retry=False)
|
||||
if response.status_code >= 400:
|
||||
raise PageIndexAPIError(
|
||||
f"MCP request failed: HTTP {response.status_code} "
|
||||
f"({response.text[:200]})"
|
||||
)
|
||||
return self._extract_result(response, request_id)
|
||||
|
||||
def _ensure_initialized(self) -> None:
|
||||
with self._lock:
|
||||
if self._initialized:
|
||||
return
|
||||
self._next_id += 1
|
||||
request_id = self._next_id
|
||||
response = self._post({
|
||||
"jsonrpc": "2.0", "id": request_id, "method": "initialize",
|
||||
"params": {
|
||||
"protocolVersion": _PROTOCOL_VERSION,
|
||||
"capabilities": {},
|
||||
"clientInfo": {"name": "pageindex-python-sdk",
|
||||
"version": sdk_version()},
|
||||
},
|
||||
})
|
||||
if response.status_code >= 400:
|
||||
raise PageIndexAPIError(
|
||||
f"Could not connect to the PageIndex MCP server: HTTP "
|
||||
f"{response.status_code} ({response.text[:200]}). Check "
|
||||
"your API key."
|
||||
)
|
||||
result = self._extract_result(response, request_id) or {}
|
||||
self._session_id = response.headers.get("Mcp-Session-Id")
|
||||
self._protocol_version = result.get("protocolVersion",
|
||||
_PROTOCOL_VERSION)
|
||||
self._instructions = result.get("instructions")
|
||||
self._initialized = True
|
||||
# Sent inside the lock so no concurrent thread can slip a
|
||||
# request between the handshake and this notification.
|
||||
try:
|
||||
self._post({"jsonrpc": "2.0",
|
||||
"method": "notifications/initialized"},
|
||||
self._session_id, self._protocol_version)
|
||||
except PageIndexAPIError:
|
||||
pass # advisory; a server that required it fails the next request
|
||||
|
||||
# ── public surface ──
|
||||
|
||||
def instructions(self) -> Optional[str]:
|
||||
"""The server's agent instructions from the initialize handshake."""
|
||||
self._ensure_initialized()
|
||||
return self._instructions
|
||||
|
||||
def list_tools(self) -> list[dict]:
|
||||
tools: list[dict] = []
|
||||
cursor: Optional[str] = None
|
||||
while True:
|
||||
params = {"cursor": cursor} if cursor else {}
|
||||
result = self._request("tools/list", params) or {}
|
||||
tools.extend(result.get("tools") or [])
|
||||
cursor = result.get("nextCursor")
|
||||
if not cursor:
|
||||
return tools
|
||||
|
||||
def call_tool(self, name: str, arguments: dict[str, Any]) -> "tuple[str, bool]":
|
||||
"""Returns (text, is_error) — is_error is the server's MCP isError
|
||||
marking, which callers must carry to their framework's own error
|
||||
channel."""
|
||||
result = self._request("tools/call",
|
||||
{"name": name, "arguments": arguments}) or {}
|
||||
is_error = bool(result.get("isError"))
|
||||
texts = []
|
||||
for block in result.get("content") or []:
|
||||
if isinstance(block, dict) and block.get("type") == "text":
|
||||
texts.append(block.get("text", ""))
|
||||
elif isinstance(block, dict) and isinstance(block.get("data"), str):
|
||||
# Base64 payloads (image/audio) become a metadata stub —
|
||||
# dumped verbatim they hand the model the raw blob. Revisit
|
||||
# if tool results ever pass through as real multimodal input.
|
||||
kind = block.get("mimeType") or block.get("type") or "binary"
|
||||
size_kb = max(1, len(block["data"]) * 3 // 4096)
|
||||
texts.append(f"[{kind} content omitted: ~{size_kb} KB]")
|
||||
else:
|
||||
texts.append(json.dumps(block, ensure_ascii=False))
|
||||
return "\n".join(texts), is_error
|
||||
+12
-1
@@ -1,6 +1,6 @@
|
||||
[tool.poetry]
|
||||
name = "pageindex"
|
||||
version = "0.2.9"
|
||||
version = "0.2.10"
|
||||
description = "Python SDK for PageIndex — reasoning-based, vectorless document retrieval, cloud and local"
|
||||
readme = "README.md"
|
||||
license = "MIT"
|
||||
@@ -38,6 +38,17 @@ sortedcontainers = ">=2.4.0"
|
||||
regex = ">=2024.0.0"
|
||||
python-dotenv = ">=1.0.0"
|
||||
pyyaml = ">=6.0"
|
||||
# Older releases break string prompts with SDK MCP servers (#597, #780).
|
||||
claude-agent-sdk = { version = ">=0.1.53", optional = true }
|
||||
# Older releases crash on current openai before the request is sent.
|
||||
openai-agents = { version = ">=0.18.1", optional = true }
|
||||
# Older releases execute a refusal turn's tool_use blocks.
|
||||
anthropic = { version = ">=0.108.0", optional = true }
|
||||
|
||||
[tool.poetry.extras]
|
||||
claude = ["claude-agent-sdk"]
|
||||
openai = ["openai-agents"]
|
||||
anthropic = ["anthropic"]
|
||||
|
||||
[tool.poetry.group.dev.dependencies]
|
||||
pytest = ">=7.0"
|
||||
|
||||
@@ -0,0 +1,215 @@
|
||||
{
|
||||
"_provenance": "Frozen copy of the PageIndex cloud MCP server's tool contract (names, input schemas, descriptions, and annotations as served via tools/list). The parity test asserts pageindex.agent_tools.TOOL_CONTRACT matches this file; update both together only when the cloud contract changes.",
|
||||
"tools": {
|
||||
"browse_documents": {
|
||||
"annotations": {
|
||||
"readOnlyHint": true,
|
||||
"openWorldHint": false
|
||||
},
|
||||
"description": "Primary document retrieval tool. After orienting with get_folder_structure() (when available), use this for all document-related questions. The bare call returns root-level sub-folders and documents; pass folder_id to drill into a sub-folder level by level. Use sort=\"relevance\" + query for semantic ranking. Do NOT jump to search_documents() first — it is an escalation path, only after browse_documents(sort=\"relevance\") has failed.",
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"folder_id": {
|
||||
"type": "string",
|
||||
"default": "root",
|
||||
"description": "Folder scope (default \"root\"). Pass a specific folder ID to scope into that folder, or \"root\" to reference the library root. The read-only \"shared-with-me\" and \"following\" folders live at the library root — pass one of those ids to browse them. Copy any folder_id verbatim from a browse/tree response, never construct one. Combine with `recursive` to control breadth."
|
||||
},
|
||||
"recursive": {
|
||||
"type": "boolean",
|
||||
"default": false,
|
||||
"description": "Whether to include documents from descendant folders. When false (default), returns the direct contents of folder_id along with its sub-folders — prefer this for level-by-level exploration so you retain folder hierarchy context. When true, flattens all descendant documents into one list and omits sub-folders — use only when a non-recursive browse of the target folder returned no relevant results and you need to widen the scope, or the user explicitly requests a flat listing."
|
||||
},
|
||||
"sort": {
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"time",
|
||||
"relevance"
|
||||
],
|
||||
"default": "time",
|
||||
"description": "Sort order. \"time\" (default) sorts by upload date (newest first); \"relevance\" orders documents by semantic relevance to `query`. Relevance also works inside the read-only shared folders — pass their folder_id — but at the library root it ranks only your own documents."
|
||||
},
|
||||
"query": {
|
||||
"type": "string",
|
||||
"description": "Search query for relevance ranking. Required when sort=\"relevance\"; must be omitted when sort=\"time\"."
|
||||
},
|
||||
"offset": {
|
||||
"type": "integer",
|
||||
"minimum": 0,
|
||||
"maximum": 9007199254740991,
|
||||
"default": 0,
|
||||
"description": "Zero-based pagination offset. Pass the value of `next_offset` from the previous response to fetch the next page."
|
||||
},
|
||||
"limit": {
|
||||
"type": "number",
|
||||
"minimum": 1,
|
||||
"maximum": 50,
|
||||
"default": 10,
|
||||
"description": "Number of documents to return per page (1-50, default 10)"
|
||||
}
|
||||
},
|
||||
"required": []
|
||||
}
|
||||
},
|
||||
"get_document": {
|
||||
"annotations": {
|
||||
"readOnlyHint": true,
|
||||
"openWorldHint": false
|
||||
},
|
||||
"description": "Check a document's processing status and metadata. `status` is one of \"pending\", \"queued\", \"processing\", \"completed\", or \"failed\" — call this before `get_document_structure()` or `get_page_content()` to confirm the document is ready.",
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"doc_name": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
|
||||
},
|
||||
"folder_id": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
],
|
||||
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
|
||||
},
|
||||
"wait_for_completion": {
|
||||
"type": "boolean",
|
||||
"default": false,
|
||||
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
|
||||
}
|
||||
},
|
||||
"required": [
|
||||
"doc_name"
|
||||
]
|
||||
}
|
||||
},
|
||||
"get_document_structure": {
|
||||
"annotations": {
|
||||
"readOnlyHint": true,
|
||||
"openWorldHint": false
|
||||
},
|
||||
"description": "Extract a document's hierarchical outline (headers, sections, page references). REQUIRED for documents over 20 pages — call this first to locate relevant sections, then pass their page numbers to `get_page_content()`. Use the `part` parameter to iterate large outlines until `pagination.has_more` is false.",
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"doc_name": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
|
||||
},
|
||||
"folder_id": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
],
|
||||
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
|
||||
},
|
||||
"part": {
|
||||
"type": "integer",
|
||||
"minimum": 1,
|
||||
"maximum": 9007199254740991,
|
||||
"default": 1,
|
||||
"description": "Part number for pagination (1-based, default 1). For large outlines, increment until the response's `pagination.has_more` becomes false."
|
||||
},
|
||||
"wait_for_completion": {
|
||||
"type": "boolean",
|
||||
"default": false,
|
||||
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
|
||||
}
|
||||
},
|
||||
"required": [
|
||||
"doc_name"
|
||||
]
|
||||
}
|
||||
},
|
||||
"get_page_content": {
|
||||
"annotations": {
|
||||
"readOnlyHint": true,
|
||||
"openWorldHint": false
|
||||
},
|
||||
"description": "Extract page content from a processed document. Use tight, targeted page ranges — never the whole document at once. For documents over 20 pages, call `get_document_structure()` first to pick relevant sections. Embedded image paths in the response feed into `get_document_image()`.",
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"doc_name": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
|
||||
},
|
||||
"folder_id": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
],
|
||||
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
|
||||
},
|
||||
"pages": {
|
||||
"type": "string",
|
||||
"minLength": 1,
|
||||
"pattern": "^(\\d+(-\\d+)?)(,\\s*\\d+(-\\d+)?)*$",
|
||||
"description": "Page specification: \"5\", \"3,7,10\", \"5-10\", or \"1-3,7,9-12\""
|
||||
},
|
||||
"wait_for_completion": {
|
||||
"type": "boolean",
|
||||
"default": false,
|
||||
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
|
||||
}
|
||||
},
|
||||
"required": [
|
||||
"doc_name",
|
||||
"pages"
|
||||
]
|
||||
}
|
||||
},
|
||||
"remove_document": {
|
||||
"annotations": {
|
||||
"readOnlyHint": false,
|
||||
"destructiveHint": true,
|
||||
"idempotentHint": true,
|
||||
"openWorldHint": false
|
||||
},
|
||||
"description": "Permanently delete documents and all associated data. Only invoke when the user explicitly names the documents AND confirms deletion. Returns `results` — one entry per requested document: `{ doc_name, status: \"deleted\" | \"not_found\" | \"failed\", error? }`. Inspect each entry for per-document failures. This action is irreversible.",
|
||||
"schema": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"doc_names": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "string",
|
||||
"minLength": 1
|
||||
},
|
||||
"minItems": 1,
|
||||
"maxItems": 10,
|
||||
"description": "Array of document names to delete. Each name must be copied verbatim from the `name` field of a browse_documents() or search_documents() response (case-sensitive, include extension). Example: [\"Q3 Report.pdf\", \"draft.pdf\"]. Max 10 per call."
|
||||
},
|
||||
"folder_id": {
|
||||
"anyOf": [
|
||||
{
|
||||
"type": "string"
|
||||
},
|
||||
{
|
||||
"type": "null"
|
||||
}
|
||||
],
|
||||
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
|
||||
}
|
||||
},
|
||||
"required": [
|
||||
"doc_names"
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
+79
-3
@@ -151,6 +151,16 @@ def test_get_page_content(local_client, indexed_doc):
|
||||
local_client.get_page_content(indexed_doc, "abc")
|
||||
|
||||
|
||||
def test_get_page_content_span_bomb_rejected(local_client, indexed_doc):
|
||||
"""An absurd range must be rejected arithmetically, not expanded into
|
||||
a billion integers in the caller's process (the tool layer already
|
||||
refused; the public client method did not)."""
|
||||
with pytest.raises(ValueError, match="spans more than 10000"):
|
||||
local_client.get_page_content(indexed_doc, "1-1000001")
|
||||
# At the bound itself the spec still parses.
|
||||
assert local_client.get_page_content(indexed_doc, "5-10004") == []
|
||||
|
||||
|
||||
def test_submit_does_not_create_cwd_logs(local_client, sample_pdf, tmp_path, monkeypatch):
|
||||
monkeypatch.chdir(tmp_path)
|
||||
def fake_page_index_main(doc, opt=None, logger=None, page_list=None):
|
||||
@@ -162,6 +172,49 @@ def test_submit_does_not_create_cwd_logs(local_client, sample_pdf, tmp_path, mon
|
||||
assert not (tmp_path / "logs").exists()
|
||||
|
||||
|
||||
def test_submit_duplicate_name_gets_suffix(local_client, sample_pdf, monkeypatch):
|
||||
"""Mirror the cloud upload: a second submit of the same file name is
|
||||
stored as name_1, not as a same-name duplicate."""
|
||||
def fake_page_index_main(doc, opt=None, logger=None, page_list=None):
|
||||
return {"doc_name": "sample.pdf", "doc_description": "d",
|
||||
"structure": json.loads(json.dumps(STRUCTURE))}
|
||||
monkeypatch.setattr(page_index_module, "page_index_main", fake_page_index_main)
|
||||
first = local_client.submit_document(sample_pdf)
|
||||
assert first["name"] == "sample.pdf"
|
||||
with pytest.warns(UserWarning, match='stored as "sample_1.pdf"'):
|
||||
second = local_client.submit_document(sample_pdf)
|
||||
assert second["name"] == "sample_1.pdf"
|
||||
names = {d["id"]: d["name"]
|
||||
for d in local_client.list_documents()["documents"]}
|
||||
assert names[first["doc_id"]] == "sample.pdf"
|
||||
assert names[second["doc_id"]] == "sample_1.pdf"
|
||||
|
||||
|
||||
def test_submit_duplicate_name_exhaustion(local_client, monkeypatch):
|
||||
api = local_client._api
|
||||
metas = ([{"name": "x.pdf"}]
|
||||
+ [{"name": f"x_{num}.pdf"} for num in range(1, 100)])
|
||||
monkeypatch.setattr(api._store, "list_metas", lambda: metas)
|
||||
with pytest.raises(PageIndexAPIError, match="Too many files"):
|
||||
api._unique_doc_name("x.pdf")
|
||||
|
||||
|
||||
def test_submit_name_exhaustion_rejects_before_indexing(
|
||||
local_client, sample_pdf, monkeypatch,
|
||||
):
|
||||
api = local_client._api
|
||||
metas = ([{"name": "sample.pdf"}]
|
||||
+ [{"name": f"sample_{num}.pdf"} for num in range(1, 100)])
|
||||
monkeypatch.setattr(api._store, "list_metas", lambda: metas)
|
||||
monkeypatch.setattr(
|
||||
page_index_module, "page_index_main",
|
||||
lambda *args, **kwargs: pytest.fail(
|
||||
"indexer ran despite name exhaustion"),
|
||||
)
|
||||
with pytest.raises(PageIndexAPIError, match="Too many files"):
|
||||
local_client.submit_document(sample_pdf)
|
||||
|
||||
|
||||
def test_submit_flash(local_client, sample_pdf, monkeypatch):
|
||||
calls = {}
|
||||
def fake_flash(pdf, summary=True, summary_model=None, **kwargs):
|
||||
@@ -432,7 +485,8 @@ def test_torn_delete_never_lists_ghost(local_client, indexed_doc, tmp_path):
|
||||
|
||||
|
||||
def test_corrupt_doc_json_is_contained(local_client, indexed_doc, sample_pdf, tmp_path):
|
||||
second = local_client.submit_document(sample_pdf)["doc_id"]
|
||||
with pytest.warns(UserWarning): # same-name resubmit → stored as sample_1.pdf
|
||||
second = local_client.submit_document(sample_pdf)["doc_id"]
|
||||
(tmp_path / "store" / "docs" / indexed_doc / "doc.json").write_text("{truncated")
|
||||
|
||||
# manifest still holds a good copy of the meta — served consistently
|
||||
@@ -599,8 +653,12 @@ def test_retrieval_endpoints_cloud_only(local_client):
|
||||
local_client.get_retrieval("any")
|
||||
|
||||
|
||||
def test_chat_completions_cloud_only(local_client):
|
||||
with pytest.raises(PageIndexAPIError, match="not yet supported in local mode"):
|
||||
def test_chat_completions_local_needs_agents_extra(local_client, monkeypatch):
|
||||
"""Local chat is implemented (see test_local_chat.py); without the
|
||||
openai-agents extra it raises the actionable install error."""
|
||||
import sys
|
||||
monkeypatch.setitem(sys.modules, "agents", None)
|
||||
with pytest.raises(PageIndexAPIError, match="pageindex\\[openai\\]"):
|
||||
local_client.chat_completions(
|
||||
messages=[{"role": "user", "content": "q"}])
|
||||
|
||||
@@ -710,3 +768,21 @@ def test_cloud_chat_stream_parsing(cloud, monkeypatch):
|
||||
messages=[{"role": "user", "content": "q"}], stream=True,
|
||||
stream_metadata=True))
|
||||
assert {"object": "chat.completion.citations", "citations": []} in chunks
|
||||
|
||||
|
||||
def test_cloud_chat_accepts_query_string(cloud):
|
||||
client, calls, fake = cloud
|
||||
fake.payload = {"choices": [{"message": {"content": "ok"}}]}
|
||||
client.chat_completions("What status?")
|
||||
assert calls[-1]["json"]["messages"] == [
|
||||
{"role": "user", "content": "What status?"}]
|
||||
with pytest.raises(PageIndexAPIError, match="non-empty string"):
|
||||
client.chat_completions(" ")
|
||||
|
||||
|
||||
def test_parse_pages_overlap_counts_union():
|
||||
from pageindex.client import _parse_pages
|
||||
pages = _parse_pages("1-5000,2000-9000")
|
||||
assert len(pages) == 9000 and pages[0] == 1 and pages[-1] == 9000
|
||||
with pytest.raises(ValueError, match="spans more than"):
|
||||
_parse_pages("1-10001")
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -60,3 +60,44 @@ def test_import_pageindex_is_lazy():
|
||||
out = subprocess.run([sys.executable, "-c", probe],
|
||||
capture_output=True, text=True, check=True)
|
||||
assert out.stdout.split() == ["clean", "function"]
|
||||
|
||||
|
||||
def test_sdk_submodules_reachable_and_dunder_probes_stay_lazy():
|
||||
"""The 0.2.10 modules resolve as attributes, and underscore probes (the
|
||||
frequent unknown names: copy/pickle/inspect dunders) raise without
|
||||
dragging in the indexing stack. A non-underscore unknown name still
|
||||
raises AttributeError — after the compat fallthrough's one classic
|
||||
import, which is the pre-0.2.10 behavior."""
|
||||
probe = (
|
||||
"import sys, pageindex\n"
|
||||
"pageindex.agent_tools; pageindex.local_chat\n"
|
||||
"pageindex.mcp_bridge; pageindex.integrations\n"
|
||||
"assert not hasattr(pageindex, '__wrapped__')\n"
|
||||
"heavy = [m for m in ('pageindex.page_index_classic', "
|
||||
"'pageindex.flash', 'pageindex.utils') if m in sys.modules]\n"
|
||||
"print(','.join(heavy) or 'clean')\n"
|
||||
"try:\n"
|
||||
" pageindex.definitely_missing\n"
|
||||
" raise SystemExit('no AttributeError')\n"
|
||||
"except AttributeError:\n"
|
||||
" pass\n"
|
||||
)
|
||||
out = subprocess.run([sys.executable, "-c", probe],
|
||||
capture_output=True, text=True, check=True)
|
||||
assert out.stdout.strip() == "clean"
|
||||
|
||||
|
||||
def test_classic_compat_surface_still_reachable():
|
||||
"""The pre-0.2.10 catch-all made every classic/utils public name a
|
||||
package attribute; dropping it broke `from pageindex import
|
||||
ConfigLoader` on upgrade with no deprecation path."""
|
||||
probe = (
|
||||
"import pageindex\n"
|
||||
"assert callable(pageindex.count_tokens)\n"
|
||||
"assert isinstance(pageindex.ConfigLoader, type)\n"
|
||||
"from pageindex import check_toc # noqa: F401\n"
|
||||
"print('ok')\n"
|
||||
)
|
||||
out = subprocess.run([sys.executable, "-c", probe],
|
||||
capture_output=True, text=True, check=True)
|
||||
assert out.stdout.strip() == "ok"
|
||||
|
||||
Reference in New Issue
Block a user