feat: agent tools and local chat for the PageIndex SDK (v0.2.10) (#396)

* feat: agent tools — the cloud MCP tool contract on the client

Four new client methods make PageIndex documents available to agent
frameworks, in both modes, with the mode decided solely by the client
constructor:

- agent_tools(): plain functions (browse_documents, get_document,
  get_document_structure, get_page_content) matching the PageIndex cloud
  MCP server's tools/list — same names, schemas, descriptions, and JSON
  response envelopes — so agent prompts port unchanged between the cloud
  MCP connection and these in-process tools. Tools never raise; errors
  come back in the same envelope. remove_document ships behind
  include_management=False.
- as_openai_tools(): the same tools wrapped for the OpenAI Agents SDK.
- as_claude_mcp(): one mcp_servers entry for the Claude Agent SDK —
  cloud clients get the remote MCP config (the framework connects to
  api.pageindex.ai/mcp and discovers the full cloud tool set), local
  clients get an in-process SDK MCP server.
- agent_instructions(doc_id=None): orchestration guidance for the
  agent's system prompt; doc_id (same shape as chat_completions) appends
  the target documents.

submit_document() gains wait=True: poll get_document status until
completed, raise on failed or after 30 minutes — the manual polling loop
every cloud caller writes today spins forever on a failed document.

Neither framework becomes a dependency: imports happen at call time with
actionable errors, and pageindex[openai] / pageindex[claude] extras are
floor-only pins. tests/data/cloud_mcp_contract.json freezes the tool
contract; a parity test guards against drift. 36 new tests (95 total),
plus a live OpenAI Agents SDK run over a seeded local store verifying
the structure-first navigation flow end to end.

* fix: agent tools review — next_steps order, resolve caching, error semantics

- Large-doc next_steps now says structure-first, consistent with tool
  descriptions and agent instructions
- _remove_document fetches document list once instead of per-name
- call_tool returns error envelope for unknown names instead of raising
- _not_ready_error timed_out flag reflects actual wait outcome
- openai_agents.py docstring corrected to match default (FunctionTools)
- Removed unused ModelSettings import from demo

* fix: agent tools review 2 — bridge thread safety, browse paging, metadata merge

- McpBridge reads session/protocol headers under the lock (now RLock:
  _ensure_initialized posts while holding it). openai-agents runs sync
  tools on threads and executes parallel tool calls concurrently, so
  bridge functions genuinely race; a torn read sent a new session id
  with a stale protocol header. Measured: one session expiry under 8
  threads cost 4 initializations before, minimal 2 after.
- Session-expiry retry also resets the negotiated protocol version, so
  the re-handshake carries no stale MCP-Protocol-Version header.
- browse_documents time sort pages list_documents natively instead of
  fetching the whole library to slice one window (relevance still needs
  the full list for scoring).
- _await_completion: a status refetch that nulls out metadata no longer
  clobbers the listing's copy (setdefault was a no-op on existing None).
- Structure tool reads the raw stored tree via a named LocalAPI
  raw_tree() seam instead of reaching into _api._store internals; drop
  the redundant deepcopy before _format_structure (store re-reads from
  disk, formatting builds fresh containers).
- Shared pageindex/_version.py replaces _sdk_version duplicated in
  mcp_bridge and the Claude integration.

Left as-is after source verification against the cloud MCP: first-page
budget bypass, pageNum falsy-zero, and the page-gap fallback text are
letter-for-letter cloud behavior — parity wins over local repair.

* fix: agent tools review 3 — page-span cap, duplicate names, wait resilience, contract drift

- _parse_page_spec bounds the requested span arithmetically (10k pages)
  before materializing it; pages="1-1000000000" previously expanded to a
  billion integers inside the caller's process.
- Local submit_document uniquifies document names the way the cloud
  upload does (taken name -> _1.._99, then reject with the cloud's own
  message). Same-name duplicates broke name-addressed tools: resolution
  always picks the newest, so older duplicates were unreachable.
- agent_instructions(doc_id=...) now fails loud when the pinned doc's
  name is shadowed by a newer same-name document (legacy stores predate
  the rename) — it previews resolution with the same _resolve_document
  the tools use, so the check cannot drift from actual behavior.
- submit_document(wait=True) tolerates transient network errors, not
  just API errors; a dropped connection at minute 25 of a 30-minute
  wait no longer kills it. Third strike wraps into PageIndexAPIError
  per the documented contract.
- The live contract-parity test compares full per-param schemas, not
  just names and descriptions. It immediately caught real drift the
  shallow check had been passing: the server now emits nullables as
  anyOf unions and stamps MAX_SAFE_INTEGER maxima on offset/part.
  Contract and snapshot updated to the served wire form; _annotation_for
  learned anyOf so bridge signatures stay Optional[str] instead of
  degrading to Any.

Adjudicated, not changed: the allowed_tools wildcard example stays
(docstring advice covers scoping; Ray's call), and raw-length response
accounting stays (letter-for-letter cloud behavior, parity wins).

* feat: surface the stored document name from submit_document

Compute PR #558 makes /doc/ return {"doc_id", "name"} carrying the
post-dedup-rename name. Mirror it end to end: local submit returns the
stored name, the client warns when it differs from the uploaded file
name (read via .get so older cloud servers stay compatible), the local
name-exhaustion check runs before indexing instead of after the LLM
spend, and the demo caches doc_id in a file instead of name-matching —
a renamed document made the name lookup re-index on every run.

* fix: add missing page_list kwarg in duplicate-name test mock

* revert: keep README.md unchanged from main — SDK section deferred

* feat: serve cloud agent instructions live from the MCP server

The cloud MCP server publishes its agent instructions in the initialize
result, adapted to each key's tool set. agent_instructions() previously
returned the SDK's local-subset text in both modes — a silently forked
copy that lacks the guidance for cloud-only tools (search_documents
escalation, folders, images) and drifts as the server's prompt evolves.

Cloud clients now serve the server's live instructions, captured from
the initialize handshake on a per-client bridge shared with
agent_tools() (one session, no extra request). An empty server response
raises instead of silently substituting the subset text — same posture
as the annotation-regression guard. The local constant stays as the
honest subset for the in-process tools, with its provenance noted and a
consistency test that every tool it names exists in the local registry.

* fix: local relevance sort answers honestly instead of imitating

sort="relevance" is cloud-side semantic ranking; the local substring
imitation could satisfy the letter of the interface while silently
missing semantically relevant documents. Per the honest-subset rule
(same treatment as folders), local now returns the "not available
here" envelope for sort="relevance" or a stray query, and the local
instructions steer discovery through name/description matching plus
full-library paging instead of prescribing a capability that does not
exist here. The tool schema keeps the cloud contract verbatim, like
folder_id: honesty lives in the runtime answer, not a forked contract.

* docs: note the cloud+Claude instructions duplication trade-off in as_claude_mcp

* fix: unsupported-capability envelopes say local-mode-yet, point to cloud

"Not available here" read as a broken feature; the honest framing is
that folders and semantic ranking exist on PageIndex cloud and are not
in local mode yet. Both envelopes now say so and name the cloud client
in next_steps, so agents relay an accurate story to the user.

* fix: local tool descriptions pre-announce cloud-only capabilities

The cloud-verbatim browse_documents description invites
sort="relevance" and folder drilling, so a local agent's first semantic
search attempt was a guaranteed dead end discovered only from the
runtime error envelope. Local registration now appends a LOCAL MODE
note to the description — the agent learns what is cloud-only before
calling; the runtime envelope stays as the backstop for prompts that
ignore descriptions. The cloud-facing contract stays byte-verbatim.

* refactor: localized tool guidance replaces the appended LOCAL MODE note

Appending a retraction to the cloud-verbatim description left the model
parsing an instruction and its negation — and kept the cloud text
recommending search_documents and get_folder_structure, tools that are
not registered locally (get_page_content likewise pointed at
get_document_image). Guidance now adapts to the local surface the way
AGENT_INSTRUCTIONS already does: schema structure stays byte-identical
to the contract (mechanically asserted by a strip-descriptions test),
while local description strings teach only what works here and point to
PageIndex cloud for the rest. A dead-reference test forbids local
guidance from naming tools outside the local registry, so a contract
refresh that reintroduces a cloud-only reference fails loudly.

* feat: hide cloud-only parameters from the local tool surface

folder_id, sort, query, and recursive were exposed locally with
localized "cloud-only" descriptions, leaving the dead-end calls
expressible and discovered at runtime. Schema constraints beat
guidance: the local surface now serves the contract minus these
parameters, so strict-schema frameworks make the calls inexpressible
and a prompt that insists on sort="relevance" degrades to the bare
call (the correct local behavior) instead of an error round-trip.

The implementations still accept the hidden parameters and answer with
the guided "works on PageIndex cloud" envelope — the backstop for
direct call_tool callers and hosts without schema enforcement.
wait_for_completion stays: seeded or torn stores can hold documents
that are genuinely not completed. The structural guard now asserts the
local schema equals the contract minus the documented hidden set,
descriptions aside.

* fix: incremental-review findings — bridge cache, guards, envelope drift

Three independent review passes over the agent-instructions increment
surfaced six fixes:

- The per-client bridge moved off the instance into a weak-keyed,
  lock-guarded module cache: cloud clients stay picklable
  (threading.RLock no longer rides on the client) and concurrent first
  calls can no longer construct duplicate bridges/sessions.
- Blank or non-string initialize.instructions now hit the same honest
  error as a missing one — a whitespace-only or structured value could
  previously become the system prompt (or crash the doc_id append with
  a raw TypeError).
- The invalid-sort envelope no longer prescribes sort="relevance" — the
  one error text that still taught the cloud-only value it would then
  reject.
- "Page through the rest of the library" is emitted only when has_more
  is true; a fully-listed library no longer instructs a pointless call.
- The mandatory full-library paging step now says limit: 50 — 6 calls
  instead of 30 on a 300-document library.
- Docstrings and comments rescoped to what is actually true: the
  never-raise contract covers invocations the signatures accept
  (unknown params fail at the Python boundary; call_tool answers them
  with the guided envelope), recursive is accepted as the identity
  rather than errored, lenient framework arg models drop hidden params
  pre-call, and the module header no longer claims full schema parity.
  The capability-phrase guard now covers every local docstring, not
  just browse_documents.

* chore: keep the demo's doc_id cache file out of the repo

* test: live envelope field-parity guard against cloud response drift

The frozen contract guards tools/list, but the response envelopes the
local tools emit were hand-built to mirror the cloud's and had no drift
detector. A key-gated live test now asserts every field local emits
exists in the live cloud response for the analogous call (top-level
keys, next_steps, document entries, structure nodes, content entries).
Guidance wording is deliberately localized and not compared. Verified
green against the live server: local and cloud field structures
currently match exactly.

* feat: local chat — three protocol surfaces over the agent tools (v0.2.10)

Local mode gains managed document QA: an agent over the #393 local tool
set, reachable through three wire protocols, each 1:1 with the backend
and with no translation layer.

- chat_completions(): standard chat.completions semantics on any
  OpenAI-compatible backend (openai-agents engine). Final answer only,
  cross-turn aggregated usage, streaming as text pieces or chunk dicts
  (the existing cloud signature, now implemented locally; model and
  max_turns are local-only additions).
- responses(): the agentic surface — OpenAI Responses format, the tool
  process is standard output items, streaming forwards native events
  (tool outputs emitted as response.output_item.done, the way the
  platform streams its own server-side tools). Round-tripping output
  into the next input keeps provider prompt-cache prefix continuity and
  the agent's memory — live-verified: the follow-up call answered from
  round-tripped tool output with zero new tool calls.
- messages(): Anthropic-native via the SDK's own tool runner (new
  pageindex[anthropic] extra, floor 0.68.0 verified for
  tool_runner/beta_tool(input_schema)). tool_use/tool_result round-trip
  is the format's native behavior; the envelope is the final message
  with aggregated usage plus the full new-turn sequence; the managed
  system blocks carry cache_control breakpoints.

Shared skeleton: thin chat header + the local AGENT_INSTRUCTIONS
(caller system content is appended, not rejected), the doc_id targeting
block as a leading context item (factored out of
build_agent_instructions), read-only toolset, structural-only
validation (no arbitrary caps — backend limits govern), sampling params
passed through, per-run tracing disabled, enable_citations rejected as
cloud-only. Design basis is industry-standard formats rather than the
cloud chat endpoint; responses()/messages() raise on cloud clients
until the cloud converges.

Tests run the real engines against scripted backends (a Model fake for
openai-agents, a mock HTTP transport under the real anthropic SDK) with
real tool execution against a seeded store, including the round-trip
prefix-extension assertions on both engines.

* fix: local-chat review findings — truncation, serialization, streams

Three independent review passes (bug scan, claims-vs-code, adversarial
runtime probes) over the local-chat increment; every fix below was
reproduced before being fixed.

messages():
- A max_turns cut no longer duplicates the final assistant turn: the
  runner has already appended it when iterations exhaust, so the
  round-trip history carried a duplicate tool_use id and ended on an
  unanswered tool_use — a guaranteed 400 on continuation. The append
  now keys on stop_reason, and truncation reads natively as
  stop_reason: "tool_use" with a continuable history.
- The envelope is JSON-serializable end to end: runner-stored turns
  carry pydantic content blocks; everything is dumped to plain dicts,
  excluding SDK-internal __api_exclude__ fields (parsed_output) that
  the API rejects on round-trip.
- Bounded by default (max_iterations 10, like the OpenAI surfaces);
  usage aggregation now preserves the final turn's native fields and
  sums the token counters None-safely; empty caller system strings are
  skipped; non-dict message entries and bad doc_id types raise
  PageIndexAPIError; anthropic < 0.68 gets an actionable version error;
  the doc block no longer spends a cache_control breakpoint.

chat_completions()/responses():
- MaxTurnsExceeded wraps into PageIndexAPIError on all four run paths.
- responses(stream=True) is one logical response: per-turn backend
  lifecycle events are collapsed (a canonical consumer previously
  stopped at turn 1's response.completed and never saw the answer),
  sequence numbers are reassigned monotonically, and the synthesized
  tool-output event carries output_index/sequence_number.
- The responses envelope carries the real request surface
  (instructions, the actual function tool definitions, tool_choice,
  parallel_tool_calls, error/incomplete_details).
- RunConfig(group_id) pins a stable prompt_cache_key: openai-agents
  otherwise stamps each run with a fresh key, tagging round-tripped
  prefixes as different cache groups and defeating the feature the
  round-trip exists for.
- Abandoning a stream now cancels the run: a watchdog task lets the
  cancellation land even while the pump awaits the backend, and the
  per-call AsyncOpenAI client is closed before its loop ends (fixes
  "Task exception was never retrieved" noise). The opening role chunk
  is emitted even for empty outputs; empty responses() input and
  enable_citations-before-extra ordering fixed.

Docs rescoped to what is true: finish_reason/status reflect loop
completion on the OpenAI surfaces (the engine does not surface per-turn
backend reasons); chat streaming yields visible narration including
pre-tool text; messages(stream=True) forwards the Anthropic SDK's
native event objects (not wire-verbatim); the doc block is a leading
conversation item on OpenAI surfaces and a system block on messages().

Tests: 25 in the file (11 new), with per-extra skip sections so a
machine with only one framework still covers the other surface;
without-frameworks matrix re-verified; live smoke re-run green with a
clean exit.

* feat: as_anthropic_tools — Anthropic tool-runner export, both modes

Fills the last cell of the agent-connection matrix: users driving their
own anthropic tool_runner loop get runnable tools directly. Cloud wraps
the live MCP tool set with input schemas passing through verbatim (MCP
inputSchema is the Messages API schema shape); local exposes the same
set messages() runs internally. The beta_tool wrapping moves from
local_chat into integrations/anthropic_sdk.py, parallel to
openai_agents.py, and messages() now consumes the shared builder.
agent_tools grows _bridge_invoker/_read_only_tools so the plain-function
and beta_tool cloud paths share invocation containment and the
read-only gate.

* fix: as_anthropic_tools review findings — async flavor, schema isolation

Adversarial + best-practice review of 4590dd8 (three independent passes)
surfaced two holes. The export was sync-only: AsyncAnthropic's runner
accepts only BetaAsyncFunctionTool and splices anything else into the
request body unserialized, so the first call died with an opaque
TypeError — asynchronous=True now builds beta_async_tool runnables
(present since the 0.68.0 floor) that run the blocking bridge/store call
in a worker thread, keeping I/O off the caller's event loop. And
beta_tool stores input_schema by reference, so cloud tools aliased the
bridge's cached metas while the local path deep-copied — the builder now
copies, and the passthrough test asserts equal-but-not-aliased so it can
no longer compare an object with itself. Docstring fixes from the same
round: the MCP-connector pointer now carries the full live-verified
shape (authorization_token was missing — following it literally gave a
401), and the manual messages.create loop's to_dict() serialization is
documented. Tests pin the runnable flavor both ways (isinstance), which
existing tests could not distinguish.

* docs: doc_id is per-call table-setting — keep it identical across a conversation

The targeting block doc_id adds is re-set on every call and sits in the
cached prompt prefix, so a round-trip that drops (or changes) doc_id
silently diverges the prefix and loses the cache continuation. State the
rule on all three chat surfaces' doc_id docs, and pin it with a prefix
test that passes the same doc_id on both calls.

* feat: every chat surface takes a bare query string

query + doc_id is the minimal PageIndex contract, so it now works
uniformly: chat_completions and messages accept a plain string (one
user message), as responses always did per its wire format. The wrap
is input sugar at the SDK surface, not a translation layer — the
outgoing wire is unchanged, and managed agent surfaces taking strings
is the ecosystem convention (Runner.run, claude_agent_sdk.query).
Cloud chat_completions gains the same acceptance; blank strings raise
on every path.

* feat: messages() defaults max_tokens to 4096

The Messages API requires a per-turn output budget on the wire, but
that is table-setting, not a PageIndex-layer user obligation — the
simple call is now a question + model + doc_id. The knob stays
overridable (passthrough intact); model stays required because no
cross-vendor default is honest to guess.

* fix: raise messages() max_tokens default to 8192

max_tokens is a cap, not consumption, so the default should be the
highest universally safe value: 4096 could truncate long-form answers
(whole-document summaries), while 8192 is the output ceiling every
non-EOL Claude model accepts and stays under the SDK's non-streaming
long-request threshold.

* fix: restore per-extra skip markers the string-input tests displaced

Inserting tests above decorated ones absorbed their @needs_agents
markers, so two tests ran (and failed) in the without-frameworks CI
job. Both simulated-bare and full runs are green again.

* fix: close 17 findings from the v0.2.10 max review

Tool layer:
- anthropic adapter: failed tool calls raise ToolError so the runner
  emits tool_result is_error:true; McpBridge.call_tool returns
  (text, is_error) and surfaces the server's MCP isError marking
- as_openai_tools builds FunctionTool with the contract/server schema
  verbatim (strict off) — function_tool() regenerated schemas from
  signatures, dropping items/enum/pattern/bounds and aborting the whole
  list on object-typed params; shared _tool_specs() feeds both adapters
- remove_document validates every name before deleting anything; call_tool
  classifies only bind-time TypeErrors as INVALID_INPUT
- unknown-tool envelope formatted with _dumps like every other envelope

Local chat:
- doc_id is enforced at the tool layer (allowlist threaded through
  call_tool and the adapters), not just prompted; the shadow check runs
  inside the scope
- _openai_model routes litellm/ and provider/ paths via LitellmModel and
  strips openai/ — the normalized retrieve_model 404'd as a raw wire name
- responses() reports the backend's real terminal status (recorded at the
  transport client; the framework discards Response.status) and wraps
  framework exceptions in PageIndexAPIError
- chat_completions streaming yields its opening chunk inside try, so an
  abandoned iterator still cancels the run and closes the backend
- prompt-cache group_id is per-conversation (model+instructions+first
  item) instead of one global constant pooling every user
- messages() max_tokens default resolves per model (claude-3 caps at 4096)

Packaging / surface:
- __init__ registers the 0.2.10 modules in _SUBMODULES; unknown names
  raise AttributeError instead of eagerly importing page_index_classic
- anthropic floor 0.84.0: first release with ToolError whose runner also
  executes the final turn's tools on a max_iterations cut
- client docstrings caught up with local chat landing

Claude Agent SDK gate:
- claude_allowed_tools(mcp_servers) derives mcp__<key>__<tool> entries
  from the caller's own registration map (live server annotations on
  cloud, the contract locally) — no name is ever spelled twice
- claude_agent_config() bundles the three slots as one-call sugar over
  the explicit form

Examples:
- demo runs against cloud again (getattr for local-only attrs) and finds
  an existing indexed copy by name before re-indexing

Tests: monkeypatches replace the consuming module's binding instead of
mutating the shared time/requests modules; 185 -> 211.

* feat: one-call config bundles for every bring-your-own-framework surface

claude_agent_config() gets two symmetric siblings, so each framework's
front door is a single splat over the same explicit primitives:

- openai_agent_config(): Agent(**...) kwargs — instructions, tools, and
  the local retrieve_model (cloud omits model for the framework default)
- anthropic_runner_config(): tool_runner(**...) kwargs — system, tools,
  and the messages() defaults (per-model max_tokens, 10-iteration bound);
  only the user's messages remain

Bundles stay pure sugar: doc_id rides agent_instructions, no extra
semantics over the explicit form, docstrings point both ways. The demo
agent shrinks to Agent(**client.openai_agent_config(doc_id=...)).

Construction is pinned against the real frameworks in tests (Agent and
tool_runner both built offline), so an upstream kwargs rename fails
loudly; 211 -> 215 tests.

* fix: three more review findings — partial-read reporting, reply correlation, output_index axis

- get_page_content: the summary is additive, not either/or — a call that
  both truncates for size and has out-of-range pages reported only the
  latter, telling the agent every in-range page was returned (#2)
- McpBridge._extract_result: strict request-id correlation only; the
  eager fallback could hand back a stale or mis-correlated JSON-RPC
  message as this call's reply (#16)
- responses() streaming: output_index now addresses the logical
  response.output — backend per-turn indexes are re-based past prior
  turns' items and the SDK-injected tool outputs take the next slot on
  that axis, instead of reusing the event-sequence counter (#15)

215 -> 217 tests.

* feat: gate the config-handoff surfaces by the read-only MCP endpoint

pageindex-chat#448 adds /mcp?tools=read — the server registers only
readOnlyHint-annotated tools — so the URL itself becomes the gate for
every surface that hands a config to a third party:

- as_claude_mcp: include_management now picks the endpoint on cloud;
  the parameter is real in both modes
- as_openai_tools(hosted=True): OpenAI connects to the read-only
  endpoint by default and require_approval simplifies to "never" — the
  approval-flow middle ground becomes hard absence, matching every
  other surface's default
- claude_allowed_tools() retired before ever shipping: with the server
  gated, allowed_tools degenerates to whole-server pre-approval, which
  claude_agent_config emits as the constant ["mcp__<name>"] — no
  setup-time bridge round-trip remains
- in-process surfaces (agent_tools, as_openai_tools, as_anthropic_tools
  over the bridge) keep bare /mcp + client-side annotation filtering:
  they materialize tools locally and hand no URL to anyone

Release ordering: 0.2.10 must ship after pageindex-chat#448 deploys —
an older server ignores unknown query params and would silently serve
the full set behind a URL that promises read-only.

* fix: chat_completions wraps framework exceptions like responses()

The AgentsException -> PageIndexAPIError wrap from the responses() fix
covered only that surface; a backend stream dying without a terminal
event (or any engine failure) still escaped chat_completions as a raw
openai-agents exception type on both its paths.

* fix: same-name documents in different folders no longer refuse doc_id targeting

The shadow check in doc_targeting_block compared names across the whole
library, so agent_instructions(doc_id=...) hard-raised for a legal cloud
layout — one file name in two folders — with advice (rename/remove) that
contradicts the contract, whose folder_id parameter exists precisely to
disambiguate this case.

Shadowing is now judged per folder: only a newer same-name document in
the SAME folder makes the name unreachable and raises. A same-name
document in another folder serves the call, and the targeting block adds
a directive to pass folder_id on every tool call — dropping the raise
alone would have traded a loud refusal for the agent silently reading
the newer document.

Local mode (folderId always None) and the scoped chat path (allowlist
resolution, fixed with the doc_id enforcement) are behaviorally
unchanged.

* Revert "fix: same-name documents in different folders no longer refuse doc_id targeting"

This reverts commit 9d16dbf6, whose premise collapsed on verification
against the cloud upload paths. Review finding #10 inferred from the
folder_id tool description that one file name in two folders is a legal
cloud layout; both upload paths actually dedup names per USER SPACE with
no folder dimension — chat's getSignedUploadUrl queries fileName +
sourceName + mode + owner (file-access.service.ts), and compute's
get_upload_url probes the S3 key (user, source, file_name) — so own
documents cannot share a name across any folders. The only legitimate
same-name source is shared mounts (shared-with-me/following), which the
api-proxy surface this SDK talks to never carries.

Cloud and local therefore share one invariant — names unique per space,
server-enforced — and the original global shadow check was the right
shape: a duplicate is an anomaly worth refusing loudly, not a layout to
accommodate with per-folder adjudication and conditional prompt notes.
The invariant is now stated in doc_targeting_block's docstring so the
finding does not get re-raised.

* docs: state the name-uniqueness invariant in library terms

* docs: trim doc_targeting_block docstring to the contract

* fix: compress out-of-range page lists in get_page_content

Two message strings enumerated every out-of-range page number one by one
while the payload fields beside them already used _format_page_spec.
Against a 2-page document, pages="3-10000" produced a 59,310-character
error whose own requested_pages field expressed the identical set as
"3-10000"; the mixed case pages="1-10000" produced 59,479. Both now
render through the helper: 431 and 600 characters.

This was inherited behaviour, not a local slip — the cloud MCP server
enumerated at the same two sites, so local reproduced it verbatim. The
cloud fixed it first (pageindex-chat #449), and this follows to keep the
strings byte-identical; the error message now matches the served one
character for character. A differential run of the two compressors over
94 inputs (empty, single, unsorted, duplicated, 10k spans, 80 random)
agrees on every one, separator included.

The new test pins all three shapes, including the non-contiguous case
("5,9" must not collapse into a range) that the compressor had no direct
coverage for.

* docs: messages() marks only the managed prefix with cache_control

The method docstring claimed the doc targeting block carries a
cache_control breakpoint too, and the doc_id note called that block part
of the cached prompt prefix. _anthropic_system deliberately marks only
the stable managed prefix — the API allows four breakpoints and the
varying doc block must not consume one — and the block is appended after
the sole breakpoint, so it is never cached.

a45b554 added the block with cache_control, making the claim true when
written; daac9d2 removed it without touching the docstring, and adb2f1f
then added the "cached prompt prefix" sentence after the fact. The same
phrase at the chat_completions and responses docstrings is correct —
there the block is a leading conversation item inside the auto-cached
prefix — so only the Messages surface is reworded.

The doc_id advice itself stands: the block is per-call table-setting and
should stay identical across a conversation. Only the caching rationale
was wrong.

* docs: _stream_sync cancels on close, not on abandonment

The docstring promised that closing "or abandoning" the iterator cancels
the run. Abandoning only works when refcounting collects the generator:
a caller that breaks out of the loop while keeping the reference never
runs the finally that sets the cancel event, so the pump thread stays
parked on the full queue and the backend client is never released.

Closing is correct and is what the dedicated test exercises. Narrowing
the promise to the behaviour the code actually provides is the honest
fix; a watchdog or finalizer would be machinery bought for a shape the
sync surface is not meant to serve, and the async client planned for
0.2.11 gets native task cancellation instead.

* fix: raise the openai-agents floor to 0.14.0

_conversation_group_id feeds RunConfig.group_id into OpenAI's
prompt_cache_key so a round-tripped prefix stays in one cache group.
That wiring first appears in openai-agents 0.14.0: 0.8.0 through 0.13.x
have no prompt_cache_key at all, and group_id there is a tracing group
id only — inert, since tracing is disabled on the line above. An install
resolving to the declared floor lost the cache continuity that the
responses() docstring sells, silently and with no test able to catch it.

The old floor's rationale (0.8.0 offloads sync tools to a thread) is
subsumed by the new one. Every symbol the package imports predates
0.14.0, so nothing else constrains the bound.

* fix: enforce doc_id at the tool layer in the framework config helpers

openai_agent_config / anthropic_runner_config / claude_agent_config
accepted doc_id but built unscoped tools, so the parameter that is a
structural allowlist on chat_completions() was prompt-only advice here —
the agent could read every document in the store regardless.

- as_openai_tools / as_anthropic_tools / as_claude_mcp take a doc_id
  tail parameter and thread it to the existing _allowed_ids channel;
  the config helpers pass it through in local mode
- cloud config helpers keep prompt-level targeting (tool scoping is
  server-side there, documented); explicit as_*(doc_id=...) raises on
  cloud instead of silently dropping the allowlist — including the
  hosted branch, which returned before _tool_specs' existing guard
- _require_local_scope consolidates the cloud rejection that was
  inlined in _tool_specs
- doc_id=[] is an empty allowlist, not "unscoped": dropped the
  `or None` at the three local chat surfaces

* fix: two chat findings — final-turn append and cache-key seeding

run_messages keyed its re-append guard on stop_reason, but the anthropic
runner executes tools whenever the turn's content carries tool_use
blocks (refusal excepted) — a max_tokens turn with complete tool_use
blocks was already appended by the runner, so the guard re-appended it,
duplicating tool_use ids and 400ing the documented verbatim
continuation. The guard now checks whether final's tool_use ids already
sit in the appended history; unexecuted tool_use blocks (refusal turns)
are stripped from the appendable history, as the SDK itself does when
rebuilding params around an unresulted turn.

_conversation_group_id seeded on items[0], which is the doc-targeting
block whenever doc_id is set — byte-identical across every conversation
about a document, so all of them pooled under one prompt_cache_key and
evicted each other's prefixes. Seed on the conversation's own first
item instead: continuations keep their key, unrelated conversations
never share one.

Also drop the dead pytestmark_openai assignment (pytest's magic name is
pytestmark; the section gate it implied never existed).

* fix: six review findings — pagination, compat, and containment

- _all_documents advances by what actually arrived and treats `total`
  as an optimization: absent/null totals and short pages silently
  truncated the library behind every name resolution
- _make_bridge_function survives description: null (the parallel
  _tool_specs path already did)
- as_openai_tools answers a malformed argument string with the guided
  error envelope instead of raising through the caller's whole run
- the pre-0.2.10 package attributes (ConfigLoader, count_tokens, ...)
  resolve again: main's underscore-guarded fallthrough is restored —
  dunder probes stay lazy, a non-underscore typo pays one classic
  import before its AttributeError
- _split_structure chunks are always lists: the structure field no
  longer changes JSON type between parts of one paginated response
- the bridge replays only session-carrying 404s (the spec's expiry
  status); 400 raises instead of re-running side effects, and the
  reset double-checks under the lock so concurrent retries cannot
  clobber a freshly re-initialized session

* fix: five secondary review findings — containment and guards

- the bridge maps content blocks individually: base64 payloads
  (image/audio) become metadata stubs instead of handing the model the
  raw blob, text blocks pass verbatim, anything else keeps the JSON
  dump (revisit if tool results become real multimodal input)
- cloud proxy annotations keep array item types (list[str], not bare
  list) so strict function calling accepts the round-trip; a type-array
  in items degrades to bare list instead of crashing the build
- run_messages raises when set_messages_params stops delivering params
  instead of silently dropping every tool turn from the envelope
- call_tool drops None-valued arguments (None ≡ omitted, the
  contract's semantics) — adapters that forward the model's nulls
  verbatim no longer trip parameter validation
- client._parse_pages bounds the span arithmetically before
  materializing it, like the tool layer: "1-999999999" raises instead
  of allocating a billion integers

* fix: three review findings — protocol honesty, model echo, containment

- responses() promised the Responses protocol ("no translation layer")
  but _openai_model ignored protocol on the LiteLLM branch: provider-
  prefixed models silently ran chat.completions under a responses-shaped
  envelope, and with no transport hook to record status (LitellmModel
  has no _client.responses) a turn truncated at the output cap reported
  status "completed". The branch now raises for protocol == "responses"
  — at agent-build time, before any backend call — naming the routes
  out: chat_completions(), messages() for Anthropic models, or
  OPENAI_BASE_URL + a bare/openai/-prefixed name for backends that
  genuinely speak /responses. Refusal, not emulation: most providers
  have no /responses endpoint to drive.
- chat_completions envelopes echoed retrieve_model verbatim, which
  carries the SDK's litellm/ routing marker after normalization — a
  name no provider catalog contains, and a different string than the
  same model passed per-call. The envelope and every streaming chunk
  now report the name the provider actually serves; routing and the
  prompt-cache group key keep the prefixed form. responses() needs no
  change (post-refusal the prefix cannot reach its envelope), and the
  user-typed openai/ prefix stays echoed as typed.
- _remove_document caught only PageIndexAPIError around the per-doc
  delete, so a bare OSError (local_store re-raises them) or a transport
  error (cloud delete_document wraps nothing) escaped mid-batch,
  discarded the entries for documents already irreversibly deleted, and
  surfaced as a generic INTERNAL_ERROR envelope inviting a retry — which
  then reports the destroyed document as not_found. The loop now catches
  Exception, keeping the per-document results the contract promises.

* fix: config bundles use the scoped shadow check their tools earned

9f67fdd made the three config helpers enforce doc_id at the tool layer
but left their instructions on doc_targeting_block's unscoped default,
so a bundle refused any doc_id whose name a newer library-wide
duplicate shadows — a raise whose message ("the tools address documents
by name and would read the newer one") had just become false: the
bundle's own tools resolve names inside the allowlist and read the
targeted document correctly. chat_completions() accepted the same
doc_id via _doc_block's scoped=True.

Each helper now computes scope = _local_doc_scope(doc_id) once and
derives both slots from it — scoped=scope is not None for the
instructions, doc_id=scope for the tools — so the check mode and the
tool allowlist come from one fact and cannot drift apart again.
build_agent_instructions grows a scoped passthrough; cloud stays on the
whole-library check (scope is None there and the tools are genuinely
unscoped), and the public agent_instructions() keeps its unscoped
default for the same reason. An in-set duplicate still raises — and in
that case the message is true on every surface that emits it.

* docs: as_openai_tools' remote-MCP note moves to the Cloud paragraph

The MCPServerStreamableHttp alternative sat in the Local: paragraph
pointing at bare {BASE_URL}/mcp — a cloud-only route (BASE_URL is the
hosted API; local has no HTTP MCP server) that as written would connect
unauthenticated to the full tool set. Now stated where it applies, in
the as_anthropic_tools connector-note form: Cloud paragraph, Bearer
auth spelled out, ?tools=read default with the drop-it escape.

* fix: nine review findings — argument coercion, scope, and honest envelopes

- call_tool coerces string booleans per the TOOL_CONTRACT schema
  ("false"/"no"/"0" read as False, not a truthy 3-minute wait) and
  survives arguments: null (json.loads("null") reaches the seam as None)
- _local_doc_scope raises on an explicitly empty doc_id on cloud: with
  no tool-layer allowlist there, dropping it silently widened an empty
  scope to the whole library
- both page-spec caps count distinct pages instead of summing parts, so
  overlapping ranges (a parent section plus its children) within the
  10k union pass again as they did in 0.2.9; the per-part arithmetic
  bound still rejects billion-page specs before materializing anything
- _remove_document deduplicates doc_names: a repeated name is one
  deletion, not a second "failed" row with an internal error string
- doc_targeting_block merges the user's metadata tags from the listing
  (local get_document keeps the 7-key cloud detail wire shape, which
  carries none) so the block delivers the metadata it promises
- _wait_until_ready folds its two raise branches into one that carries
  the doc_id: a poll that dies no longer discards the handle to an
  uploaded, billed document
- _reported_model strips both routing prefixes (litellm/ and openai/)
  and responses() now reports it too, instead of echoing a model id the
  provider never served
- _openai_model wraps AsyncOpenAI() construction so a missing backend
  credential surfaces as PageIndexAPIError like every other gate on the
  chat surfaces (and builds the client once for both protocols)
- _browse_documents advances its cursor by the rows that actually
  arrived and guards a null/absent total — the same hazards
  _all_documents already guards — and an empty window ends pagination
  instead of freezing the cursor

* fix: two chat findings — protocol terminal states, provider error types

- responses(stream=True) raised PageIndexAPIError when the backend
  ended the response with response.failed / response.incomplete:
  openai-agents yields the terminal lifecycle event, then re-raises it
  as ModelBehaviorError, so the generic AgentsException wrap
  short-circuited the emit the agen's tail was built for — its
  failed/incomplete terminal mapping was dead code against the real
  engine, and the caller lost both the partial output and the real
  status. The wrap now steps aside when the recorded terminal state is
  failed/incomplete, and the stream ends with the honest terminal
  event (committed output, real status, error/incomplete_details) —
  the backend's terminal state is a protocol event, not an engine
  failure. Non-stream was already honest for incomplete via the
  transport recorder; a failed response arrives there as an HTTP
  error, covered below. Known ceiling: the truncated final turn's
  partial text was already streamed as deltas but is not reconstructed
  into the terminal event's output (the engine commits items only on
  turn completion).
- Provider exceptions (network, auth, rate limit) leaked as raw
  openai/anthropic types through every chat surface, against the
  layer's own "never raw engine types" contract. Every engine boundary
  now wraps its vendor's base exception into PageIndexAPIError
  (chained): the four OpenAI-engine sites catch openai.OpenAIError —
  LiteLLM's exception types subclass openai's, so one handler covers
  both routing paths — and messages() catches anthropic.AnthropicError
  around the batch drive and the stream generator.

* fix: guided failure for unknown LiteLLM providers, non-object call_tool args

- _openai_model pre-checks the first path segment against
  litellm.provider_list (fail-open if the attribute ever disappears):
  a HuggingFace repo id like Qwen/Qwen2.5-7B-Instruct on an
  OpenAI-compatible server now fails at build time with the escape
  spelled out — 'openai/<id>' plus OPENAI_BASE_URL — instead of at
  request time inside LiteLLM with "LLM Provider NOT provided". The
  slash-means-provider routing convention itself is unchanged; the
  retrieve_model and chat_completions docstrings now document it where
  they promise "any OpenAI-compatible server works"
- call_tool answers a non-dict arguments value (a JSON array or scalar
  from a misbehaving caller) with the guided INVALID_INPUT envelope
  instead of raising AttributeError through the agent loop, matching
  the openai adapter's own non-object guard

* fix: wrap litellm import in PageIndexAPIError when not installed

* fix: silence CodeQL findings — merge implicit string concat, drop unused vars

* fix: two external review findings — init-notification race, SDK floor

notifications/initialized moves inside the bridge lock: a concurrent
first use could send tools/list between the handshake and the
notification, which strict MCP servers reject with a 400 the bridge
never replays. Regression test races two threads through a stalled
notification window.

claude-agent-sdk floor rises to 0.1.53 — below it, string prompts with
SDK MCP servers (the documented local-mode flow) hit invisible
registration (#597) and a deadlock (#780).

* refactor: drop the unused exc parameter from _wrap_max_turns

The parameter was dead from the moment it was introduced (daac9d2):
the body reads only max_turns, and every call site already carries the
cause via `raise ... from exc`. The signature implied the helper
inspected the engine exception, which it never did.

No behavior change — message text and __cause__ chaining verified
identical across all four call sites (chat_completions and responses,
stream and non-stream).

* fix: raise the anthropic and openai-agents floors past broken releases

Both declared floors named a version that cannot work, and CI never
caught either because it installs the latest.

anthropic >=0.84.0 -> >=0.108.0. Probed against a mock transport: on a
turn with stop_reason="refusal" carrying a tool_use block, 0.84.0,
0.92.0 and 0.100.0 all execute the tool and post the tool_result back;
0.108.0 and later stop at the refusal. test_messages_refusal_with_
tool_use_stays_appendable asserts the latter, so that test was false at
the floor. messages() is unaffected in practice (it never passes
include_management, so remove_document is not registered), but
as_anthropic_tools(include_management=True) hands it to a caller's own
runner.

openai-agents >=0.14.0 -> >=0.18.1. 0.14.0 and 0.16.0 raise pydantic
ValidationError on InputTokensDetails.cache_write_tokens before any
request reaches the transport when paired with openai 2.54.0 — and they
declare openai <3,>=2.26.0, so pip resolves exactly that pair. 0.18.1 is
clean. The 0.14.0 rationale (RunConfig.group_id -> prompt_cache_key)
still holds above the new floor.

The three extras' floor comments are cut to the binding constraint; the
reasoning lives here.

* test: cover max_turns wrapping on every chat surface

test_chat_completions_max_turns_wrapped only drove chat_completions, so
the two responses() call sites had no coverage, and no test asserted
that the engine exception survives as __cause__. Parametrized over both
surfaces and both stream modes; the non-positive max_turns rejection
splits out, since it is input validation rather than wrapping.

* fix: four review findings — envelope size honesty, contained tool errors

- _dumps drops indent=2: emission now matches _serialized_size's compact
  accounting, so the pagination budget bounds what is actually sent
  (indented parts measured under 95k but emitted ~1.8x the 100k cap)
- call_tool builds the _allowed_ids frozenset inside the guarded block:
  a non-iterable doc_id returns the INVALID_INPUT envelope instead of
  raising into the agent loop; same move for _bridge_invoker's
  arguments normalization
- next_steps strings qualify submit_document() as
  PageIndexClient.submit_document() (three sites), matching the one
  already-qualified site — it is a client method, not a registered tool
- tests: import httpx at module scope (guaranteed via the hard openai
  dependency) so agents-gated tests survive an install without the
  anthropic extra; formatting assertion follows the compact envelope
This commit is contained in:
Ray
2026-08-13 07:32:11 +08:00
committed by GitHub
parent d375c00a5a
commit 4e41acdc68
21 changed files with 7404 additions and 100 deletions
+1 -1
View File
@@ -31,5 +31,5 @@ jobs:
cache: pip
- run: pip install -r requirements.txt pytest
- if: matrix.agent-frameworks == 'with'
run: pip install openai-agents claude-agent-sdk
run: pip install openai-agents claude-agent-sdk anthropic
- run: python -m pytest -q
+1
View File
@@ -6,3 +6,4 @@ __pycache__
logs/
.pageindex/
dist/
*.doc_id
+32 -51
View File
@@ -6,20 +6,21 @@ local mode and the OpenAI Agents SDK. Instead of vector similarity search and
chunking, PageIndex builds a hierarchical tree index and uses agentic LLM
reasoning for human-like, context-aware retrieval.
Agent tools:
- get_document() — document metadata (status, page count, etc.)
- get_document_structure() — tree structure index of a document
- get_page_content() — retrieve text content of specific pages
The agent tools come straight from the SDK — ``client.as_openai_tools()``
exposes the PageIndex tool contract (browse_documents, get_document,
get_document_structure, get_page_content) and ``client.agent_instructions()``
provides the retrieval playbook, so the whole agent is a few lines. Swap
``PageIndexLocalClient()`` for ``PageIndexCloudClient(api_key=...)`` and the
same code runs against the cloud.
Steps:
1 — Index a PDF locally and view its tree structure index
2 — View document metadata
3 — Ask a question (agent reasons over the index and auto-calls tools)
Requirements: pip install openai-agents; OPENAI_API_KEY in the environment.
Requirements: pip install "pageindex[openai]"; OPENAI_API_KEY in the environment.
"""
import sys
import json
import asyncio
import concurrent.futures
from pathlib import Path
@@ -27,62 +28,30 @@ import requests
sys.path.insert(0, str(Path(__file__).parent.parent))
from agents import Agent, Runner, function_tool, set_tracing_disabled
from agents.model_settings import ModelSettings
from agents import Agent, Runner, set_tracing_disabled
from agents.stream_events import RawResponsesStreamEvent, RunItemStreamEvent
from openai.types.responses import ResponseTextDeltaEvent, ResponseReasoningSummaryTextDeltaEvent
from pageindex import PageIndexClient
from pageindex import PageIndexAPIError, PageIndexLocalClient
import pageindex.utils as utils
PDF_URL = "https://arxiv.org/pdf/2603.15031"
_EXAMPLES_DIR = Path(__file__).parent
PDF_PATH = _EXAMPLES_DIR / "documents" / "attention-residuals.pdf"
DOC_ID_PATH = _EXAMPLES_DIR / "documents" / "attention-residuals.doc_id"
STORAGE_PATH = _EXAMPLES_DIR / ".pageindex"
AGENT_SYSTEM_PROMPT = """
You are PageIndex, a document QA assistant.
TOOL USE:
- Call get_document() first to confirm status and page count.
- Call get_document_structure() to identify relevant page ranges.
- Call get_page_content(pages="5-7") with tight ranges; never fetch the whole document.
- Before each tool call, output one short sentence explaining the reason.
Answer based only on tool output. Be concise.
"""
def query_agent(client: PageIndexClient, doc_id: str, prompt: str, verbose: bool = False) -> str:
def query_agent(client: PageIndexLocalClient, doc_id: str, prompt: str, verbose: bool = False) -> str:
"""Run a document QA agent using the OpenAI Agents SDK.
Streams text output token-by-token and returns the full answer string.
Tool calls are always printed; verbose=True also prints arguments and output previews.
"""
@function_tool
def get_document() -> str:
"""Get document metadata: status, page count, name, and description."""
return json.dumps(client.get_document(doc_id))
@function_tool
def get_document_structure() -> str:
"""Get the document's full tree structure (without text) to find relevant sections."""
return json.dumps(client.get_document_structure(doc_id), ensure_ascii=False)
@function_tool
def get_page_content(pages: str) -> str:
"""
Get the text content of specific pages.
Use tight ranges: e.g. '5-7' for pages 5 to 7, '3,8' for pages 3 and 8, '12' for page 12.
"""
return json.dumps(client.get_page_content(doc_id, pages), ensure_ascii=False)
agent = Agent(
name="PageIndex",
instructions=AGENT_SYSTEM_PROMPT,
tools=[get_document, get_document_structure, get_page_content],
model=getattr(client, "retrieve_model", None),
# model_settings=ModelSettings(reasoning={"effort": "low", "summary": "auto"}), # Uncomment to enable reasoning
**client.openai_agent_config(doc_id=doc_id),
# model_settings=ModelSettings(reasoning={"effort": "low", "summary": "auto"}), # from agents.model_settings import ModelSettings
)
async def _run():
@@ -152,21 +121,33 @@ if __name__ == "__main__":
print("Download complete.\n")
# Setup: local mode — no PageIndex API key needed, your LLM key does the work
client = PageIndexClient(storage_path=str(STORAGE_PATH))
client = PageIndexLocalClient(storage_path=str(STORAGE_PATH))
# Step 1: Index PDF and view tree structure
print("=" * 60)
print("Step 1: Index PDF and view tree structure")
print("=" * 60)
doc_id = next(
(doc["id"] for doc in client.list_documents(limit=100)["documents"]
if doc["name"] == PDF_PATH.name),
None,
)
doc_id = None
if DOC_ID_PATH.exists():
cached = DOC_ID_PATH.read_text().strip()
try:
client.get_document(cached)
doc_id = cached
except PageIndexAPIError:
DOC_ID_PATH.unlink()
if doc_id is None:
# The .doc_id cache is gitignored — on a fresh clone with an
# existing store, find the already-indexed copy by name instead of
# re-indexing it.
doc_id = next(
(doc["id"] for doc in client.list_documents(limit=100)["documents"]
if doc["name"] == PDF_PATH.name), None)
if doc_id:
DOC_ID_PATH.write_text(doc_id)
print(f"\nLoaded cached doc_id: {doc_id}")
else:
doc_id = client.submit_document(str(PDF_PATH))["doc_id"]
doc_id = client.submit_document(str(PDF_PATH), wait=True)["doc_id"]
DOC_ID_PATH.write_text(doc_id)
print(f"\nIndexed. doc_id: {doc_id}")
print("\nTree Structure (top-level sections):")
structure = client.get_tree(doc_id, node_summary=True)["result"]
+16 -6
View File
@@ -18,26 +18,36 @@ __all__ = [
]
_LAZY = {
"page_index": ".page_index_classic",
"page_index_main": ".page_index_classic",
"page_index_flash": ".flash",
"optimize_tree": ".tree_optimize",
"md_to_tree": ".page_index_md",
}
_SUBMODULES = {"client", "cloud_api", "errors", "flash", "local_api",
"local_store", "page_index_classic", "page_index_md", "tree_optimize",
"utils"}
_SUBMODULES = {"agent_tools", "client", "cloud_api", "errors", "flash",
"integrations", "local_api", "local_chat", "local_store",
"mcp_bridge", "page_index_classic", "page_index_md",
"tree_optimize", "utils"}
def __getattr__(name):
if name.startswith("_"):
# Dunder probes (copy, pickle, inspect) are the frequent unknown
# names — they must not trigger the classic import below.
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
import importlib
if name in _SUBMODULES:
return importlib.import_module(f".{name}", __name__)
module = importlib.import_module(_LAZY.get(name, ".page_index_classic"), __name__)
# Pre-0.2.10 compat: unknown names fall through to the classic module,
# whose public surface (ConfigLoader, count_tokens, ...) resolved as
# package attributes. A non-underscore typo pays one classic import
# before its AttributeError — not worth an allowlist.
module = importlib.import_module(_LAZY.get(name, ".page_index_classic"),
__name__)
try:
value = getattr(module, name)
except AttributeError:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
raise AttributeError(
f"module {__name__!r} has no attribute {name!r}") from None
globals()[name] = value
return value
+10
View File
@@ -0,0 +1,10 @@
"""Installed-package version, shared by every surface that reports it upstream."""
from __future__ import annotations
def sdk_version() -> str:
try:
from importlib.metadata import version
return version("pageindex")
except Exception:
return "0.0.0"
File diff suppressed because it is too large Load Diff
+588 -35
View File
@@ -1,23 +1,36 @@
"""PageIndex SDK client: the 0.2.x cloud surface, now with a local mode."""
from __future__ import annotations
from typing import Any, Iterator, Optional, Union
import os
import time
import warnings
from typing import Any, Callable, Iterator, Optional, Union
from .errors import PageIndexAPIError
def _parse_pages(pages: str) -> list[int]:
result = []
result: set[int] = set()
too_many = (f"Page specification '{pages}' spans more than "
"10000 pages; request a narrower range")
for part in pages.split(","):
part = part.strip()
if "-" in part:
start, end = (int(x) for x in part.split("-", 1))
if start > end:
raise ValueError(f"Invalid range '{part}': start must be <= end")
result.extend(range(start, end + 1))
else:
result.append(int(part))
return sorted(set(result))
start = end = int(part)
# Bound each part arithmetically before materializing it — a spec
# like "1-999999999" would otherwise expand to a billion integers.
# The cap is on distinct pages, so overlapping parts (a parent
# section plus its children) don't double-count.
if end - start + 1 > 10_000:
raise ValueError(too_many)
result.update(range(start, end + 1))
if len(result) > 10_000:
raise ValueError(too_many)
return sorted(result)
def _normalize_retrieve_model(model: str) -> str:
@@ -47,10 +60,12 @@ class PageIndexClient:
trees. Defaults to the packaged config (see pageindex/config.yaml).
summary_model (str, optional): Local mode only — LLM used for node
summaries and document descriptions.
retrieve_model (str, optional): Local mode only — exposed as
``client.retrieve_model`` (the agent demo reads it); the SDK
itself consumes it once agent-based local chat lands in a
later release.
retrieve_model (str, optional): Local mode only — the model the
local chat surfaces (``chat_completions``, ``responses``)
default to, exposed as ``client.retrieve_model``.
``provider/model`` names route through LiteLLM; for an
OpenAI-compatible server that itself serves slashed model ids
(vLLM, TGI), prefix ``openai/`` (e.g. ``openai/Qwen/...``).
storage_path (str, optional): Local mode only — directory where
indexed documents are stored. Defaults to ``./.pageindex``.
@@ -62,10 +77,9 @@ class PageIndexClient:
instead of inferring it from api_key.
Local mode differences (all documented per method): indexing is
synchronous, only PDFs are supported, and ``chat_completions`` (until
agent-based local chat lands in a later release) / folders /
``beta_headers`` / the deprecated retrieval API (``submit_query``,
``get_retrieval``) are cloud-only.
synchronous, only PDFs are supported, and folders / ``beta_headers`` /
the deprecated retrieval API (``submit_query``, ``get_retrieval``) are
cloud-only.
"""
BASE_URL = "https://api.pageindex.ai"
@@ -126,12 +140,14 @@ class PageIndexClient:
beta_headers: Optional[list[str]] = None,
folder_id: Optional[str] = None,
metadata: Optional[dict] = None,
wait: bool = False,
) -> dict[str, Any]:
"""
Submit a PDF document for processing. Returns {'doc_id': ...}.
Submit a PDF document for processing. Returns {'doc_id': ..., 'name': ...}.
Cloud: uploads the file; processing is asynchronous — poll
``is_retrieval_ready(doc_id)`` before retrieving.
Cloud: uploads the file; processing is asynchronous. Pass
``wait=True`` to block until the document is ready, or poll
``get_document(doc_id)['status']`` yourself.
Local: indexes the document in this call (it blocks while your LLM
builds the tree — minutes for a standard index of a long document),
@@ -151,14 +167,67 @@ class PageIndexClient:
metadata (dict, optional): Your own JSON-serializable tags for the
document; returned in get_tree/get_ocr responses and
list_documents entries (both modes).
wait (bool): Return only once the document is ready for use.
Cloud: polls status until "completed" (raises on "failed" or
after 30 minutes). Local: indexing is synchronous already, so
this changes nothing. Leave False to submit many documents
concurrently and poll afterwards.
Returns:
dict: {'doc_id': ...}
dict: {'doc_id': ..., 'name': ...}. 'name' is the stored document
name: a taken name gains a numeric suffix (name_1..name_99)
and a UserWarning is emitted. Older cloud servers omit 'name'.
"""
return self._api.submit_document(
result = self._api.submit_document(
file_path=file_path, mode=mode,
beta_headers=beta_headers, folder_id=folder_id, metadata=metadata,
)
stored = result.get("name")
if stored and stored != os.path.basename(file_path):
warnings.warn(
f'Document "{os.path.basename(file_path)}" was stored as '
f'"{stored}".',
stacklevel=2,
)
if wait:
self._wait_until_ready(result["doc_id"])
return result
def _wait_until_ready(self, doc_id: str, timeout: float = 1800.0) -> None:
import requests
interval = 2.0
deadline = time.monotonic() + timeout
poll_failures = 0
while True:
try:
status = self.get_document(doc_id).get("status")
poll_failures = 0
except (PageIndexAPIError, requests.RequestException) as exc:
# Tolerate transient poll failures; a 30-minute wait should
# not die on one 502 or dropped connection.
poll_failures += 1
if poll_failures >= 3:
raise PageIndexAPIError(
f"Could not poll document status (doc_id: {doc_id}): "
f"{exc}. Processing continues in the cloud — poll "
"get_document(doc_id) for status."
) from exc
status = None
if status == "completed":
return
if status == "failed":
raise PageIndexAPIError(
f"Document processing failed (doc_id: {doc_id})."
)
if time.monotonic() >= deadline:
raise PageIndexAPIError(
f"Timed out after {int(timeout)}s waiting for document "
f"processing (doc_id: {doc_id}, last status: {status}). "
"Processing continues in the cloud — poll "
"get_document(doc_id) for status."
)
time.sleep(interval)
interval = min(interval * 1.5, 15.0)
# ---------- OCR FUNCTIONALITY ----------
@@ -256,11 +325,11 @@ class PageIndexClient:
Cloud-only: the cloud API marks this endpoint deprecated in favor of
chat completions, so local mode does not implement it — raises
PageIndexAPIError. Use ``chat_completions`` (cloud) instead.
PageIndexAPIError. Use ``chat_completions`` instead.
"""
return self._require_cloud(
"submit_query is cloud-only — the retrieval API is deprecated in "
"favor of chat completions; use chat_completions in cloud mode."
"favor of chat completions; use chat_completions instead."
).submit_query(doc_id=doc_id, query=query, thinking=thinking)
def get_retrieval(self, retrieval_id: str) -> dict[str, Any]:
@@ -269,55 +338,217 @@ class PageIndexClient:
Cloud-only: the cloud API marks this endpoint deprecated in favor of
chat completions, so local mode does not implement it — raises
PageIndexAPIError. Use ``chat_completions`` (cloud) instead.
PageIndexAPIError. Use ``chat_completions`` instead.
"""
return self._require_cloud(
"get_retrieval is cloud-only — the retrieval API is deprecated in "
"favor of chat completions; use chat_completions in cloud mode."
"favor of chat completions; use chat_completions instead."
).get_retrieval(retrieval_id=retrieval_id)
# ---------- CHAT COMPLETIONS ----------
def chat_completions(
self,
messages: list[dict[str, str]],
messages: Union[str, list[dict[str, str]]],
stream: bool = False,
doc_id: Optional[Union[str, list[str]]] = None,
temperature: Optional[float] = None,
stream_metadata: bool = False,
enable_citations: bool = False,
model: Optional[str] = None,
max_turns: Optional[int] = None,
) -> Union[dict[str, Any], Iterator[str], Iterator[dict[str, Any]]]:
"""
PageIndex Chat Completions, scoped to specific PageIndex documents.
PageIndex Chat Completions: document QA in one call.
Cloud: the hosted chat endpoint. Local: a managed document-QA agent
run over the local tools against your own LLM backend's
/chat/completions (requires ``pageindex[openai]``; the OpenAI SDK's
usual env config — OPENAI_API_KEY, OPENAI_BASE_URL — selects the
backend, so any OpenAI-compatible server works; a ``/`` in the
model name means LiteLLM provider routing, so prefix ``openai/``
when the backend itself serves slashed ids, e.g.
``openai/Qwen/...`` on vLLM). The non-stream
response carries the final answer only; streaming yields the
agent's visible text as it is produced, including narration before
tool calls. ``finish_reason`` reports loop completion ("stop") —
the engine does not surface per-turn backend finish reasons. For
the tool-use process and prompt-cache round-trip use
``responses()`` or ``messages()``.
Args:
messages: Conversation messages with 'role' and 'content' keys.
messages: Conversation messages with 'role' and 'content' keys,
or a bare query string (it becomes a single user message).
Local also accepts system/developer messages — their content
is appended to the managed system prompt.
stream: Enable streaming responses.
doc_id: Document ID or list of IDs to scope the conversation.
temperature: Sampling temperature (0.0-1.0).
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call and is part
of the cached prompt prefix.
temperature: Sampling temperature, passed through to the model.
stream_metadata: With stream=True, yield chunk dicts instead of
text pieces.
enable_citations: Enable citation instructions in responses.
enable_citations: Cloud-only — local mode raises (citations need
block-level OCR data local mode does not store).
model: Local only — backend model name (defaults to
``retrieve_model``). The cloud endpoint selects its own.
max_turns: Local only — cap on agent turns per call.
Returns:
- stream=False: complete response dict ({'id', 'object', 'created',
'choices', 'usage'})
- stream=True, stream_metadata=False: iterator of text chunks
- stream=True, stream_metadata=True: iterator of chunk dicts
Local: not yet supported — raises PageIndexAPIError. Agent-based
local chat arrives in a later release.
"""
return self._require_cloud(
"chat_completions is not yet supported in local mode — it arrives "
"in a later release. Create the client with an api_key to use "
"cloud chat."
).chat_completions(
if isinstance(messages, str):
if not messages.strip():
raise PageIndexAPIError(
"messages must be a non-empty string or a list of "
"message dicts.")
messages = [{"role": "user", "content": messages}]
from .cloud_api import CloudAPI
if not isinstance(self._api, CloudAPI):
from .local_chat import run_chat_completions
return run_chat_completions(
self, messages, stream=stream, doc_id=doc_id,
temperature=temperature, stream_metadata=stream_metadata,
enable_citations=enable_citations, model=model,
max_turns=max_turns,
)
if model is not None or max_turns is not None:
raise PageIndexAPIError(
"model and max_turns are local-mode parameters — the cloud "
"chat endpoint selects its own model."
)
return self._api.chat_completions(
messages=messages, stream=stream, doc_id=doc_id,
temperature=temperature, stream_metadata=stream_metadata,
enable_citations=enable_citations,
)
def responses(
self,
input: Union[str, list[dict[str, Any]]],
model: Optional[str] = None,
stream: bool = False,
doc_id: Optional[Union[str, list[str]]] = None,
instructions: Optional[str] = None,
temperature: Optional[float] = None,
top_p: Optional[float] = None,
max_turns: Optional[int] = None,
) -> Union[dict[str, Any], Iterator[dict[str, Any]]]:
"""
Document QA over the OpenAI Responses protocol — the agentic surface.
Local only for now. Drives your backend's /responses end to end (no
translation layer), so the ``output`` carries the whole process as
standard items — messages, function calls, and function outputs
(the SDK executes the tools). Append the returned ``output`` to your
next call's ``input`` verbatim to keep provider prompt-cache prefix
continuity and the agent's memory of what it already read.
Requires ``pageindex[openai]`` and a backend that supports the
Responses API; backends that only speak chat.completions should use
``chat_completions()``. Provider-prefixed models (``anthropic/…``)
route through LiteLLM's chat.completions adapter and are therefore
refused here — use ``chat_completions()`` or ``messages()`` for
those.
Args:
input: A user message string, or a list of Responses input items
(round-trip prior ``output`` items here).
model: Backend model name (defaults to ``retrieve_model``).
stream: Yield Responses stream events as dicts — one logical
response per call: per-turn backend lifecycle events are
collapsed, sequence numbers are reassigned monotonically,
and ``output_index`` is re-based onto the single logical
``output``; tool outputs are emitted as
``response.output_item.done`` events and the single final
event is the terminal ``response.*`` for the run's status.
doc_id: Document ID or list of IDs to scope the conversation.
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call and is part
of the cached prompt prefix.
instructions: Appended to the managed system prompt.
temperature / top_p: Passed through to the model.
max_turns: Cap on agent turns per call.
"""
from .cloud_api import CloudAPI
if isinstance(self._api, CloudAPI):
raise PageIndexAPIError(
"responses is not available on PageIndex cloud yet — it is "
"a local-mode surface for now."
)
from .local_chat import run_responses
return run_responses(
self, input, model=model, stream=stream, doc_id=doc_id,
instructions=instructions, temperature=temperature, top_p=top_p,
max_turns=max_turns,
)
def messages(
self,
messages: Union[str, list[dict[str, Any]]],
model: str,
max_tokens: Optional[int] = None,
stream: bool = False,
doc_id: Optional[Union[str, list[str]]] = None,
system: Optional[Union[str, list[dict[str, Any]]]] = None,
temperature: Optional[float] = None,
top_p: Optional[float] = None,
top_k: Optional[int] = None,
stop_sequences: Optional[list[str]] = None,
max_turns: Optional[int] = None,
) -> Union[dict[str, Any], Iterator[Any]]:
"""
Document QA over the Anthropic Messages protocol — Claude-native.
Local only for now. Drives Anthropic's /v1/messages via the
Anthropic SDK's own tool runner (requires ``pageindex[anthropic]``;
ANTHROPIC_API_KEY selects the backend). ``tool_use``/``tool_result``
round-trip is the format's native behavior: the response is the
final message envelope with cross-turn aggregated ``usage`` plus a
``messages`` field — the full new turn sequence, valid for verbatim
append to your history. The managed system prompt carries a
``cache_control`` breakpoint.
Args:
messages: Native Messages-format history (including prior
tool_use/tool_result blocks on round-trip), or a bare query
string (it becomes a single user message).
model: Required — there is no cross-vendor default to guess.
max_tokens: Per-turn output budget the Messages API requires on
the wire; the default is resolved per model (8192, or 4096
for the claude-3 generation whose ceiling is lower) so the
simple call needs only a question. Passed through.
stream: Yield the Anthropic SDK's event stream across turns
(its native event objects, including SDK-synthesized
convenience events), one message sequence per turn.
doc_id: Document ID or list of IDs to scope the conversation.
Keep it identical across a conversation's calls — the
targeting block it adds is re-set each call.
system: Appended after the managed system blocks.
temperature / top_p / top_k / stop_sequences: Passed through.
max_turns: Cap on agent turns per call (default 10, like the
OpenAI surfaces). A truncated run reports
``stop_reason: "tool_use"`` and its ``messages`` remain
valid for continuation.
"""
from .cloud_api import CloudAPI
if isinstance(self._api, CloudAPI):
raise PageIndexAPIError(
"messages is not available on PageIndex cloud yet — it is "
"a local-mode surface for now."
)
from .local_chat import run_messages
return run_messages(
self, messages, model=model, max_tokens=max_tokens,
stream=stream, doc_id=doc_id, system=system,
temperature=temperature, top_p=top_p, top_k=top_k,
stop_sequences=stop_sequences, max_turns=max_turns,
)
# ---------- DOCUMENT MANAGEMENT ----------
def get_document(self, doc_id: str) -> dict[str, Any]:
@@ -365,6 +596,328 @@ class PageIndexClient:
"""
return self._api.list_documents(limit=limit, offset=offset, folder_id=folder_id)
# ---------- AGENT INTEGRATION ----------
def agent_tools(self, include_management: bool = False) -> list[Callable[..., str]]:
"""
Plain functions for any agent framework (LangChain, PydanticAI, ...).
For the OpenAI / Claude Agent SDKs, prefer ``as_openai_tools()`` /
``as_claude_mcp()``.
Cloud: the full cloud tool set, discovered live from the PageIndex
MCP server when this method is called — one function per tool,
signature and docstring synthesized from the server's schemas, calls
executed from your process over MCP. Raises PageIndexAPIError if the
server cannot be reached. Local: the built-in tools over the local
store (``browse_documents``, ``get_document``,
``get_document_structure``, ``get_page_content``).
Each function takes JSON-serializable arguments, returns a JSON
string, and reports failures inside that JSON instead of raising.
Args:
include_management (bool): Also expose tools that modify the
library. Local: adds ``remove_document``. Cloud: by default
only tools the server marks read-only are exposed; True
exposes the server's complete list (upload, delete, ...).
"""
from .agent_tools import build_agent_tools
return build_agent_tools(self, include_management)
def as_openai_tools(self, include_management: bool = False,
hosted: bool = False,
doc_id: Optional[Union[str, list[str]]] = None) -> list:
"""
Tools for the OpenAI Agents SDK — pass to ``Agent(tools=...)``
(or ``openai_agent_config()`` for all the Agent slots in one
call).
Cloud (default): the full live read tool set (search, folders,
images — as enabled for your key) as plain function tools,
discovered from the PageIndex MCP server and executed from your
process — works with any model backend. Pass ``hosted=True`` to
hand the connection to OpenAI instead: one hosted MCP tool, tool
calls executed server-side (lowest latency; requires an
OpenAI-hosted model on the Responses API). The framework's own
``MCPServerStreamableHttp`` — ``params={"url":
f"{BASE_URL}/mcp?tools=read", "headers": {"Authorization":
"Bearer <your PageIndex API key>"}}`` (drop ``?tools=read`` for
the full tool set) — is the async-native alternative for its
``mcp_servers=`` slot.
Local: the in-process tools, any model backend; ``hosted`` does
not apply.
Requires ``openai-agents`` (``pip install 'pageindex[openai]'``),
imported only when this method is called.
Args:
include_management (bool): Also expose tools that modify the
library (delete, upload). Default off: the in-process
cloud default serves only server-annotated read-only
tools, and ``hosted=True`` connects OpenAI to the
read-only endpoint (``/mcp?tools=read``) instead.
hosted (bool): Cloud only — hand the MCP connection to OpenAI
for server-side tool execution (OpenAI models only).
doc_id: Local only — restrict the tools to this document ID
(or list of IDs), enforced at the tool layer: out-of-scope
lookups return NOT_FOUND. Raises on cloud, where scoping
is server-side.
"""
from .integrations.openai_agents import build_openai_tools
return build_openai_tools(self, include_management, hosted,
doc_ids=doc_id)
def _local_doc_scope(self, doc_id):
"""doc_id for the tool layer: passed through locally (structural
allowlist), dropped on cloud where scoping is server-side and the
config helpers keep prompt-level targeting."""
if not getattr(self, "api_key", None):
return doc_id
if doc_id is not None and not doc_id:
# Cloud has no tool-layer allowlist to make an empty scope mean
# "nothing"; dropping it would silently mean "everything".
raise PageIndexAPIError(
"doc_id is empty. Pass one or more document IDs, or omit "
"doc_id to give the agent the whole library.")
return None
def openai_agent_config(
self,
doc_id: Optional[Union[str, list[str]]] = None,
include_management: bool = False,
model: Optional[str] = None,
) -> dict[str, Any]:
"""
Document QA ``Agent`` kwargs for the OpenAI Agents SDK in one
call::
agent = Agent(**client.openai_agent_config())
Sugar over the explicit form — ``agent_instructions`` (with
``doc_id`` targeting) as the instructions and
``as_openai_tools`` as the tools; local clients also carry their
configured ``retrieve_model`` (cloud omits ``model`` so the
framework default applies). To customize further, switch to
those methods directly.
Args:
doc_id: Document ID or list of IDs to target, as in
``agent_instructions``. Local: also enforced at the tool
layer, not just prompted. Cloud: prompt-level targeting
(tool scoping is server-side).
include_management (bool): Also expose tools that modify the
library.
model: Backend model name; overrides the local default.
"""
from .agent_tools import build_agent_instructions
scope = self._local_doc_scope(doc_id)
config: dict[str, Any] = {
"name": "PageIndex",
"instructions": build_agent_instructions(self, doc_id,
scoped=scope is not None),
"tools": self.as_openai_tools(include_management, doc_id=scope),
}
model = model or getattr(self, "retrieve_model", None)
if model:
config["model"] = model
return config
def as_anthropic_tools(self, include_management: bool = False,
asynchronous: bool = False,
doc_id: Optional[Union[str, list[str]]] = None,
) -> list:
"""
Runnable tools for the Anthropic SDK's tool runner — pass to
``client.beta.messages.tool_runner(tools=...)`` (or
``anthropic_runner_config()`` for the whole setup in one call).
The default flavor is for the sync ``Anthropic`` client; pass
``asynchronous=True`` for ``AsyncAnthropic``. For a manual
``messages.create`` loop, serialize with
``[tool.to_dict() for tool in ...]``.
Cloud: the full live read tool set (search, folders, images — as
enabled for your key), discovered from the PageIndex MCP server
and executed from your process; the server's input schemas pass
through verbatim (MCP and the Messages API share the schema
shape). The server-side alternative is the Messages API's beta
MCP connector — ``mcp_servers=[{"type": "url", "name":
"pageindex", "url": f"{BASE_URL}/mcp?tools=read",
"authorization_token": <your PageIndex API key>}]`` (drop
``?tools=read`` for the full tool set) — with no client-side
tools involved. Local: the in-process tools — the same set
``messages()`` runs internally.
Requires ``anthropic>=0.84.0``
(``pip install 'pageindex[anthropic]'``), imported only when this
method is called.
Args:
include_management (bool): Also expose tools that modify the
library. Local: adds ``remove_document``. Cloud: by default
only tools the server marks read-only are exposed; True
exposes the server's complete list (upload, delete, ...).
asynchronous (bool): Build ``beta_async_tool`` runnables for
``AsyncAnthropic`` (each tool call runs in a worker
thread, keeping blocking I/O off your event loop). The
sync and async runners each accept only their own flavor.
doc_id: Local only — restrict the tools to this document ID
(or list of IDs), enforced at the tool layer: out-of-scope
lookups return NOT_FOUND. Raises on cloud, where scoping
is server-side.
"""
from .integrations.anthropic_sdk import build_anthropic_tools
return build_anthropic_tools(self, include_management, asynchronous,
doc_ids=doc_id)
def anthropic_runner_config(
self,
model: str,
doc_id: Optional[Union[str, list[str]]] = None,
include_management: bool = False,
asynchronous: bool = False,
max_tokens: Optional[int] = None,
max_turns: Optional[int] = None,
) -> dict[str, Any]:
"""
Document QA ``tool_runner`` kwargs for the Anthropic SDK in one
call — only your ``messages`` remain::
runner = anthropic_client.beta.messages.tool_runner(
**client.anthropic_runner_config(model="claude-sonnet-4-5"),
messages=[{"role": "user", "content": "..."}],
)
Sugar over the explicit form — ``agent_instructions`` (with
``doc_id`` targeting) as the system prompt and
``as_anthropic_tools`` as the tools — plus the same defaults
``messages()`` applies: a per-model ``max_tokens`` and a
``max_iterations`` bound of 10. To customize further, switch to
those methods directly.
Args:
model: Backend model name (also resolves the ``max_tokens``
default).
doc_id: Document ID or list of IDs to target, as in
``agent_instructions``. Local: also enforced at the tool
layer, not just prompted. Cloud: prompt-level targeting
(tool scoping is server-side).
include_management (bool): Also expose tools that modify the
library.
asynchronous (bool): Build async runnables for
``AsyncAnthropic``.
max_tokens: Per-turn output budget; default resolved per
model.
max_turns: Agent-loop bound; default 10.
"""
from .agent_tools import build_agent_instructions
from .local_chat import _default_max_tokens
scope = self._local_doc_scope(doc_id)
return {
"model": model,
"max_tokens": (max_tokens if max_tokens is not None
else _default_max_tokens(model)),
"system": build_agent_instructions(self, doc_id,
scoped=scope is not None),
"tools": self.as_anthropic_tools(include_management, asynchronous,
doc_id=scope),
"max_iterations": max_turns if max_turns is not None else 10,
}
def as_claude_mcp(self, include_management: bool = False,
doc_id: Optional[Union[str, list[str]]] = None):
"""
``mcp_servers`` entry for the Claude Agent SDK.
Cloud: returns the remote PageIndex MCP config.
``include_management`` picks the endpoint, so the URL itself is
the gate — the default connects to the read-only endpoint
(``/mcp?tools=read``: the server registers only read-only tools),
``True`` connects to the full tool set. Local: returns an
in-process SDK MCP server exposing the agent tools, gated the
same way at registration (requires ``claude-agent-sdk``;
``pip install 'pageindex[claude]'``). ``doc_id`` (local only)
restricts those tools to that document ID (or list), enforced at
the tool layer; it raises on cloud, where scoping is server-side.
Cloud hosts that surface MCP server instructions receive the same
guidance ``agent_instructions()`` returns natively — passing both
duplicates the text (harmless). ``system_prompt`` stays the
recommended channel: it is guaranteed delivery, carries ``doc_id``
targeting, and is the only channel local mode has.
Usage (or ``claude_agent_config()`` for all three slots in one
call)::
options = ClaudeAgentOptions(
system_prompt=client.agent_instructions(),
mcp_servers={"pageindex": client.as_claude_mcp()},
# Pre-approval only — the server itself is already gated.
allowed_tools=["mcp__pageindex"],
)
"""
from .integrations.claude_agent_sdk import build_claude_mcp
return build_claude_mcp(self, include_management, doc_ids=doc_id)
def claude_agent_config(
self,
doc_id: Optional[Union[str, list[str]]] = None,
include_management: bool = False,
server_name: str = "pageindex",
) -> dict[str, Any]:
"""
Document QA ``ClaudeAgentOptions`` kwargs in one call::
options = ClaudeAgentOptions(**client.claude_agent_config())
Sugar over the explicit form — the managed system prompt
(``agent_instructions``) and the server entry (``as_claude_mcp``,
itself the tool gate) with its ``allowed_tools`` pre-approval,
one ``include_management`` and ``server_name`` applied
everywhere. To customize (your own system prompt, extra
servers), switch to those methods directly.
Args:
doc_id: Document ID or list of IDs to target, as in
``agent_instructions``. Local: also enforced at the tool
layer, not just prompted. Cloud: prompt-level targeting
(tool scoping is server-side).
include_management (bool): Also allow tools that modify the
library.
server_name (str): Key the server is registered under.
"""
from .agent_tools import build_agent_instructions
scope = self._local_doc_scope(doc_id)
return {
"system_prompt": build_agent_instructions(self, doc_id,
scoped=scope is not None),
"mcp_servers": {server_name: self.as_claude_mcp(
include_management, doc_id=scope)},
# Pre-approval only — the server itself is already gated (the
# read-only endpoint on cloud, the registered set locally).
"allowed_tools": [f"mcp__{server_name}"],
}
def agent_instructions(self, doc_id: Optional[Union[str, list[str]]] = None) -> str:
"""
Orchestration guidance for document QA agents — pass as the agent's
system prompt (or append to your own).
Cloud: the live instructions the PageIndex MCP server serves for
your key's tool set, fetched over the same session as
``agent_tools()`` — server-side guidance updates arrive without an
SDK release. Raises PageIndexAPIError if the server cannot be
reached. Local: the built-in guidance for the in-process tools.
With ``doc_id`` (str or list, same shape as ``chat_completions``),
appends the target documents' names and metadata and directs the
agent to work within them. Raises PageIndexAPIError if a doc_id does
not exist, or if its name is shadowed by a newer same-name document
(the name-addressed tools could not reach it).
"""
from .agent_tools import build_agent_instructions
return build_agent_instructions(self, doc_id)
# ---------- FOLDER MANAGEMENT ----------
def create_folder(
+3 -1
View File
@@ -55,7 +55,9 @@ class CloudAPI:
returned in get_tree/get_ocr responses and list_documents entries. Defaults to None.
Returns:
dict: {'doc_id': ...}
dict: {'doc_id': ...} — plus 'name', the stored document name
(a taken name gains a numeric suffix), when the server
returns it.
"""
data = {'if_retrieval': True}
if mode is not None:
+5
View File
@@ -0,0 +1,5 @@
"""Framework adapters for the agent tools layer.
These modules import their target frameworks lazily, at call time — the
frameworks are never required to install or import pageindex.
"""
+59
View File
@@ -0,0 +1,59 @@
"""Anthropic SDK adapter for the tool runner's tools=... slot.
Cloud clients get one runnable tool per live cloud MCP tool — the server's
input schemas pass through verbatim (MCP inputSchema and Messages API
input_schema are the same shape), calls proxied over MCP. Local clients get
the in-process tools — the same set messages() runs internally. Failed
calls raise ToolError so the runner emits the tool_result with
``is_error: true`` and the envelope as its content.
"""
from __future__ import annotations
import asyncio
from typing import Any
from ..errors import PageIndexAPIError
def build_anthropic_tools(client, include_management: bool = False,
asynchronous: bool = False, doc_ids=None) -> list:
try:
from anthropic import beta_async_tool, beta_tool
from anthropic.lib.tools import ToolError
except ImportError as exc:
raise PageIndexAPIError(
"as_anthropic_tools requires the Anthropic SDK tool runner "
"(anthropic>=0.84.0) — pip install -U anthropic (or pip install "
"'pageindex[anthropic]')."
) from exc
from ..agent_tools import _tool_specs
def wrap(name, description, schema, invoke):
"""One runnable tool in the caller's flavor: the sync runner and the
async runner each accept only their own kind, and the async variant
moves the blocking bridge/store call into a worker thread so it
never blocks the caller's event loop."""
def run(kwargs: dict) -> str:
text, is_error = invoke(kwargs)
if is_error:
raise ToolError(text)
return text
if asynchronous:
async def _afn(**kwargs: Any) -> str:
return await asyncio.to_thread(run, kwargs)
_afn.__name__ = name
return beta_async_tool(_afn, name=name, description=description,
input_schema=schema)
def _fn(**kwargs: Any) -> str:
return run(kwargs)
_fn.__name__ = name
return beta_tool(_fn, name=name, description=description,
input_schema=schema)
return [wrap(*spec)
for spec in _tool_specs(client, include_management,
doc_ids=doc_ids)]
@@ -0,0 +1,69 @@
"""Claude Agent SDK adapter: one value for the mcp_servers slot.
Cloud clients get the remote PageIndex MCP config — the framework connects
directly, and include_management picks the endpoint (the read-only
``?tools=read`` URL by default); local clients get an in-process SDK MCP
server over the same tool contract, gated the same way at registration.
"""
from __future__ import annotations
import asyncio
from typing import Any
from .._version import sdk_version
from ..errors import PageIndexAPIError
def build_claude_mcp(client, include_management: bool = False, doc_ids=None):
from ..agent_tools import _require_local_scope
# The cloud branch returns a URL config — reject cloud doc_ids so they
# are never silently dropped.
_require_local_scope(client, doc_ids)
if getattr(client, "api_key", None):
# include_management picks the endpoint — the URL itself is the
# gate (?tools=read serves only readOnlyHint-annotated tools).
suffix = "" if include_management else "?tools=read"
return {
"type": "http",
"url": f"{client.BASE_URL}/mcp{suffix}",
"headers": {"Authorization": f"Bearer {client.api_key}"},
}
try:
from claude_agent_sdk import create_sdk_mcp_server, tool
except ImportError as exc:
raise PageIndexAPIError(
"as_claude_mcp in local mode requires the Claude Agent SDK — "
"pip install claude-agent-sdk (or pip install 'pageindex[claude]')."
) from exc
from ..agent_tools import (TOOL_CONTRACT, _local_description,
_local_schema, call_tool, tool_names)
def make_handler(name: str):
async def handler(arguments: dict[str, Any]) -> dict[str, Any]:
text, is_error = await asyncio.to_thread(
call_tool, client, name, arguments or {}, doc_ids
)
result: dict[str, Any] = {"content": [{"type": "text", "text": text}]}
if is_error:
result["is_error"] = True
return result
return handler
def tool_kwargs(name: str) -> dict:
annotations = TOOL_CONTRACT[name].get("annotations")
if not annotations:
return {}
try:
from claude_agent_sdk import ToolAnnotations
except ImportError:
return {}
return {"annotations": ToolAnnotations(**annotations)}
tools = [
tool(name, _local_description(name),
_local_schema(name), **tool_kwargs(name))(make_handler(name))
for name in tool_names(include_management)
]
return create_sdk_mcp_server(name="pageindex", version=sdk_version(),
tools=tools)
+80
View File
@@ -0,0 +1,80 @@
"""OpenAI Agents SDK adapter for the Agent(tools=...) slot.
Cloud clients default to the live read tool set as plain FunctionTools via
the MCP bridge; pass hosted=True to use a single HostedMCPTool instead
(the model connects to the PageIndex cloud MCP server from OpenAI's side —
the read-only ``?tools=read`` endpoint by default). Local clients get the
in-process tools wrapped as FunctionTools. Tools are built as FunctionTool
directly so the contract/server JSON schema goes to the model verbatim —
function_tool() would regenerate it from a Python signature, dropping
items/enum/pattern/bounds and rejecting object-typed parameters.
"""
from __future__ import annotations
import asyncio
import json
from typing import Any
from ..errors import PageIndexAPIError
def build_openai_tools(client, include_management: bool = False,
hosted: bool = False, doc_ids=None) -> list:
try:
from agents import FunctionTool, HostedMCPTool
except ImportError as exc:
raise PageIndexAPIError(
"as_openai_tools requires the OpenAI Agents SDK — "
"pip install openai-agents (or pip install 'pageindex[openai]')."
) from exc
from ..agent_tools import (_dumps, _failure, _require_local_scope,
_tool_specs)
# The hosted branch returns before _tool_specs — reject cloud doc_ids
# here so they are never silently dropped.
_require_local_scope(client, doc_ids)
if getattr(client, "api_key", None) and hosted:
# include_management picks the endpoint — the URL itself is the
# gate (?tools=read serves only readOnlyHint-annotated tools), so
# nothing needs the Responses API approval flow.
suffix = "" if include_management else "?tools=read"
return [HostedMCPTool(tool_config={
"type": "mcp",
"server_label": "pageindex",
"server_url": f"{client.BASE_URL}/mcp{suffix}",
"headers": {"Authorization": f"Bearer {client.api_key}"},
"require_approval": "never",
})]
def wrap(name, description, schema, invoke):
async def on_invoke_tool(ctx: Any, args_json: str) -> str:
# strict_json_schema is off, so the provider never validates the
# payload; a malformed or non-object argument string must come
# back as the guided error envelope — raising here aborts the
# caller's whole run (hand-built FunctionTools have no
# failure_error_function to hand the error back to the model).
try:
parsed = json.loads(args_json) if args_json else {}
except ValueError:
parsed = None
if not isinstance(parsed, dict):
payload, _ = _failure(
f"Invalid arguments for {name}: expected a JSON object, "
f"got: {(args_json or '')[:200]!r}", None,
{"summary": "Malformed tool arguments",
"options": [f"Re-send the {name} call with a JSON "
"object of its parameters"]},
"INVALID_INPUT")
return _dumps(payload)
arguments = {key: value for key, value in parsed.items()
if value is not None}
text, _ = await asyncio.to_thread(invoke, arguments)
return text
return FunctionTool(name=name, description=description,
params_json_schema=schema,
on_invoke_tool=on_invoke_tool,
strict_json_schema=False)
return [wrap(*spec)
for spec in _tool_specs(client, include_management,
doc_ids=doc_ids)]
+26 -2
View File
@@ -97,6 +97,9 @@ class LocalAPI:
raise PageIndexAPIError(
"Failed to submit document: PDF has no content. All pages are blank."
)
# Fail before paying for indexing when _1.._99 are all taken; the
# binding name resolution happens again at save.
self._unique_doc_name(os.path.basename(file_path))
try:
if mode == "flash":
@@ -115,7 +118,7 @@ class LocalAPI:
doc_id = "pi-" + uuid.uuid4().hex
meta = {
"id": doc_id,
"name": os.path.basename(file_path),
"name": self._unique_doc_name(os.path.basename(file_path)),
"description": description,
"status": "completed",
"createdAt": _now_iso(),
@@ -129,7 +132,23 @@ class LocalAPI:
from .utils import remove_fields
self._store.save_document(
doc_id, meta, remove_fields(structure, fields=["text"]), pages)
return {"doc_id": doc_id}
return {"doc_id": doc_id, "name": meta["name"]}
def _unique_doc_name(self, name: str) -> str:
"""Mirror the cloud upload: a taken name gets _1.._99 appended,
beyond that the submit is rejected."""
taken = {meta.get("name") for meta in self._store.list_metas()}
if name not in taken:
return name
base, ext = os.path.splitext(name)
for num in range(1, 100):
candidate = f"{base}_{num}{ext}"
if candidate not in taken:
return candidate
raise PageIndexAPIError(
"Failed to submit document: Too many files with similar names. "
"Please use a different file name."
)
@staticmethod
def _extract_page_texts(file_path: str) -> list[str]:
@@ -190,6 +209,11 @@ class LocalAPI:
add_node_text(structure, pdf_pages)
return structure
def raw_tree(self, doc_id: str) -> list | None:
"""Stored tree verbatim — keeps start_index/end_index, which
get_tree's cloud wire shape renames and drops."""
return self._store.get_tree(doc_id)
def get_tree(self, doc_id: str, node_summary: bool = False,
include_text: bool = True) -> dict[str, Any]:
meta = self._require_doc(doc_id, "Failed to get tree result")
+840
View File
@@ -0,0 +1,840 @@
"""Managed local chat: document-QA agents over the local tools.
Three methods, three backend protocols, routed 1:1: ``chat_completions``
drives the backend's /chat/completions (any OpenAI-compatible backend,
final answer only), ``responses`` drives /responses (process items are
standard output; round-trip them for provider prompt-cache continuation and
agent memory), ``messages`` drives Anthropic's /v1/messages via the SDK's
own tool runner (tool_use/tool_result round-trip is the format's native
behavior).
Content passes through untouched — the caller's messages, the model's
answers, tool outputs. Native stop reasons pass through on ``messages``;
the OpenAI engine's abstraction does not surface per-turn finish reasons,
so ``chat_completions`` reports loop completion as ``"stop"``, while
``responses`` reports the backend's terminal ``status`` where the wire
surfaces one (recorded at the transport layer — the framework discards
it). The SDK owns gatekeeping (structural validation), table-setting
(managed instructions, tools, doc targeting), tool execution, and billing
(usage aggregation, envelope ids).
"""
from __future__ import annotations
import asyncio
import concurrent.futures
import hashlib
import json
import queue
import threading
import time
import uuid
from typing import Any, Iterator, Optional, Union
from .agent_tools import AGENT_INSTRUCTIONS, doc_targeting_block
from .errors import PageIndexAPIError
CHAT_HEADER = (
"You are PageIndex by Vectify AI, a document-focused assistant. "
"Be concise, never use emojis, and do not expose tool names."
)
# ── shared: prompt, doc targeting, validation, sync bridges ──
def _managed_instructions(extra_system: list[str]) -> str:
return "\n\n".join([CHAT_HEADER, AGENT_INSTRUCTIONS, *extra_system])
def _doc_block(client, doc_id) -> Optional[str]:
if doc_id is None:
return None
if not isinstance(doc_id, (str, list)):
raise PageIndexAPIError("doc_id must be a string or a list of "
"strings.")
doc_ids = [doc_id] if isinstance(doc_id, str) else list(doc_id)
missing = []
for one_id in doc_ids:
try:
client.get_document(one_id)
except PageIndexAPIError:
missing.append(str(one_id))
if missing:
raise PageIndexAPIError(
"Documents not found or access denied: " + ", ".join(missing)
)
# scoped: the chat surfaces also pass doc_id into the tool layer, so
# name resolution happens inside the allowlist — only a duplicate name
# within the targeted set shadows.
return doc_targeting_block(client, doc_id, scoped=True)
def _system_text(content: Any) -> str:
"""Text of a system/developer message: a string, or text parts joined."""
if isinstance(content, str):
return content
if isinstance(content, list):
texts = [part.get("text") for part in content
if isinstance(part, dict) and isinstance(part.get("text"), str)]
if texts:
return "\n".join(texts)
raise PageIndexAPIError(
"system message content must be a string or a list of text parts."
)
def _split_chat_messages(messages) -> "tuple[list[str], list[dict]]":
"""Validate the chat_completions surface's messages: system/developer
content joins the managed instructions; user/assistant history passes
through. Tool-history round-trips belong to responses()/messages()."""
if not isinstance(messages, list) or not messages:
raise PageIndexAPIError("messages must be a non-empty list.")
system_texts: list[str] = []
history: list[dict] = []
for message in messages:
if not isinstance(message, dict) or "role" not in message:
raise PageIndexAPIError(
"Each message must be a dict with 'role' and 'content'.")
role = message["role"]
if role in ("system", "developer"):
system_texts.append(_system_text(message.get("content")))
elif role in ("user", "assistant"):
content = message.get("content")
if not isinstance(content, str):
raise PageIndexAPIError(
"chat_completions content must be a string; for "
"structured items use responses() or messages()."
)
history.append({"role": role, "content": content})
else:
raise PageIndexAPIError(
f"Unsupported role for chat_completions: {role!r}. Tool "
"history round-trips belong to responses() or messages()."
)
if not history:
raise PageIndexAPIError("messages must contain a user or assistant "
"message.")
return system_texts, history
def _run_sync(coro):
try:
asyncio.get_running_loop()
except RuntimeError:
return asyncio.run(coro)
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as pool:
return pool.submit(asyncio.run, coro).result()
_SENTINEL = object()
def _stream_sync(agen_factory) -> Iterator[Any]:
"""Drive an async generator from a background thread; yield synchronously.
Closing the iterator cancels the run between items: the pump stops, and
the async generator's cleanup cancels the underlying agent task, so no
further model turns or tool executions start. An in-flight backend
request cannot be aborted mid-turn.
"""
items: "queue.Queue[Any]" = queue.Queue(maxsize=32)
cancelled = threading.Event()
def deliver(item) -> bool:
while not cancelled.is_set():
try:
items.put(item, timeout=0.1)
return True
except queue.Full:
continue
return False
def pump():
async def consume():
agen = agen_factory()
async def drain():
async for item in agen:
if not deliver(item):
break
# The watchdog lets cancellation land even while drain() is
# awaiting the backend — a plain async-for would only notice
# between items.
task = asyncio.ensure_future(drain())
try:
while not task.done():
if cancelled.is_set():
task.cancel()
break
await asyncio.sleep(0.05)
try:
await task
except asyncio.CancelledError:
pass
finally:
await agen.aclose()
try:
asyncio.run(consume())
except BaseException as exc: # re-raised on the consumer thread
deliver(exc)
return
deliver(_SENTINEL)
threading.Thread(target=pump, daemon=True).start()
try:
while True:
item = items.get()
if item is _SENTINEL:
return
if isinstance(item, BaseException):
raise item
yield item
finally:
cancelled.set()
# ── OpenAI engine (chat_completions / responses) ──
def _require_openai_agents(method: str) -> None:
try:
import agents # noqa: F401
except ImportError as exc:
raise PageIndexAPIError(
f"{method} in local mode requires the OpenAI Agents SDK — "
"pip install openai-agents (or pip install 'pageindex[openai]')."
) from exc
def _openai_model(protocol: str, model_name: str):
"""The backend protocol driver — the seam tests replace with a fake.
``litellm/<provider>/<model>`` (the client's normalized retrieve_model
form) and bare ``<provider>/<model>`` paths drive the provider through
LiteLLM — chat.completions only, so the responses protocol refuses them
instead of silently downgrading; a first segment LiteLLM does not know
(a HuggingFace repo id like ``Qwen/...``) is refused with the
``openai/`` escape instead of failing inside LiteLLM at request time;
an ``openai/`` prefix strips to the OpenAI SDK; bare names go to the
OpenAI SDK as-is."""
if "/" in model_name and not model_name.startswith("openai/"):
if protocol == "responses":
raise PageIndexAPIError(
f"responses() cannot drive "
f"'{model_name.removeprefix('litellm/')}': provider-prefixed "
"models route through LiteLLM, which speaks chat.completions, "
"not the Responses API. Use chat_completions() (or messages() "
"for Anthropic models), or point OPENAI_BASE_URL at a "
"Responses-capable backend and use a bare or "
"'openai/'-prefixed model name."
)
try:
from agents.extensions.models.litellm_model import LitellmModel
import litellm
except ImportError:
raise PageIndexAPIError(
f"'{model_name}' routes through LiteLLM, but litellm is not "
"installed. Run: pip install 'litellm>=1.30'"
)
wire = model_name.removeprefix("litellm/")
providers = getattr(litellm, "provider_list", None)
if providers and wire.split("/", 1)[0] not in providers:
raise PageIndexAPIError(
f"'{wire}' routes through LiteLLM, but "
f"'{wire.split('/', 1)[0]}' is not a LiteLLM provider. For an "
"OpenAI-compatible server (vLLM, TGI, Ollama) serving this "
f"model id, use 'openai/{wire}' and point OPENAI_BASE_URL "
"at the server."
)
return LitellmModel(wire)
import openai
model_name = model_name.removeprefix("openai/")
try:
backend = openai.AsyncOpenAI()
except openai.OpenAIError as exc:
raise PageIndexAPIError(
f"The OpenAI backend is not configured: {exc}") from exc
if protocol == "chat":
from agents.models.openai_chatcompletions import (
OpenAIChatCompletionsModel)
return OpenAIChatCompletionsModel(model_name, backend)
from agents.models.openai_responses import OpenAIResponsesModel
return OpenAIResponsesModel(model_name, openai_client=backend)
def _reported_model(model_name: str) -> str:
"""The name the provider actually serves — routing prefixes stripped."""
return model_name.removeprefix("litellm/").removeprefix("openai/")
def _openai_agent(client, protocol: str, model_name: str, instructions: str,
temperature, top_p, doc_ids=None):
from agents import Agent, ModelSettings
from .integrations.openai_agents import build_openai_tools
return Agent(
name="PageIndex",
instructions=instructions,
tools=build_openai_tools(client, doc_ids=doc_ids),
model=_openai_model(protocol, model_name),
model_settings=ModelSettings(temperature=temperature, top_p=top_p),
)
def _validate_max_turns(max_turns) -> None:
if max_turns is not None and (not isinstance(max_turns, int)
or max_turns < 1):
raise PageIndexAPIError("max_turns must be a positive integer.")
def _conversation_group_id(model_name: str, instructions: str, items) -> str:
"""Stable per-conversation cache-routing key: openai-agents hashes
RunConfig.group_id into the OpenAI prompt_cache_key, and without one it
stamps every run with a fresh key, tagging a round-tripped prefix as a
different cache group. Keyed on the prefix identity — model,
instructions, first conversation item — so a conversation's
continuations share one route without pooling unrelated conversations.
Callers pass the conversation's own items, never the SDK-prepended
doc-targeting block: that block is byte-identical for every
conversation about a document and would pool them all under one key."""
seed = json.dumps([model_name, instructions,
items[0] if items else None],
sort_keys=True, default=str)
return "pageindex-" + hashlib.sha256(seed.encode()).hexdigest()[:16]
def _run_kwargs(max_turns, group_id: str) -> dict:
# Managed runs never export traces — the caller opted into document QA,
# not telemetry.
from agents import RunConfig
kwargs: dict = {"run_config": RunConfig(tracing_disabled=True,
group_id=group_id)}
if max_turns is not None:
kwargs["max_turns"] = max_turns
return kwargs
def _record_response_status(agent, recorded: dict) -> None:
"""Capture each turn's terminal Response status at the transport client:
openai-agents' non-streaming path discards Response.status, so a final
turn truncated at the output cap would otherwise report as a clean
completion. No-op for backends without an OpenAI responses resource
(the streaming path records from lifecycle events instead)."""
responses = getattr(getattr(getattr(agent, "model", None), "_client", None),
"responses", None)
create = getattr(responses, "create", None)
if create is None:
return
async def recording_create(*args, **kwargs):
response = await create(*args, **kwargs)
if getattr(response, "status", None):
recorded["status"] = response.status
for field in ("incomplete_details", "error"):
value = getattr(response, field, None)
recorded[field] = (value.model_dump(mode="json")
if hasattr(value, "model_dump") else value)
return response
responses.create = recording_create
async def _aclose_backend(agent) -> None:
"""Close the per-call AsyncOpenAI client before its event loop ends —
otherwise httpx tears down pooled connections on a closed loop and
emits 'Task exception was never retrieved' noise."""
backend = getattr(getattr(agent, "model", None), "_client", None)
close = getattr(backend, "close", None)
if close is not None:
try:
await close()
except Exception:
pass
async def _run_closing(agent, coro):
try:
return await coro
finally:
await _aclose_backend(agent)
def _wrap_max_turns(max_turns) -> PageIndexAPIError:
limit = max_turns if max_turns is not None else "the default limit"
return PageIndexAPIError(
f"The agent did not finish within max_turns ({limit}). Raise "
"max_turns, or narrow the question."
)
def _openai_usage(raw_responses) -> dict:
prompt = sum(r.usage.input_tokens for r in raw_responses)
completion = sum(r.usage.output_tokens for r in raw_responses)
return {"prompt_tokens": prompt, "completion_tokens": completion,
"total_tokens": prompt + completion}
def run_chat_completions(client, messages, stream: bool = False,
doc_id=None, temperature: Optional[float] = None,
stream_metadata: bool = False,
enable_citations: bool = False,
model: Optional[str] = None,
max_turns: Optional[int] = None,
) -> Union[dict, Iterator[str], Iterator[dict]]:
if enable_citations:
raise PageIndexAPIError(
"enable_citations is cloud-only — citations need block-level OCR "
"data that local mode does not store."
)
_require_openai_agents("chat_completions")
_validate_max_turns(max_turns)
system_texts, history = _split_chat_messages(messages)
block = _doc_block(client, doc_id)
items = ([{"role": "user", "content": block}] if block else []) + history
model_name = model or client.retrieve_model
# litellm/ and openai/ are the SDK's routing markers, not model names —
# report the name the provider actually serves.
reported_model = _reported_model(model_name)
managed = _managed_instructions(system_texts)
agent = _openai_agent(client, "chat", model_name, managed,
temperature, None, doc_ids=doc_id)
run_kwargs = _run_kwargs(max_turns,
_conversation_group_id(model_name, managed,
history))
import openai
from agents import Runner
from agents.exceptions import AgentsException, MaxTurnsExceeded
if not stream:
try:
result = _run_sync(_run_closing(agent,
Runner.run(agent, input=items, **run_kwargs)))
except MaxTurnsExceeded as exc:
raise _wrap_max_turns(max_turns) from exc
except AgentsException as exc:
raise PageIndexAPIError(
f"The agent backend failed: {exc}") from exc
except openai.OpenAIError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
return {
"id": f"chatcmpl-{uuid.uuid4().hex}",
"object": "chat.completion",
"created": int(time.time()),
"model": reported_model,
"choices": [{
"index": 0,
"message": {"role": "assistant",
"content": result.final_output or ""},
"finish_reason": "stop",
}],
"usage": _openai_usage(result.raw_responses),
}
chat_id = f"chatcmpl-{uuid.uuid4().hex}"
created = int(time.time())
def chunk(delta: dict, finish=None) -> dict:
return {
"id": chat_id, "object": "chat.completion.chunk",
"created": created, "model": reported_model,
"choices": [{"index": 0, "delta": delta,
"finish_reason": finish}],
}
async def agen():
from openai.types.responses import ResponseTextDeltaEvent
streamed = Runner.run_streamed(agent, input=items, **run_kwargs)
completed = False
# First yield inside the try: a consumer that stops on the opening
# chunk must still tear the run down via the finally below.
try:
yield chunk({"role": "assistant", "content": ""})
async for event in streamed.stream_events():
if (event.type == "raw_response_event"
and isinstance(event.data, ResponseTextDeltaEvent)):
yield chunk({"content": event.data.delta})
completed = True
except MaxTurnsExceeded as exc:
raise _wrap_max_turns(max_turns) from exc
except AgentsException as exc:
raise PageIndexAPIError(
f"The agent backend failed: {exc}") from exc
except openai.OpenAIError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
finally:
if not completed and hasattr(streamed, "cancel"):
streamed.cancel() # abandoned/failed: stop the agent task
await _aclose_backend(agent)
yield chunk({}, finish="stop")
yield {
"id": chat_id, "object": "chat.completion.chunk",
"created": created, "model": reported_model, "choices": [],
"usage": _openai_usage(streamed.raw_responses),
}
if stream_metadata:
return _stream_sync(agen)
return (piece["choices"][0]["delta"]["content"]
for piece in _stream_sync(agen)
if piece.get("choices")
and "content" in piece["choices"][0]["delta"]
and piece["choices"][0]["delta"]["content"])
def run_responses(client, input, model: Optional[str] = None,
stream: bool = False, doc_id=None,
instructions: Optional[str] = None,
temperature: Optional[float] = None,
top_p: Optional[float] = None,
max_turns: Optional[int] = None,
) -> Union[dict, Iterator[dict]]:
_require_openai_agents("responses")
_validate_max_turns(max_turns)
if isinstance(input, str) and input.strip():
items = [{"role": "user", "content": input}]
elif (isinstance(input, list) and input
and all(isinstance(item, dict) for item in input)):
items = list(input)
else:
raise PageIndexAPIError("input must be a non-empty string or list "
"of item dicts.")
block = _doc_block(client, doc_id)
conversation = items
if block:
items = [{"role": "user", "content": block}] + items
extra = [instructions] if instructions else []
model_name = model or client.retrieve_model
managed = _managed_instructions(extra)
agent = _openai_agent(client, "responses", model_name, managed,
temperature, top_p, doc_ids=doc_id)
run_kwargs = _run_kwargs(max_turns,
_conversation_group_id(model_name, managed,
conversation))
recorded: dict = {}
import openai
from agents import Runner
from agents.exceptions import AgentsException, MaxTurnsExceeded
def envelope(output: list, raw_responses) -> dict:
usage = _openai_usage(raw_responses)
return {
"id": f"resp_{uuid.uuid4().hex}",
"object": "response",
"created_at": int(time.time()),
"model": _reported_model(model_name),
"status": recorded.get("status") or "completed",
"output": output,
"usage": {"input_tokens": usage["prompt_tokens"],
"output_tokens": usage["completion_tokens"],
"total_tokens": usage["total_tokens"]},
"instructions": managed,
"tools": [{"type": "function", "name": tool.name,
"description": tool.description,
"parameters": tool.params_json_schema,
"strict": getattr(tool, "strict_json_schema", True)}
for tool in agent.tools],
"tool_choice": "auto",
"parallel_tool_calls": True,
"temperature": temperature,
"top_p": top_p,
"max_output_tokens": None,
"error": recorded.get("error"),
"incomplete_details": recorded.get("incomplete_details"),
"metadata": None,
}
if not stream:
_record_response_status(agent, recorded)
try:
result = _run_sync(_run_closing(agent,
Runner.run(agent, input=[dict(item) for item in items],
**run_kwargs)))
except MaxTurnsExceeded as exc:
raise _wrap_max_turns(max_turns) from exc
except AgentsException as exc:
raise PageIndexAPIError(
f"The agent backend failed: {exc}") from exc
except openai.OpenAIError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
output = result.to_input_list()[len(items):]
return envelope(output, result.raw_responses)
# One logical response per call: per-turn backend lifecycle events
# (created/completed/...) are collapsed — forwarding them verbatim would
# end a canonical consumer at the first turn — and sequence numbers are
# reassigned monotonically across the whole run.
lifecycle = {"response.created", "response.in_progress",
"response.completed", "response.failed",
"response.incomplete", "response.queued"}
async def agen():
streamed = Runner.run_streamed(agent,
input=[dict(item) for item in items],
**run_kwargs)
sequence = 0
# output_index addresses an item's position in the logical
# response.output (the final envelope's list). Backend events
# carry per-turn indexes that restart at 0 each turn, so they are
# re-based by the count of items already committed by prior turns
# — and the tool outputs the SDK injects between turns take the
# next slot on that same axis.
output_offset = 0
completed = False
try:
async for event in streamed.stream_events():
if event.type == "raw_response_event":
data = event.data.model_dump(exclude_unset=True)
if data.get("type") in lifecycle:
if data["type"] in ("response.completed",
"response.incomplete",
"response.failed"):
# Per-turn terminal state; the last turn's wins
# and feeds the final envelope below.
state = data.get("response") or {}
for field in ("status", "incomplete_details",
"error"):
recorded[field] = state.get(field)
output_offset += len(state.get("output") or [])
continue
if isinstance(data.get("output_index"), int):
data["output_index"] += output_offset
sequence += 1
data["sequence_number"] = sequence
yield data
elif (event.type == "run_item_stream_event"
and event.item.type == "tool_call_output_item"):
# We are the tool executor, so we emit the output item
# the way the platform streams its own server-side tools.
sequence += 1
yield {"type": "response.output_item.done",
"output_index": output_offset,
"sequence_number": sequence,
"item": dict(event.item.to_input_item())}
output_offset += 1
completed = True
except MaxTurnsExceeded as exc:
raise _wrap_max_turns(max_turns) from exc
except AgentsException as exc:
if recorded.get("status") not in ("failed", "incomplete"):
raise PageIndexAPIError(
f"The agent backend failed: {exc}") from exc
# response.failed / response.incomplete: the engine re-raises
# the backend's terminal state as an exception — it is a
# protocol event, emitted as the terminal event below.
completed = True
except openai.OpenAIError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
finally:
if not completed and hasattr(streamed, "cancel"):
streamed.cancel() # abandoned/failed: stop the agent task
await _aclose_backend(agent)
output = streamed.to_input_list()[len(items):]
sequence += 1
status = recorded.get("status") or "completed"
terminal = {"incomplete": "response.incomplete",
"failed": "response.failed"}.get(status,
"response.completed")
yield {"type": terminal, "sequence_number": sequence,
"response": envelope(output, streamed.raw_responses)}
return _stream_sync(agen)
# ── Anthropic engine (messages) ──
def _require_anthropic() -> None:
try:
import anthropic # noqa: F401
except ImportError as exc:
raise PageIndexAPIError(
"messages in local mode requires the Anthropic SDK — "
"pip install anthropic (or pip install 'pageindex[anthropic]')."
) from exc
try:
from anthropic import beta_tool # noqa: F401
from anthropic.lib.tools import ToolError # noqa: F401
except ImportError as exc:
raise PageIndexAPIError(
"messages in local mode requires anthropic >= 0.84.0 (the tool "
"runner with ToolError) — pip install -U anthropic."
) from exc
def _anthropic_client():
"""The backend client — the seam tests replace with a fake transport."""
import anthropic
return anthropic.Anthropic()
def _anthropic_system(extra_system, block: Optional[str]) -> list[dict]:
"""System blocks: cache_control marks the stable managed prefix only
(the API allows 4 breakpoints total — the varying doc block and caller
blocks must not consume the budget); the doc block and caller system
content follow as their own blocks."""
blocks = [{"type": "text",
"text": CHAT_HEADER + "\n\n" + AGENT_INSTRUCTIONS,
"cache_control": {"type": "ephemeral"}}]
if block:
blocks.append({"type": "text", "text": block})
if extra_system is None:
return blocks
if isinstance(extra_system, str):
if extra_system.strip():
blocks.append({"type": "text", "text": extra_system})
return blocks
if isinstance(extra_system, list):
return blocks + list(extra_system)
raise PageIndexAPIError("system must be a string or a list of blocks.")
def _dump_block(block) -> Any:
"""A content block as a plain JSON dict, minus SDK-internal fields the
API rejects (ParsedBetaTextBlock.__api_exclude__, e.g. parsed_output)."""
if hasattr(block, "model_dump"):
exclude = getattr(type(block), "__api_exclude__", None)
return block.model_dump(mode="json",
exclude=set(exclude) if exclude else None)
return block
def _dump_message(message) -> dict:
message = dict(message)
content = message.get("content")
if isinstance(content, list):
message["content"] = [_dump_block(item) for item in content]
return message
def _anthropic_usage(turns, final_usage: dict) -> dict:
"""The final turn's native usage dict with the token counters replaced
by cross-turn sums (None-safe); all other native fields survive."""
totals = dict(final_usage)
for field in ("input_tokens", "output_tokens",
"cache_creation_input_tokens", "cache_read_input_tokens"):
values = [getattr(turn.usage, field, None) for turn in turns]
counted = [value for value in values if isinstance(value, int)]
if counted:
totals[field] = sum(counted)
return totals
_CLAUDE_4096_MODELS = ("claude-3-opus", "claude-3-sonnet", "claude-3-haiku",
"claude-3-5-sonnet-20240620")
def _default_max_tokens(model: str) -> int:
"""The wire-required per-turn budget when the caller sets none: 8192,
except the claude-3 generation whose output ceiling is 4096."""
return 4096 if model.startswith(_CLAUDE_4096_MODELS) else 8192
def run_messages(client, messages, model: str,
max_tokens: Optional[int] = None,
stream: bool = False, doc_id=None, system=None,
temperature: Optional[float] = None,
top_p: Optional[float] = None,
top_k: Optional[int] = None,
stop_sequences: Optional[list[str]] = None,
max_turns: Optional[int] = None,
) -> Union[dict, Iterator[Any]]:
from .integrations.anthropic_sdk import build_anthropic_tools
_require_anthropic()
import anthropic
_validate_max_turns(max_turns)
if isinstance(messages, str) and messages.strip():
messages = [{"role": "user", "content": messages}]
if (not isinstance(messages, list) or not messages
or not all(isinstance(message, dict) for message in messages)):
raise PageIndexAPIError("messages must be a non-empty string or a "
"list of message dicts.")
block = _doc_block(client, doc_id)
prepared = [dict(message) for message in messages]
passthrough = {key: value for key, value in {
"temperature": temperature, "top_p": top_p, "top_k": top_k,
"stop_sequences": stop_sequences,
}.items() if value is not None}
runner = _anthropic_client().beta.messages.tool_runner(
max_tokens=(max_tokens if max_tokens is not None
else _default_max_tokens(model)),
messages=prepared,
model=model,
tools=build_anthropic_tools(client, doc_ids=doc_id),
system=_anthropic_system(system, block),
stream=stream,
# Bounded like the OpenAI surfaces (their framework default is 10).
max_iterations=max_turns if max_turns is not None else 10,
**passthrough,
)
if stream:
def events() -> Iterator[Any]:
try:
for turn_stream in runner:
for event in turn_stream:
yield event
except anthropic.AnthropicError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
return events()
try:
turns = [turn for turn in runner]
except anthropic.AnthropicError as exc:
raise PageIndexAPIError(
f"The model backend failed: {exc}") from exc
if not turns:
raise PageIndexAPIError("The model returned no response.")
captured: dict = {}
def capture(params):
captured.update(params)
return params
runner.set_messages_params(capture)
if not captured.get("messages"):
# The conversation is read back through a mutator; if a vendor
# change stops it delivering params, the envelope would silently
# lose the tool turns — fail loudly instead.
raise PageIndexAPIError(
"Could not read the conversation back from the anthropic tool "
"runner — the installed anthropic version is incompatible with "
"this pageindex release."
)
conversation = list(captured["messages"])
final = turns[-1]
envelope = final.model_dump(mode="json")
envelope["content"] = [_dump_block(item) for item in final.content]
envelope["usage"] = _anthropic_usage(turns, envelope.get("usage") or {})
# The full turn sequence (assistant tool_use + user tool_result + final),
# valid for verbatim append to the caller's history. The runner appends
# a turn to its params only when it executed tools from it — content
# carried tool_use blocks and the turn was not a refusal. stop_reason
# alone cannot tell: a max_tokens turn with complete tool_use blocks
# still executes. Whether final's tool_use ids already sit in the
# history is the ground truth for "already appended".
new_messages = [_dump_message(message)
for message in conversation[len(prepared):]]
final_blocks = [_dump_block(item) for item in final.content]
final_ids = {block["id"] for block in final_blocks
if block.get("type") == "tool_use"}
history_ids = {block.get("id")
for message in new_messages
if (message.get("role") == "assistant"
and isinstance(message.get("content"), list))
for block in message["content"]
if (isinstance(block, dict)
and block.get("type") == "tool_use")}
if not final_ids or not final_ids <= history_ids:
# Unexecuted tool_use blocks (refusal turns) have no tool_result,
# so they cannot enter an appendable history — strip them, as the
# SDK itself does when it rebuilds params around such a turn.
appendable = [block for block in final_blocks
if block.get("type") != "tool_use"]
if appendable:
new_messages = new_messages + [
{"role": "assistant", "content": appendable}]
envelope["messages"] = new_messages
return envelope
+210
View File
@@ -0,0 +1,210 @@
"""Minimal MCP client (streamable HTTP) for the PageIndex cloud MCP server.
Backs the cloud branches of ``client.agent_tools()`` and
``client.agent_instructions()``: ``tools/list`` discovers the live tool set,
``tools/call`` executes a tool, and the ``initialize`` handshake carries the
server's agent instructions. Synchronous, requests-only.
Works against both stateful and stateless servers: a session id returned by
``initialize`` is echoed back, and a session-carrying request rejected with
HTTP 404 (the spec's expired-session status) re-initializes once and
retries; a 400 is an ordinary bad request and is never replayed.
"""
from __future__ import annotations
import json
import threading
from typing import Any, Optional
import requests
from ._version import sdk_version
from .errors import PageIndexAPIError
_PROTOCOL_VERSION = "2025-06-18"
_TIMEOUT = (10, 240) # tools may wait server-side (wait_for_completion: 3 min)
def _parse_sse(text: str) -> list[dict]:
"""JSON-RPC messages out of a text/event-stream body."""
messages = []
text = text.replace("\r\n", "\n").replace("\r", "\n")
for block in text.split("\n\n"):
data_lines = [line[5:].removeprefix(" ") for line in block.splitlines()
if line.startswith("data:")]
if not data_lines:
continue
try:
messages.append(json.loads("\n".join(data_lines)))
except ValueError:
continue
return messages
class McpBridge:
def __init__(self, url: str, headers: dict[str, str]):
self._url = url
self._auth_headers = dict(headers)
self._session_id: Optional[str] = None
self._protocol_version: Optional[str] = None
self._instructions: Optional[str] = None
self._initialized = False
self._lock = threading.RLock()
self._next_id = 0
# ── JSON-RPC over streamable HTTP ──
def _post(self, payload: dict, session_id: Optional[str] = None,
protocol_version: Optional[str] = None) -> requests.Response:
headers = {
"Content-Type": "application/json",
"Accept": "application/json, text/event-stream",
**self._auth_headers,
}
if session_id:
headers["Mcp-Session-Id"] = session_id
if protocol_version:
headers["MCP-Protocol-Version"] = protocol_version
try:
return requests.post(self._url, json=payload, headers=headers,
timeout=_TIMEOUT)
except requests.RequestException as exc:
raise PageIndexAPIError(
f"Could not reach the PageIndex MCP server: {exc}"
) from exc
def _extract_result(self, response: requests.Response, request_id: int) -> Any:
content_type = response.headers.get("Content-Type", "")
if "text/event-stream" in content_type:
# SSE is UTF-8 by spec; requests guesses latin-1 for charset-less
# text/* and would mojibake every non-ASCII character.
messages = _parse_sse(response.content.decode("utf-8",
errors="replace"))
else:
try:
messages = [response.json()]
except ValueError as exc:
raise PageIndexAPIError(
f"MCP server returned a non-JSON response "
f"(HTTP {response.status_code})."
) from exc
# Strict id correlation only — accepting any result-bearing message
# would return a stale or mis-correlated reply as this call's.
reply = next((m for m in messages if m.get("id") == request_id), None)
if reply is None:
raise PageIndexAPIError(
"MCP server response contained no reply matching the request."
)
if "error" in reply:
error = reply["error"] or {}
raise PageIndexAPIError(
f"MCP error {error.get('code')}: {error.get('message')}"
)
return reply.get("result")
def _request(self, method: str, params: Optional[dict] = None,
_retry: bool = True) -> Any:
self._ensure_initialized()
with self._lock:
self._next_id += 1
request_id = self._next_id
session_id = self._session_id
protocol_version = self._protocol_version
payload: dict[str, Any] = {"jsonrpc": "2.0", "id": request_id,
"method": method}
if params is not None:
payload["params"] = params
response = self._post(payload, session_id, protocol_version)
if response.status_code == 404 and session_id and _retry:
# Session expired (stateful servers; the spec's 404): the server
# refused the request at session validation, so replaying it is
# safe. 400 is an ordinary bad request — replaying one would
# re-run side effects. Reset only if no other thread has already
# re-initialized, then retry once on the fresh session.
with self._lock:
if self._session_id == session_id:
self._initialized = False
self._session_id = None
self._protocol_version = None
return self._request(method, params, _retry=False)
if response.status_code >= 400:
raise PageIndexAPIError(
f"MCP request failed: HTTP {response.status_code} "
f"({response.text[:200]})"
)
return self._extract_result(response, request_id)
def _ensure_initialized(self) -> None:
with self._lock:
if self._initialized:
return
self._next_id += 1
request_id = self._next_id
response = self._post({
"jsonrpc": "2.0", "id": request_id, "method": "initialize",
"params": {
"protocolVersion": _PROTOCOL_VERSION,
"capabilities": {},
"clientInfo": {"name": "pageindex-python-sdk",
"version": sdk_version()},
},
})
if response.status_code >= 400:
raise PageIndexAPIError(
f"Could not connect to the PageIndex MCP server: HTTP "
f"{response.status_code} ({response.text[:200]}). Check "
"your API key."
)
result = self._extract_result(response, request_id) or {}
self._session_id = response.headers.get("Mcp-Session-Id")
self._protocol_version = result.get("protocolVersion",
_PROTOCOL_VERSION)
self._instructions = result.get("instructions")
self._initialized = True
# Sent inside the lock so no concurrent thread can slip a
# request between the handshake and this notification.
try:
self._post({"jsonrpc": "2.0",
"method": "notifications/initialized"},
self._session_id, self._protocol_version)
except PageIndexAPIError:
pass # advisory; a server that required it fails the next request
# ── public surface ──
def instructions(self) -> Optional[str]:
"""The server's agent instructions from the initialize handshake."""
self._ensure_initialized()
return self._instructions
def list_tools(self) -> list[dict]:
tools: list[dict] = []
cursor: Optional[str] = None
while True:
params = {"cursor": cursor} if cursor else {}
result = self._request("tools/list", params) or {}
tools.extend(result.get("tools") or [])
cursor = result.get("nextCursor")
if not cursor:
return tools
def call_tool(self, name: str, arguments: dict[str, Any]) -> "tuple[str, bool]":
"""Returns (text, is_error) — is_error is the server's MCP isError
marking, which callers must carry to their framework's own error
channel."""
result = self._request("tools/call",
{"name": name, "arguments": arguments}) or {}
is_error = bool(result.get("isError"))
texts = []
for block in result.get("content") or []:
if isinstance(block, dict) and block.get("type") == "text":
texts.append(block.get("text", ""))
elif isinstance(block, dict) and isinstance(block.get("data"), str):
# Base64 payloads (image/audio) become a metadata stub —
# dumped verbatim they hand the model the raw blob. Revisit
# if tool results ever pass through as real multimodal input.
kind = block.get("mimeType") or block.get("type") or "binary"
size_kb = max(1, len(block["data"]) * 3 // 4096)
texts.append(f"[{kind} content omitted: ~{size_kb} KB]")
else:
texts.append(json.dumps(block, ensure_ascii=False))
return "\n".join(texts), is_error
+12 -1
View File
@@ -1,6 +1,6 @@
[tool.poetry]
name = "pageindex"
version = "0.2.9"
version = "0.2.10"
description = "Python SDK for PageIndex — reasoning-based, vectorless document retrieval, cloud and local"
readme = "README.md"
license = "MIT"
@@ -38,6 +38,17 @@ sortedcontainers = ">=2.4.0"
regex = ">=2024.0.0"
python-dotenv = ">=1.0.0"
pyyaml = ">=6.0"
# Older releases break string prompts with SDK MCP servers (#597, #780).
claude-agent-sdk = { version = ">=0.1.53", optional = true }
# Older releases crash on current openai before the request is sent.
openai-agents = { version = ">=0.18.1", optional = true }
# Older releases execute a refusal turn's tool_use blocks.
anthropic = { version = ">=0.108.0", optional = true }
[tool.poetry.extras]
claude = ["claude-agent-sdk"]
openai = ["openai-agents"]
anthropic = ["anthropic"]
[tool.poetry.group.dev.dependencies]
pytest = ">=7.0"
+215
View File
@@ -0,0 +1,215 @@
{
"_provenance": "Frozen copy of the PageIndex cloud MCP server's tool contract (names, input schemas, descriptions, and annotations as served via tools/list). The parity test asserts pageindex.agent_tools.TOOL_CONTRACT matches this file; update both together only when the cloud contract changes.",
"tools": {
"browse_documents": {
"annotations": {
"readOnlyHint": true,
"openWorldHint": false
},
"description": "Primary document retrieval tool. After orienting with get_folder_structure() (when available), use this for all document-related questions. The bare call returns root-level sub-folders and documents; pass folder_id to drill into a sub-folder level by level. Use sort=\"relevance\" + query for semantic ranking. Do NOT jump to search_documents() first — it is an escalation path, only after browse_documents(sort=\"relevance\") has failed.",
"schema": {
"type": "object",
"properties": {
"folder_id": {
"type": "string",
"default": "root",
"description": "Folder scope (default \"root\"). Pass a specific folder ID to scope into that folder, or \"root\" to reference the library root. The read-only \"shared-with-me\" and \"following\" folders live at the library root — pass one of those ids to browse them. Copy any folder_id verbatim from a browse/tree response, never construct one. Combine with `recursive` to control breadth."
},
"recursive": {
"type": "boolean",
"default": false,
"description": "Whether to include documents from descendant folders. When false (default), returns the direct contents of folder_id along with its sub-folders — prefer this for level-by-level exploration so you retain folder hierarchy context. When true, flattens all descendant documents into one list and omits sub-folders — use only when a non-recursive browse of the target folder returned no relevant results and you need to widen the scope, or the user explicitly requests a flat listing."
},
"sort": {
"type": "string",
"enum": [
"time",
"relevance"
],
"default": "time",
"description": "Sort order. \"time\" (default) sorts by upload date (newest first); \"relevance\" orders documents by semantic relevance to `query`. Relevance also works inside the read-only shared folders — pass their folder_id — but at the library root it ranks only your own documents."
},
"query": {
"type": "string",
"description": "Search query for relevance ranking. Required when sort=\"relevance\"; must be omitted when sort=\"time\"."
},
"offset": {
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991,
"default": 0,
"description": "Zero-based pagination offset. Pass the value of `next_offset` from the previous response to fetch the next page."
},
"limit": {
"type": "number",
"minimum": 1,
"maximum": 50,
"default": 10,
"description": "Number of documents to return per page (1-50, default 10)"
}
},
"required": []
}
},
"get_document": {
"annotations": {
"readOnlyHint": true,
"openWorldHint": false
},
"description": "Check a document's processing status and metadata. `status` is one of \"pending\", \"queued\", \"processing\", \"completed\", or \"failed\" — call this before `get_document_structure()` or `get_page_content()` to confirm the document is ready.",
"schema": {
"type": "object",
"properties": {
"doc_name": {
"type": "string",
"minLength": 1,
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
},
"folder_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
},
"wait_for_completion": {
"type": "boolean",
"default": false,
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
}
},
"required": [
"doc_name"
]
}
},
"get_document_structure": {
"annotations": {
"readOnlyHint": true,
"openWorldHint": false
},
"description": "Extract a document's hierarchical outline (headers, sections, page references). REQUIRED for documents over 20 pages — call this first to locate relevant sections, then pass their page numbers to `get_page_content()`. Use the `part` parameter to iterate large outlines until `pagination.has_more` is false.",
"schema": {
"type": "object",
"properties": {
"doc_name": {
"type": "string",
"minLength": 1,
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
},
"folder_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
},
"part": {
"type": "integer",
"minimum": 1,
"maximum": 9007199254740991,
"default": 1,
"description": "Part number for pagination (1-based, default 1). For large outlines, increment until the response's `pagination.has_more` becomes false."
},
"wait_for_completion": {
"type": "boolean",
"default": false,
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
}
},
"required": [
"doc_name"
]
}
},
"get_page_content": {
"annotations": {
"readOnlyHint": true,
"openWorldHint": false
},
"description": "Extract page content from a processed document. Use tight, targeted page ranges — never the whole document at once. For documents over 20 pages, call `get_document_structure()` first to pick relevant sections. Embedded image paths in the response feed into `get_document_image()`.",
"schema": {
"type": "object",
"properties": {
"doc_name": {
"type": "string",
"minLength": 1,
"description": "Copy the `name` field verbatim from a browse_documents() or search_documents() response (case-sensitive, include extension). Example: \"Q3 Report.pdf\". If the response shows two documents with the same name, pass `folder_id` alongside to disambiguate."
},
"folder_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
},
"pages": {
"type": "string",
"minLength": 1,
"pattern": "^(\\d+(-\\d+)?)(,\\s*\\d+(-\\d+)?)*$",
"description": "Page specification: \"5\", \"3,7,10\", \"5-10\", or \"1-3,7,9-12\""
},
"wait_for_completion": {
"type": "boolean",
"default": false,
"description": "If true and document is processing, automatically wait up to 3 minutes until completed. Reduces repeated tool calls."
}
},
"required": [
"doc_name",
"pages"
]
}
},
"remove_document": {
"annotations": {
"readOnlyHint": false,
"destructiveHint": true,
"idempotentHint": true,
"openWorldHint": false
},
"description": "Permanently delete documents and all associated data. Only invoke when the user explicitly names the documents AND confirms deletion. Returns `results` — one entry per requested document: `{ doc_name, status: \"deleted\" | \"not_found\" | \"failed\", error? }`. Inspect each entry for per-document failures. This action is irreversible.",
"schema": {
"type": "object",
"properties": {
"doc_names": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"minItems": 1,
"maxItems": 10,
"description": "Array of document names to delete. Each name must be copied verbatim from the `name` field of a browse_documents() or search_documents() response (case-sensitive, include extension). Example: [\"Q3 Report.pdf\", \"draft.pdf\"]. Max 10 per call."
},
"folder_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"description": "Disambiguator for same-name documents. Copy the `folder_id` from the intended browse/search result; use \"root\" for root-level documents, or \"shared-with-me\"/\"following\" for the read-only folders at the library root; omit if `doc_name` is unique. Copy any folder_id verbatim from a browse_documents()/get_folder_structure() response, never construct one."
}
},
"required": [
"doc_names"
]
}
}
}
}
File diff suppressed because it is too large Load Diff
+79 -3
View File
@@ -151,6 +151,16 @@ def test_get_page_content(local_client, indexed_doc):
local_client.get_page_content(indexed_doc, "abc")
def test_get_page_content_span_bomb_rejected(local_client, indexed_doc):
"""An absurd range must be rejected arithmetically, not expanded into
a billion integers in the caller's process (the tool layer already
refused; the public client method did not)."""
with pytest.raises(ValueError, match="spans more than 10000"):
local_client.get_page_content(indexed_doc, "1-1000001")
# At the bound itself the spec still parses.
assert local_client.get_page_content(indexed_doc, "5-10004") == []
def test_submit_does_not_create_cwd_logs(local_client, sample_pdf, tmp_path, monkeypatch):
monkeypatch.chdir(tmp_path)
def fake_page_index_main(doc, opt=None, logger=None, page_list=None):
@@ -162,6 +172,49 @@ def test_submit_does_not_create_cwd_logs(local_client, sample_pdf, tmp_path, mon
assert not (tmp_path / "logs").exists()
def test_submit_duplicate_name_gets_suffix(local_client, sample_pdf, monkeypatch):
"""Mirror the cloud upload: a second submit of the same file name is
stored as name_1, not as a same-name duplicate."""
def fake_page_index_main(doc, opt=None, logger=None, page_list=None):
return {"doc_name": "sample.pdf", "doc_description": "d",
"structure": json.loads(json.dumps(STRUCTURE))}
monkeypatch.setattr(page_index_module, "page_index_main", fake_page_index_main)
first = local_client.submit_document(sample_pdf)
assert first["name"] == "sample.pdf"
with pytest.warns(UserWarning, match='stored as "sample_1.pdf"'):
second = local_client.submit_document(sample_pdf)
assert second["name"] == "sample_1.pdf"
names = {d["id"]: d["name"]
for d in local_client.list_documents()["documents"]}
assert names[first["doc_id"]] == "sample.pdf"
assert names[second["doc_id"]] == "sample_1.pdf"
def test_submit_duplicate_name_exhaustion(local_client, monkeypatch):
api = local_client._api
metas = ([{"name": "x.pdf"}]
+ [{"name": f"x_{num}.pdf"} for num in range(1, 100)])
monkeypatch.setattr(api._store, "list_metas", lambda: metas)
with pytest.raises(PageIndexAPIError, match="Too many files"):
api._unique_doc_name("x.pdf")
def test_submit_name_exhaustion_rejects_before_indexing(
local_client, sample_pdf, monkeypatch,
):
api = local_client._api
metas = ([{"name": "sample.pdf"}]
+ [{"name": f"sample_{num}.pdf"} for num in range(1, 100)])
monkeypatch.setattr(api._store, "list_metas", lambda: metas)
monkeypatch.setattr(
page_index_module, "page_index_main",
lambda *args, **kwargs: pytest.fail(
"indexer ran despite name exhaustion"),
)
with pytest.raises(PageIndexAPIError, match="Too many files"):
local_client.submit_document(sample_pdf)
def test_submit_flash(local_client, sample_pdf, monkeypatch):
calls = {}
def fake_flash(pdf, summary=True, summary_model=None, **kwargs):
@@ -432,7 +485,8 @@ def test_torn_delete_never_lists_ghost(local_client, indexed_doc, tmp_path):
def test_corrupt_doc_json_is_contained(local_client, indexed_doc, sample_pdf, tmp_path):
second = local_client.submit_document(sample_pdf)["doc_id"]
with pytest.warns(UserWarning): # same-name resubmit → stored as sample_1.pdf
second = local_client.submit_document(sample_pdf)["doc_id"]
(tmp_path / "store" / "docs" / indexed_doc / "doc.json").write_text("{truncated")
# manifest still holds a good copy of the meta — served consistently
@@ -599,8 +653,12 @@ def test_retrieval_endpoints_cloud_only(local_client):
local_client.get_retrieval("any")
def test_chat_completions_cloud_only(local_client):
with pytest.raises(PageIndexAPIError, match="not yet supported in local mode"):
def test_chat_completions_local_needs_agents_extra(local_client, monkeypatch):
"""Local chat is implemented (see test_local_chat.py); without the
openai-agents extra it raises the actionable install error."""
import sys
monkeypatch.setitem(sys.modules, "agents", None)
with pytest.raises(PageIndexAPIError, match="pageindex\\[openai\\]"):
local_client.chat_completions(
messages=[{"role": "user", "content": "q"}])
@@ -710,3 +768,21 @@ def test_cloud_chat_stream_parsing(cloud, monkeypatch):
messages=[{"role": "user", "content": "q"}], stream=True,
stream_metadata=True))
assert {"object": "chat.completion.citations", "citations": []} in chunks
def test_cloud_chat_accepts_query_string(cloud):
client, calls, fake = cloud
fake.payload = {"choices": [{"message": {"content": "ok"}}]}
client.chat_completions("What status?")
assert calls[-1]["json"]["messages"] == [
{"role": "user", "content": "What status?"}]
with pytest.raises(PageIndexAPIError, match="non-empty string"):
client.chat_completions(" ")
def test_parse_pages_overlap_counts_union():
from pageindex.client import _parse_pages
pages = _parse_pages("1-5000,2000-9000")
assert len(pages) == 9000 and pages[0] == 1 and pages[-1] == 9000
with pytest.raises(ValueError, match="spans more than"):
_parse_pages("1-10001")
File diff suppressed because it is too large Load Diff
+41
View File
@@ -60,3 +60,44 @@ def test_import_pageindex_is_lazy():
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.split() == ["clean", "function"]
def test_sdk_submodules_reachable_and_dunder_probes_stay_lazy():
"""The 0.2.10 modules resolve as attributes, and underscore probes (the
frequent unknown names: copy/pickle/inspect dunders) raise without
dragging in the indexing stack. A non-underscore unknown name still
raises AttributeError — after the compat fallthrough's one classic
import, which is the pre-0.2.10 behavior."""
probe = (
"import sys, pageindex\n"
"pageindex.agent_tools; pageindex.local_chat\n"
"pageindex.mcp_bridge; pageindex.integrations\n"
"assert not hasattr(pageindex, '__wrapped__')\n"
"heavy = [m for m in ('pageindex.page_index_classic', "
"'pageindex.flash', 'pageindex.utils') if m in sys.modules]\n"
"print(','.join(heavy) or 'clean')\n"
"try:\n"
" pageindex.definitely_missing\n"
" raise SystemExit('no AttributeError')\n"
"except AttributeError:\n"
" pass\n"
)
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "clean"
def test_classic_compat_surface_still_reachable():
"""The pre-0.2.10 catch-all made every classic/utils public name a
package attribute; dropping it broke `from pageindex import
ConfigLoader` on upgrade with no deprecation path."""
probe = (
"import pageindex\n"
"assert callable(pageindex.count_tokens)\n"
"assert isinstance(pageindex.ConfigLoader, type)\n"
"from pageindex import check_toc # noqa: F401\n"
"print('ok')\n"
)
out = subprocess.run([sys.executable, "-c", probe],
capture_output=True, text=True, check=True)
assert out.stdout.strip() == "ok"