mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-10-02 07:44:37 +08:00
bc1c1740ffb6d1b0222b69dda8d30216a72abec5
365
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bc1c1740ff |
feat: model knobs, LiteLLM-verbatim chat lane, per-door passthrough params (#409)
* docs: drop the demo's install step — openai-agents ships with the SDK now
* feat: the chat lane routes every model through LiteLLM — bare names included
The direct-OpenAI special case existed to dodge LiteLLM's import cost,
and it made OpenAI's own Responses-first models fail on the front door:
gpt-5.6-sol 400s on chatcmpl+tools while reasoning is on (server-side
policy — wire-captured with no reasoning_effort in our request).
LiteLLM 1.97 translates such calls onto /v1/responses; 1.84 does not,
so the sol-class 400 now carries its two exits (upgrade litellm /
responses()).
Routing after the flip: chat protocol — bare names are OpenAI-compatible
shorthand (wire form openai/<name>; OPENAI_API_KEY / OPENAI_BASE_URL
still select the backend, and the missing key stays a build-time
failure), litellm/ strips, openai/ opts out to the OpenAI SDK directly;
responses protocol unchanged (OpenAI-SDK native, LiteLLM refused).
The import cost is handled instead of dodged: local clients preload
litellm on a background thread (first call then perceives 0.0s), and
pageindex sets LITELLM_LOCAL_MODEL_COST_MAP=True via setdefault —
LiteLLM's import otherwise blocks on a network fetch of its price map
(fresh venv: 5.6s -> 1.3s; offline it hangs to the timeout).
Also restores prompt_cache_key delivery, found dead during the flip's
gating verification: openai-agents 0.20 no longer derives it from
RunConfig.group_id, so both lanes sent nothing. ModelSettings.extra_body
is the one channel all three model classes put on the wire (the bare
kwarg is dropped by LiteLLM; extra_args[extra_body] collides with the
responses model's own parameter — both wire-verified), and it is scoped
to OpenAI destinations: LiteLLM plants extra_body as a literal field in
other providers' bodies, and Anthropic rejects unknown fields — the
anthropic wire test now pins the absence.
Verified before landing: mock-server matrix (OPENAI_BASE_URL + bare
name works through LiteLLM; gpt-named self-hosted models are NOT
bridged off a custom base_url; prompt_cache_key on the wire in every
OpenAI lane with distinct per-conversation keys; anthropic body clean)
and live (sol answers through chat(), gpt-5.4 unchanged, responses()
bare unchanged with the key on its wire).
* refactor: no prefix-triggered direct lane — chat model names are LiteLLM's, verbatim
Ray's ruling on the flip's remaining carve-out: a routing decision must
never hide in a model-name prefix. openai/ now means what LiteLLM says
it means (its openai provider), like every other name on the chat lane —
the grammar is LiteLLM's with zero exceptions.
The two defenses for keeping a direct carve-out had no concrete victim:
debugging isolation (litellm is unavoidable in indexing anyway, and
responses() IS the OpenAI-SDK-native door), and endpoint determinism
(litellm sends chatcmpl for openai-provider models except the gpt-5
bridge, which never fires against a custom base_url — wire-verified).
If a direct escape is ever needed, it will be a declared parameter,
never name grammar.
openai/-prefixed names keep the build-time OPENAI_API_KEY check for
parity with bare names; responses() is untouched (bare and openai/
still drive the OpenAI SDK — LiteLLM cannot speak that protocol).
* fix: third-audit findings — extra_body naming drift, litellm floor hint
The prompt_cache_key delivery channel went through three iterations and
settled on ModelSettings.extra_body; the _conversation_cache_key
docstring still named extra_args from the middle iteration.
The litellm install hint said >=1.30, below both our own pyproject floor
(>=1.84.0) and the floor openai-agents' litellm extra declares (>=1.83).
A user in a broken environment following it would land on a version the
package itself rules out. The hint now matches the declared floor.
* test: pin the bundle door's Agents-SDK model grammar; rename the cache-key test to extra_body
The config bundle hands its model string to the Agents SDK's own
MultiProvider grammar, which refuses unknown prefixes (probe on 0.20:
'anthropic/x' -> UserError: Unknown prefix). _normalize_retrieve_model's
litellm/ spelling is what keeps that door working — a link the existing
self-referential assert (config["model"] == client.retrieve_model)
could not catch. Pinned with a provider-slashed name.
Also renames the cache-key delivery test to its real channel,
extra_body — the extra_args name survived from the superseded delivery
attempt.
* refactor: name the retrieve_model helper for its reason — the Agents SDK's grammar
_normalize_retrieve_model said what it does, not why. The litellm/
spelling exists because the Agents SDK resolves raw model strings with
its own prefix grammar and refuses unknown prefixes — the name now
points at that constraint.
* test: the bundle-grammar test skips without openai-agents, like its file's siblings
Every agents-dependent test in this file importorskips; without the
guard this one errors where the others skip.
* feat: index_model + chat_model — two-knob model surface with full legacy fallback
The documented surface becomes two role knobs: index_model builds the
index, chat_model answers on the chat surfaces. model turns into the
set-both umbrella (its 0.2.8 indexing semantics are a strict subset, so
old configs run unchanged); summary_model and retrieve_model stay
accepted as legacy role names.
Resolution lives in ConfigLoader.load(), the one seam every consumer
already passes through (client, CLI standard/md paths, flash's
summary fallback, tree_optimize's default_model): new names win over
old, specific over general, model sets every role, and code constants
close each chain. The packaged yaml no longer ships model keys — key
presence is what separates a user's explicit choice from a built-in
default, and _validate_keys accepts the five model names explicitly.
Consequences: with no config at all, classic-mode structure extraction
now uses DEFAULT_INDEX_MODEL (gpt-5.6-luna) instead of the yaml's old
gpt-4o-2024-11-20 line (ratified; flash-default users see no change).
client.retrieve_model becomes a read-only alias for client.chat_model.
The resolution matrix test pins one row per released generation:
0.2.8 (model), 0.3.0.dev (model+retrieve_model), 0.2.10.dev (all three
legacy names), the new pair, umbrella-only, and mixed.
* feat: --index-model on the CLI; README flag docs follow
The CLI leads with --index-model; --model stays as its legacy synonym
(the CLI only indexes, so the umbrella and the index role coincide).
The flash branch's summary fallback gains the index position, and the
standard branch now forwards --summary-model, which it had silently
ignored — the flag's help always claimed it worked there. The md
branch's unfiltered model=None no longer clobbers the default: the
resolver treats None as unset.
* feat: default chat model becomes gpt-5.6-sol
Ray's pick for the out-of-box QA default; indexing stays on luna. sol
runs tools on the Responses lane — current litellm bridges chat()
there automatically; older litellm gets the guided 400 naming both
exits.
* chore: trim the model-keys comment to the constraint
The per-generation history lives in 7244ee4's message.
* feat: per-door reasoning passthrough — reasoning_effort / reasoning / thinking
Each chat door gains its own protocol's native thinking control,
forwarded verbatim with no invented vocabulary and no default of ours:
chat_completions(reasoning_effort=...), responses(reasoning={...}),
messages(thinking={...}). Unset sends nothing, so backend defaults
(sol: medium, adaptive) are untouched. chat() stays answer-only.
Delivery channels, each verified: the chat door rides
extra_args["reasoning_effort"] — LiteLLM's own top-level kwarg on
every supported openai-agents version, admitting non-enum values
("none"); newer openai-agents promotes it to the top-level argument
and pops the duplicate. Wire-captured on a mock backend
(/v1/chat/completions body carries it) and coexists with the Claude
cache marker in one dict. The responses door rides
ModelSettings.reasoning — coerced to the typed openai Reasoning object
and forwarded verbatim by the Responses model; the envelope echoes the
caller's dict. The messages door joins the existing anthropic
passthrough dict, asserted through the real tool runner.
LiteLLM semantics observed and accepted as-is: unknown models refuse
the param loudly with LiteLLM's own remedies, and gpt-5.4+ names with
an explicit effort route to /v1/responses even against a custom
api_base (its documented pre-existing arm). The sol-class 400 guidance
now names the third exit — an explicit effort routes on older litellm
releases too.
Cloud chat_completions rejects the new parameter like model/max_turns;
responses()/messages() are local-only already.
* feat: extra_body escape hatch on the three protocol doors
The industry-standard per-request extension channel (openai/anthropic
SDK trio): a dict merged verbatim into the backend request, last, so
caller keys win over SDK-set ones. Routing per door: OpenAI-compatible
destinations get a true body merge (ModelSettings.extra_body / the
anthropic SDK's native extra_body); LiteLLM-routed providers take the
keys as LiteLLM's own top-level kwargs instead, since LiteLLM plants
extra_body as literal fields other providers reject. Cloud mode rejects
it like the other local-only knobs. chat() stays answer-only.
* chore: litellm floor 1.84 -> 1.97.0
1.97.0 is where the unset-effort chatcmpl->responses bridge landed
(responses_api_bridge_check's on_constraint_enforcing_endpoint arm,
A/B-verified against 1.96.2), so sol-class models work through the
chat lane out of the box instead of 400ing until a manual upgrade.
Three spots move together: the pyproject floor, the requirements.txt
CI pin, and the install hint.
* feat: named top_p/max_tokens on the chat door, max_output_tokens on responses
extra_body could not carry these: openai-agents' LitellmModel passes
every ModelSettings sampling field as an explicit keyword and unpacks
extra_args into the same call, so the common knobs collided with a bare
TypeError on the LiteLLM lane (reproduced against a stub — Python call
semantics, callee-independent). Named params ride ModelSettings fields,
the one channel clean on every lane; responses() uses the protocol's
own name (openai_responses maps ModelSettings.max_tokens to
max_output_tokens on the wire) and the envelope now echoes the real
value instead of a constant None. Both caps bound each backend call in
the agent loop, not the whole run — documented. Cloud rejects them like
the other local-only knobs; the long tail (frequency_penalty etc.)
stays extra_body-blocked-loudly on that lane by choice.
* chore: the missing-key error names chat_model as the other exit
The default chat_model is what put keyless users on the OpenAI lane,
so the error now points at the knob that picks a different backend.
|
||
|
|
08a912c69c |
feat: openai-agents becomes a base dependency (#406)
feat: openai-agents becomes a base dependency — the chat engine ships with the SDK chat() is the SDK's front door, and its engine lived behind a vendor-named extra: pip install pageindex could index a document but failed on the first chat call, and chatting with Claude required installing '[openai]'. Measured before moving: the base tree already carries litellm (75 MB) + openai (13 MB), openai-agents adds ~15 MB (agents 8.1 + mcp 1.7 + griffe 1.4 + small pure-python deps), and current litellm's openai range (>=2.20,<3) intersects cleanly with openai-agents' (>=2.45,<3). The [openai] extra stays declared but empty, so existing pip install 'pageindex[openai]' commands keep resolving. Error messages and docstrings drop the extra; requirements.txt gains the dependency, so CI now runs the openai-agents test lane instead of skipping it. Extras now mean exactly one thing: a vendor's own SDK surface ([anthropic] for messages()/tool runner, [claude] for the Claude Agent SDK).v0.2.10.dev3 |
||
|
|
ec4851059f |
feat: chat() front door and Claude cache marking on LiteLLM routes (#405)
* feat: chat() — the answer-out front door over chat_completions * feat: cache-mark the managed prefix on anthropic-routed LiteLLM models * refactor: cache predicate asks litellm's own provider resolution * feat: extend cache marking to Claude on Bedrock and Vertex — both live-verified * fix: point anthropic-extra users at messages(); guard the two silent vendor chains * docs: the max-tokens table is a closed set — litellm's map prunes EOL'd entries, so it cannot replace it * chore: trim the max-tokens docstring to the contract * docs: state the local text-only history contract — cloud forwards tool turns, local rejects; extra fields dropv0.2.10.dev2 |
||
|
|
7c6e3c1c8f |
feat: Flash with full optimization becomes the default local indexing mode (#404)
* fix: break the phantom exception chain in _run_sync
Move asyncio.run(coro) out of the except RuntimeError block so real
errors no longer carry a bogus "no running event loop" context in
their traceback.
* feat: Flash with full optimization becomes the default local indexing mode
Every entrance now defaults to Flash with the full optimize pass
(deterministic merge, then LLM expand), replacing the standard LLM-built
tree as the default:
- submit_document(): mode=None now means "flash"; pass mode="standard"
for the LLM-built tree. _index_flash runs optimize="full" with the
expand model = summary_model, and fails fast with the missing key
name(s) via litellm.validate_environment before any work.
- page_index_flash(): optimize takes "full" (default) / "merge" / False;
True is accepted as "full" for compatibility, unknown values raise
instead of silently degrading to merge-only. optimize_expand stays
honored for legacy callers.
- CLI: --mode {flash,standard} replaces --flash (kept as a hidden
compatibility alias that forces flash). --optimize defaults to full in
flash mode with an `off` choice; explicitly passing it outside flash
still errors. Standard-only tuning flags (--toc-check-pages,
--max-*-per-node, --if-add-*) now error in flash mode instead of being
silently ignored, mirroring the existing flash-only flag errors. The
key pre-check runs only when an LLM will actually be called, so
--no-summary --optimize off|merge works keyless. Output drops the
_structure_flash suffix — always <name>_structure.json.
On the Disney earnings PDF the optimized default is also faster than
unoptimized flash (fewer nodes to summarize) and fixes hierarchy
mistakes; both modes emit identical schemas end to end.
Docs updated to match (mode flag, defaults, LLM usage honesty); tests
pin the new defaults: stored mode == "flash", optimize passthrough, and
the unknown-optimize rejection.
v0.2.10.dev1
|
||
|
|
ec342c4e28 |
fix: post-merge review fixes for v0.2.10
- Conformant responses() envelope: official output (model items only) + items (full transcript for round-trip), usage details aggregated across turns; verified against real OpenAI API - Python floor corrected to >=3.10 (litellm stable requires it) - Three stale anthropic>=0.84.0 hints updated to 0.108.0 - Image-stub behavior disclosed on tool-builder docstrings - Non-essential comments trimmed |
||
|
|
4e41acdc68 |
feat: agent tools and local chat for the PageIndex SDK (v0.2.10) (#396)
* feat: agent tools — the cloud MCP tool contract on the client Four new client methods make PageIndex documents available to agent frameworks, in both modes, with the mode decided solely by the client constructor: - agent_tools(): plain functions (browse_documents, get_document, get_document_structure, get_page_content) matching the PageIndex cloud MCP server's tools/list — same names, schemas, descriptions, and JSON response envelopes — so agent prompts port unchanged between the cloud MCP connection and these in-process tools. Tools never raise; errors come back in the same envelope. remove_document ships behind include_management=False. - as_openai_tools(): the same tools wrapped for the OpenAI Agents SDK. - as_claude_mcp(): one mcp_servers entry for the Claude Agent SDK — cloud clients get the remote MCP config (the framework connects to api.pageindex.ai/mcp and discovers the full cloud tool set), local clients get an in-process SDK MCP server. - agent_instructions(doc_id=None): orchestration guidance for the agent's system prompt; doc_id (same shape as chat_completions) appends the target documents. submit_document() gains wait=True: poll get_document status until completed, raise on failed or after 30 minutes — the manual polling loop every cloud caller writes today spins forever on a failed document. Neither framework becomes a dependency: imports happen at call time with actionable errors, and pageindex[openai] / pageindex[claude] extras are floor-only pins. tests/data/cloud_mcp_contract.json freezes the tool contract; a parity test guards against drift. 36 new tests (95 total), plus a live OpenAI Agents SDK run over a seeded local store verifying the structure-first navigation flow end to end. * fix: agent tools review — next_steps order, resolve caching, error semantics - Large-doc next_steps now says structure-first, consistent with tool descriptions and agent instructions - _remove_document fetches document list once instead of per-name - call_tool returns error envelope for unknown names instead of raising - _not_ready_error timed_out flag reflects actual wait outcome - openai_agents.py docstring corrected to match default (FunctionTools) - Removed unused ModelSettings import from demo * fix: agent tools review 2 — bridge thread safety, browse paging, metadata merge - McpBridge reads session/protocol headers under the lock (now RLock: _ensure_initialized posts while holding it). openai-agents runs sync tools on threads and executes parallel tool calls concurrently, so bridge functions genuinely race; a torn read sent a new session id with a stale protocol header. Measured: one session expiry under 8 threads cost 4 initializations before, minimal 2 after. - Session-expiry retry also resets the negotiated protocol version, so the re-handshake carries no stale MCP-Protocol-Version header. - browse_documents time sort pages list_documents natively instead of fetching the whole library to slice one window (relevance still needs the full list for scoring). - _await_completion: a status refetch that nulls out metadata no longer clobbers the listing's copy (setdefault was a no-op on existing None). - Structure tool reads the raw stored tree via a named LocalAPI raw_tree() seam instead of reaching into _api._store internals; drop the redundant deepcopy before _format_structure (store re-reads from disk, formatting builds fresh containers). - Shared pageindex/_version.py replaces _sdk_version duplicated in mcp_bridge and the Claude integration. Left as-is after source verification against the cloud MCP: first-page budget bypass, pageNum falsy-zero, and the page-gap fallback text are letter-for-letter cloud behavior — parity wins over local repair. * fix: agent tools review 3 — page-span cap, duplicate names, wait resilience, contract drift - _parse_page_spec bounds the requested span arithmetically (10k pages) before materializing it; pages="1-1000000000" previously expanded to a billion integers inside the caller's process. - Local submit_document uniquifies document names the way the cloud upload does (taken name -> _1.._99, then reject with the cloud's own message). Same-name duplicates broke name-addressed tools: resolution always picks the newest, so older duplicates were unreachable. - agent_instructions(doc_id=...) now fails loud when the pinned doc's name is shadowed by a newer same-name document (legacy stores predate the rename) — it previews resolution with the same _resolve_document the tools use, so the check cannot drift from actual behavior. - submit_document(wait=True) tolerates transient network errors, not just API errors; a dropped connection at minute 25 of a 30-minute wait no longer kills it. Third strike wraps into PageIndexAPIError per the documented contract. - The live contract-parity test compares full per-param schemas, not just names and descriptions. It immediately caught real drift the shallow check had been passing: the server now emits nullables as anyOf unions and stamps MAX_SAFE_INTEGER maxima on offset/part. Contract and snapshot updated to the served wire form; _annotation_for learned anyOf so bridge signatures stay Optional[str] instead of degrading to Any. Adjudicated, not changed: the allowed_tools wildcard example stays (docstring advice covers scoping; Ray's call), and raw-length response accounting stays (letter-for-letter cloud behavior, parity wins). * feat: surface the stored document name from submit_document Compute PR #558 makes /doc/ return {"doc_id", "name"} carrying the post-dedup-rename name. Mirror it end to end: local submit returns the stored name, the client warns when it differs from the uploaded file name (read via .get so older cloud servers stay compatible), the local name-exhaustion check runs before indexing instead of after the LLM spend, and the demo caches doc_id in a file instead of name-matching — a renamed document made the name lookup re-index on every run. * fix: add missing page_list kwarg in duplicate-name test mock * revert: keep README.md unchanged from main — SDK section deferred * feat: serve cloud agent instructions live from the MCP server The cloud MCP server publishes its agent instructions in the initialize result, adapted to each key's tool set. agent_instructions() previously returned the SDK's local-subset text in both modes — a silently forked copy that lacks the guidance for cloud-only tools (search_documents escalation, folders, images) and drifts as the server's prompt evolves. Cloud clients now serve the server's live instructions, captured from the initialize handshake on a per-client bridge shared with agent_tools() (one session, no extra request). An empty server response raises instead of silently substituting the subset text — same posture as the annotation-regression guard. The local constant stays as the honest subset for the in-process tools, with its provenance noted and a consistency test that every tool it names exists in the local registry. * fix: local relevance sort answers honestly instead of imitating sort="relevance" is cloud-side semantic ranking; the local substring imitation could satisfy the letter of the interface while silently missing semantically relevant documents. Per the honest-subset rule (same treatment as folders), local now returns the "not available here" envelope for sort="relevance" or a stray query, and the local instructions steer discovery through name/description matching plus full-library paging instead of prescribing a capability that does not exist here. The tool schema keeps the cloud contract verbatim, like folder_id: honesty lives in the runtime answer, not a forked contract. * docs: note the cloud+Claude instructions duplication trade-off in as_claude_mcp * fix: unsupported-capability envelopes say local-mode-yet, point to cloud "Not available here" read as a broken feature; the honest framing is that folders and semantic ranking exist on PageIndex cloud and are not in local mode yet. Both envelopes now say so and name the cloud client in next_steps, so agents relay an accurate story to the user. * fix: local tool descriptions pre-announce cloud-only capabilities The cloud-verbatim browse_documents description invites sort="relevance" and folder drilling, so a local agent's first semantic search attempt was a guaranteed dead end discovered only from the runtime error envelope. Local registration now appends a LOCAL MODE note to the description — the agent learns what is cloud-only before calling; the runtime envelope stays as the backstop for prompts that ignore descriptions. The cloud-facing contract stays byte-verbatim. * refactor: localized tool guidance replaces the appended LOCAL MODE note Appending a retraction to the cloud-verbatim description left the model parsing an instruction and its negation — and kept the cloud text recommending search_documents and get_folder_structure, tools that are not registered locally (get_page_content likewise pointed at get_document_image). Guidance now adapts to the local surface the way AGENT_INSTRUCTIONS already does: schema structure stays byte-identical to the contract (mechanically asserted by a strip-descriptions test), while local description strings teach only what works here and point to PageIndex cloud for the rest. A dead-reference test forbids local guidance from naming tools outside the local registry, so a contract refresh that reintroduces a cloud-only reference fails loudly. * feat: hide cloud-only parameters from the local tool surface folder_id, sort, query, and recursive were exposed locally with localized "cloud-only" descriptions, leaving the dead-end calls expressible and discovered at runtime. Schema constraints beat guidance: the local surface now serves the contract minus these parameters, so strict-schema frameworks make the calls inexpressible and a prompt that insists on sort="relevance" degrades to the bare call (the correct local behavior) instead of an error round-trip. The implementations still accept the hidden parameters and answer with the guided "works on PageIndex cloud" envelope — the backstop for direct call_tool callers and hosts without schema enforcement. wait_for_completion stays: seeded or torn stores can hold documents that are genuinely not completed. The structural guard now asserts the local schema equals the contract minus the documented hidden set, descriptions aside. * fix: incremental-review findings — bridge cache, guards, envelope drift Three independent review passes over the agent-instructions increment surfaced six fixes: - The per-client bridge moved off the instance into a weak-keyed, lock-guarded module cache: cloud clients stay picklable (threading.RLock no longer rides on the client) and concurrent first calls can no longer construct duplicate bridges/sessions. - Blank or non-string initialize.instructions now hit the same honest error as a missing one — a whitespace-only or structured value could previously become the system prompt (or crash the doc_id append with a raw TypeError). - The invalid-sort envelope no longer prescribes sort="relevance" — the one error text that still taught the cloud-only value it would then reject. - "Page through the rest of the library" is emitted only when has_more is true; a fully-listed library no longer instructs a pointless call. - The mandatory full-library paging step now says limit: 50 — 6 calls instead of 30 on a 300-document library. - Docstrings and comments rescoped to what is actually true: the never-raise contract covers invocations the signatures accept (unknown params fail at the Python boundary; call_tool answers them with the guided envelope), recursive is accepted as the identity rather than errored, lenient framework arg models drop hidden params pre-call, and the module header no longer claims full schema parity. The capability-phrase guard now covers every local docstring, not just browse_documents. * chore: keep the demo's doc_id cache file out of the repo * test: live envelope field-parity guard against cloud response drift The frozen contract guards tools/list, but the response envelopes the local tools emit were hand-built to mirror the cloud's and had no drift detector. A key-gated live test now asserts every field local emits exists in the live cloud response for the analogous call (top-level keys, next_steps, document entries, structure nodes, content entries). Guidance wording is deliberately localized and not compared. Verified green against the live server: local and cloud field structures currently match exactly. * feat: local chat — three protocol surfaces over the agent tools (v0.2.10) Local mode gains managed document QA: an agent over the #393 local tool set, reachable through three wire protocols, each 1:1 with the backend and with no translation layer. - chat_completions(): standard chat.completions semantics on any OpenAI-compatible backend (openai-agents engine). Final answer only, cross-turn aggregated usage, streaming as text pieces or chunk dicts (the existing cloud signature, now implemented locally; model and max_turns are local-only additions). - responses(): the agentic surface — OpenAI Responses format, the tool process is standard output items, streaming forwards native events (tool outputs emitted as response.output_item.done, the way the platform streams its own server-side tools). Round-tripping output into the next input keeps provider prompt-cache prefix continuity and the agent's memory — live-verified: the follow-up call answered from round-tripped tool output with zero new tool calls. - messages(): Anthropic-native via the SDK's own tool runner (new pageindex[anthropic] extra, floor 0.68.0 verified for tool_runner/beta_tool(input_schema)). tool_use/tool_result round-trip is the format's native behavior; the envelope is the final message with aggregated usage plus the full new-turn sequence; the managed system blocks carry cache_control breakpoints. Shared skeleton: thin chat header + the local AGENT_INSTRUCTIONS (caller system content is appended, not rejected), the doc_id targeting block as a leading context item (factored out of build_agent_instructions), read-only toolset, structural-only validation (no arbitrary caps — backend limits govern), sampling params passed through, per-run tracing disabled, enable_citations rejected as cloud-only. Design basis is industry-standard formats rather than the cloud chat endpoint; responses()/messages() raise on cloud clients until the cloud converges. Tests run the real engines against scripted backends (a Model fake for openai-agents, a mock HTTP transport under the real anthropic SDK) with real tool execution against a seeded store, including the round-trip prefix-extension assertions on both engines. * fix: local-chat review findings — truncation, serialization, streams Three independent review passes (bug scan, claims-vs-code, adversarial runtime probes) over the local-chat increment; every fix below was reproduced before being fixed. messages(): - A max_turns cut no longer duplicates the final assistant turn: the runner has already appended it when iterations exhaust, so the round-trip history carried a duplicate tool_use id and ended on an unanswered tool_use — a guaranteed 400 on continuation. The append now keys on stop_reason, and truncation reads natively as stop_reason: "tool_use" with a continuable history. - The envelope is JSON-serializable end to end: runner-stored turns carry pydantic content blocks; everything is dumped to plain dicts, excluding SDK-internal __api_exclude__ fields (parsed_output) that the API rejects on round-trip. - Bounded by default (max_iterations 10, like the OpenAI surfaces); usage aggregation now preserves the final turn's native fields and sums the token counters None-safely; empty caller system strings are skipped; non-dict message entries and bad doc_id types raise PageIndexAPIError; anthropic < 0.68 gets an actionable version error; the doc block no longer spends a cache_control breakpoint. chat_completions()/responses(): - MaxTurnsExceeded wraps into PageIndexAPIError on all four run paths. - responses(stream=True) is one logical response: per-turn backend lifecycle events are collapsed (a canonical consumer previously stopped at turn 1's response.completed and never saw the answer), sequence numbers are reassigned monotonically, and the synthesized tool-output event carries output_index/sequence_number. - The responses envelope carries the real request surface (instructions, the actual function tool definitions, tool_choice, parallel_tool_calls, error/incomplete_details). - RunConfig(group_id) pins a stable prompt_cache_key: openai-agents otherwise stamps each run with a fresh key, tagging round-tripped prefixes as different cache groups and defeating the feature the round-trip exists for. - Abandoning a stream now cancels the run: a watchdog task lets the cancellation land even while the pump awaits the backend, and the per-call AsyncOpenAI client is closed before its loop ends (fixes "Task exception was never retrieved" noise). The opening role chunk is emitted even for empty outputs; empty responses() input and enable_citations-before-extra ordering fixed. Docs rescoped to what is true: finish_reason/status reflect loop completion on the OpenAI surfaces (the engine does not surface per-turn backend reasons); chat streaming yields visible narration including pre-tool text; messages(stream=True) forwards the Anthropic SDK's native event objects (not wire-verbatim); the doc block is a leading conversation item on OpenAI surfaces and a system block on messages(). Tests: 25 in the file (11 new), with per-extra skip sections so a machine with only one framework still covers the other surface; without-frameworks matrix re-verified; live smoke re-run green with a clean exit. * feat: as_anthropic_tools — Anthropic tool-runner export, both modes Fills the last cell of the agent-connection matrix: users driving their own anthropic tool_runner loop get runnable tools directly. Cloud wraps the live MCP tool set with input schemas passing through verbatim (MCP inputSchema is the Messages API schema shape); local exposes the same set messages() runs internally. The beta_tool wrapping moves from local_chat into integrations/anthropic_sdk.py, parallel to openai_agents.py, and messages() now consumes the shared builder. agent_tools grows _bridge_invoker/_read_only_tools so the plain-function and beta_tool cloud paths share invocation containment and the read-only gate. * fix: as_anthropic_tools review findings — async flavor, schema isolation Adversarial + best-practice review of |
||
|
|
d375c00a5a |
feat: local mode for the PageIndex SDK (v0.2.9) (#389)
* feat: add local mode to the PageIndex SDK client
One PageIndexClient, two backends. With api_key: the 0.2.x cloud SDK,
request for request (with the reviewed fixes: bounded timeouts on JSON
endpoints, none on uploads, URL-encoded ids, 401 key hint, empty DELETE
body tolerated). Without api_key: the same methods run locally —
page_index builds the tree in submit_document (mode="flash" uses
PageIndex Flash), documents are stored as plain JSON per doc under
storage_path, submit_query is LLM tree search with retrieve_model, and
chat_completions answers over the retrieved nodes with OpenAI-style
responses and streaming.
Local responses mirror the cloud wire shapes verified against the server
source: tree nodes rename start_index to page_index and drop end_index,
a non-leaf summary becomes prefix_summary, and the metadata/list/delete/
retrieval envelopes match key for key. Cloud-only features (folders,
beta_headers, enable_citations) raise instead of pretending.
Replaces the demo-only workspace client (index/get_document_structure/
get_page_content had no real users) and its retrieve.py helpers.
page_index_main gains an optional logger param so the SDK can keep
./logs out of the caller's working directory; pymupdf import is now lazy
(only the optional PyMuPDF parser path needs it).
* chore: package pageindex 0.3.0.dev4 for PyPI
Poetry packaging for the combined SDK + local pipeline: every production
import is a declared dependency (openai and requests join requirements.txt
for the same reason), config.yaml and the flash data tables ship in the
wheel, the benchmark PNG does not. pymupdf drops to an optional note now
that its import is lazy. dev4 follows the already-published 0.3.0.dev1-3;
pip still resolves plain 'pip install pageindex' to 0.2.8 until a final
0.3.0 — install with --pre.
* docs: add SDK section to README; move the agentic demo onto the SDK
The demo keeps its flow and the post-cutoff demo paper, swapping the
removed workspace client for PageIndexClient local mode (list_documents
for the doc-id cache, get_tree/get_ocr behind the agent tools). The old
examples/workspace JSONs demoed the removed format and go with it.
* refactor: rebuild cloud_api on the 0.2.8 client text
Ray's rule for the cloud half: his 0.2.8 code is the base; a Kylin-lineage
change survives only when strictly better — invisible on healthy traffic
while fixing a real failure mode. Kept under that bar: request timeouts
(dead connections hung forever; uploads still pass none, exactly like
0.2.8), the upload handle closed via with (leaked on request errors),
URL-encoded path ids (a crafted id could reroute the URL), the
empty-DELETE-body guard, and stream hardening (choices guard,
response.close in finally). Reverted as not strictly better: the
lowercase summary param (the server accepts both spellings), the 401
message hint (visible text change; AUTH_HINT dropped from errors.py), and
the _request/requests.request reorganization — every method body,
docstring, and section comment is 0.2.8's text again.
diff -w against ../pageindex_sdk/pageindex/client.py now reads as that
surgical patch plus plumbing: CloudAPI reads BASE_URL/api_key through the
owning client, PageIndexAPIError comes from errors.py, and
is_retrieval_ready lives verbatim on PageIndexClient shared by both
modes. A mocked-requests harness driving 0.2.8 and this file through 19
identical calls shows the only remaining request-level difference is
timeout.
* refactor: drop the local retrieval endpoints — cloud-only, deprecated
Ray: the cloud already marks POST /retrieval/ and GET /retrieval/{id}/
deprecated in favor of chat completions, so local mode should not grow a
fresh implementation of a retiring surface. submit_query/get_retrieval
now raise in local mode with a pointer to chat_completions; cloud mode is
untouched (the endpoint still works there and 0.2.8 code keeps running).
The tree search that backed them stays as chat_completions' retrieval
engine; the retrievals/ storage goes away. This also closes the one real
cross-mode parity gap — the retrieved_nodes inner shape — by removing
its local half.
* feat: manifest.json — one-file document listings for the local store
Ray wanted a metafile that shows every document in one place instead of
per-directory reads. It is a cache, never a second source of truth:
writers update it best-effort after save/delete (atomic replace, no
locks), and list_metas trusts it only while its id set matches the docs/
directory names — documents are immutable, so matching names imply valid
content. Any mismatch (lost concurrent update, crash, corrupt or deleted
manifest) rebuilds it from the doc.json files, reading only the missing
entries. Incomplete dirs (no doc.json) stay invisible and are never
recorded, so a save that completes later is still picked up.
1000-doc listing: 37ms of per-dir reads -> 2.3ms warm (scandir names +
one manifest read); one-time rebuild 160ms.
* fix: align local doc_id prefix and createdAt format with the cloud
Local doc ids now carry the cloud's pi- namespace prefix (random token
stays uuid4 hex — nothing parses cuid internals, the prefix is the
contract; chat ids already mirrored chatcmpl-). createdAt now matches
the server byte for byte: the cloud emits the DB datetime's bare
isoformat — naive UTC, second precision — while we emitted microseconds
plus +00:00.
* feat: optional metadata tags on submit_document (both modes)
Ray's call after the alignment review: the metadata key in tree/OCR
envelopes and list entries should carry real data, and the only honest
way is exposing the field the server already accepts. Cloud mode
forwards it as the existing metadata form field; local mode validates it
early (a JSON-serializable dict, checked before any LLM spend), stores
it in doc.json, and returns it from the same three places. get_document
still omits it, mirroring the server, whose metadata-endpoint SQL never
selects that column. Scope stays deliberately narrow: set at submit and
read back — no metadata_filter, no update API.
* docs: state that createdAt is UTC and show how to localize it
The value is naive UTC in both modes (the cloud column is timestamp
DEFAULT CURRENT_TIMESTAMP on a UTC server, emitted via bare isoformat).
Wall-clock display is the consumer's layer: emitting local time under
the same format would silently change meaning per machine, and adding an
offset marker would break both the byte-format parity and string-order
sorting.
* fix: createdAt carries milliseconds, matching the cloud's datetime(3)
The earlier second-precision alignment was reasoned from the postgres
schema file, but production is MySQL (DATABASE_BACKEND defaults to
mysql) and its FilePageIndex.createdAt is datetime(3) DEFAULT
CURRENT_TIMESTAMP(3) — so the server isoformat()s a millisecond-
precision naive-UTC datetime, emitting .XXX000 fractions (bare seconds
only when the millisecond happens to be zero). Local now generates
through the same mechanism. The docs' 2024-01-15T10:30:00.000Z sample is
a JS-style string Python isoformat cannot produce — not evidence.
* fix: createdAt at millisecond precision, matching the timestamp(3) column
The cloud column is timestamp(3); its datetimes render through bare
isoformat() as six fractional digits ending in 000 (or no fraction when
the millisecond is exactly zero). Truncate to the millisecond and render
the same way, replacing the second-precision guess from
|
||
|
|
b723c9f0a7 |
ci: run the test suite on every push and pull request
Publishing was the only automation touching code: a tag builds and ships to PyPI without ever running a test, and pull requests get no checks at all. This runs pytest on a small matrix — Python 3.10 (the floor) and 3.13, each with and without the agent frameworks installed, so the lazy-import contract (the package must work with neither framework present) is enforced rather than assumed. |
||
|
|
d5c4e62c20 |
Exclude common words from roman numeral conversion (#387)
Update PageIndex Flash |
||
|
|
933df35b2f |
Add a Flash summary flag (#386)
* Update PageIndex Flash * Add a summary flag |
||
|
|
502b763876 | Update README (#385) | ||
|
|
3c4c7e8353 |
Use embedded PDF bookmarks in Flash trees (#384)
* Update PageIndex Flash * Add embedded ToC CLI flag and README note |
||
|
|
fb3f6e441a | Fail fast on rejected keys and unknown models (#381) | ||
|
|
03cf0fa0ef | Sync PageIndex Flash from private branch (#380) | ||
|
|
d409566464 | Sync PageIndex Flash from private branch (#379) | ||
|
|
1b2fdadf51 | Keep merge inside --optimize (#377) | ||
|
|
2a29ac5aa0 |
Merge pull request #375 from VectifyAI/fix/default-merge
Refine the Flash tree with merge |
||
|
|
9e789babf3 | Add benchmark to Flash README (#376) | ||
|
|
e17dac2053 | Strip internal marker from optimize output | ||
|
|
278cc7f28a |
improve tree optimization
Carried from bench/flash-cost; benchmark artifacts left there. |
||
|
|
4a116daa2f | Replace thinning with merge | ||
|
|
3f33a53b50 | Add tree optimization and node summaries for Flash | ||
|
|
afb3100232 |
Merge pull request #374 from VectifyAI/fix/flash-requirements
Add Flash dependencies to requirements |
||
|
|
6e248f8f05 | Add Flash dependencies to requirements | ||
|
|
d178bc557d |
Merge pull request #372 from VectifyAI/openai-sdk-fast-path
Use OpenAI SDK directly for OpenAI models |
||
|
|
548d3201b7 | Use OpenAI SDK directly for OpenAI models | ||
|
|
66d5b9ac9b |
Merge pull request #371 from VectifyAI/fix-flash-readme
Add --flash CLI flag and update README |
||
|
|
faf3c57fc8 | Add --flash flag to CLI | ||
|
|
a77301e5c2 |
Merge pull request #370 from VectifyAI/fix-flash-readme
Update Flash README links |
||
|
|
165fcea3db | Update Flash links to main branch | ||
|
|
c9a01fad13 |
Merge pull request #369 from VectifyAI/pageindex-flash
Add PageIndex Flash |
||
|
|
fffc2c5512 | Add PageIndex Flash | ||
|
|
0e4b68c92d |
Merge pull request #367 from VectifyAI/add-summary-model-config
Add summary_model config |
||
|
|
acfcde2c58 | Remove summary_model from page_index() signature | ||
|
|
820e8bf103 | Use summary_model for doc description, expose in page_index() | ||
|
|
26eb0e12dc | Add summary_model config with fallback to model | ||
|
|
a4c51126c4 |
Merge pull request #366 from VectifyAI/add-page-level-thinning
Add page-level thinning to PDF pipeline |
||
|
|
1246d56fae | Add page-level thinning to PDF pipeline | ||
|
|
982514ab40 | Trim comments (#365) | ||
|
|
df8c9283da | Lazy-load litellm (#364) | ||
|
|
39121c4d34 | Update README (#363) | ||
|
|
8ac5f70b68 | Add PageIndex Flash section to README (#362) | ||
|
|
36cc37deef | Add PageIndex Flash preview to README (#361) | ||
|
|
190f8b378b |
Merge commit from fork
fix(security): patch prompt injection in PDF ingestion pipeline |
||
|
|
c58cd62b50 |
Merge pull request #250 from VectifyAI/feat/md-bold-heading-recognition
Recognize whole-line bold as level-1 heading in markdown parser |
||
|
|
01dbcbf47f | fix(markdown): skip empty bold headings | ||
|
|
b625906c78 | fix(security): harden document boundary validation | ||
|
|
3367144d50 | fix(security): validate TOC identity before index updates | ||
|
|
9bcc7cff89 | Refactor TOC entry update logic in page_index.py | ||
|
|
9470b63960 | Fix indentation and formatting in page_index.py |