Commit Graph
424 Commits
Author SHA1 Message Date
Ray d0d6aa39e8 Tool results reach the frameworks as MCP content; the frameworks render it (#487)
* Tool results reach the frameworks as MCP content; the frameworks render it

get_document_image returns an MCP image block, and the SDK flattened every
tool result to text on its way to a framework, so an own-model chat received
"[image/png content omitted: ~256 KB]" where the hosted lanes (HostedMCPTool,
the http MCP config, the managed chat) received the page.

The rule now: the SDK carries MCP types and renders nothing. Each framework
gets the tool set the way it consumes MCP, and converts the content itself.

- McpBridge.call_tool returns the content blocks untouched; the invoke
  contract behind _tool_specs is (content blocks, is_error) on the cloud and
  local paths alike (local tools send one text block).
- openai-agents: an in-process MCPServer over the tool specs, and the
  FunctionTools are MCPUtil.to_function_tool over it. The framework's own
  MCP conversion builds the tools (schema verbatim, strict off) and renders
  results (text items, image data URLs); its failure pipeline answers
  malformed arguments. The hand-built FunctionTools are gone.
- Anthropic tool runner: results validated as CallToolResult and rendered by
  the Anthropic SDK's mcp_content (text blocks, base64 image blocks), on the
  error channel too.
- Claude Agent SDK: the content passes through, it is MCP already.
- Plain functions (agent_tools): render_text, the former flattening, keeps
  the size stub for binary content; a string cannot carry an image.

Wire consequences: tool results on the openai-agents lanes are the
framework's structured items (a single text item for a text result) rather
than a bare string, and the Anthropic tool_result content is a block list.
ChatStream tool_result events carry that structured output; the woven
display shows its text.

mcp becomes a declared dependency: it already arrives with openai-agents,
and both adapters use its types.

Verified live on a 15-page paper, page 1, across chat() on the LiteLLM lane
(gpt-5.6-luna), chat(protocol="responses"), and chat(protocol="messages",
claude-sonnet-4-6): each answer described the red attribution text and the
arXiv stamp rotated along the left margin.

Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW

* Relay test asserts the wire aliases: mcp 1.x and 2.x differ in attribute names

CI installs mcp 2.x, where CallToolResult exposes is_error and content
blocks exposes mime_type as attributes, with the wire names (isError,
mimeType) as aliases; mcp 1.x uses the wire names as attributes. The
adapters only ever validate from and dump to the wire shape, so they run on
both; the test now does the same.

Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
v0.2.15
2026-09-06 19:13:20 +08:00
Mingtian Zhang 348231c5e3 Merge pull request #479 from VectifyAI/zmtomorrow-patch-2
Update platform link to cloud link in README
2026-09-05 15:31:48 +08:00
Mingtian Zhang ad49eb3792 Update platform link to cloud link in README 2026-09-05 08:18:37 +01:00
Ray 2858384a75 docs: the chart images load by absolute URL so PyPI renders them (#475)
pyproject publishes README.md as the PyPI description, and PyPI does not
resolve a relative path against the GitHub repo — the five charts render
as broken images on the project page (verified on 0.2.14). Every chart
now loads from raw.githubusercontent.com; the ten other images in the
file already did.

PyPI's sanitizer drops <source>, so only the <img> fallbacks matter
there; the srcset URLs change too, to keep GitHub's dark variants
spelled the same way.
2026-09-04 15:30:44 +08:00
Ray 0400366d35 litellm terminal noise: quiet as a default, and the no-fetch stamp on every lane (#473)
* fix: flip litellm's banner switch at the provider lookups, not after them

_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.

Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* refactor: the preload thread only imports — every litellm entry quiets itself now

Since 30652ca each get_llm_provider caller and every completion helper
flips the banner switch and clamps litellm's logger before touching
litellm, so the background import's own _quiet_litellm() had no path
left that depended on it; it was also the one place that mutated
process-global logging at a moment the caller could not predict. The
env stamps stay: import-time records still need them. The config-time
repair test no longer needs the preload guard pinned.

Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9

* fix: litellm quiet is a default, not a clamp — and the no-fetch stamp covers every lane

_quiet_litellm() set litellm's loggers to ERROR on every entry, overriding whatever the host, litellm._turn_on_debug(), or LITELLM_LOG=DEBUG had set. It now sets a level only while a logger is still NOTSET: the default stays ERROR, anything set explicitly wins, and there is no per-call setLevel churn.

LITELLM_LOCAL_MODEL_COST_MAP moves from the client's preload to utils' import. The CLI pipeline and openai_agent_config on a cloud client import litellm without the preload and fetched the remote model map on first use (offline: a WARNING). Every litellm lane imports utils first.

The LITELLM_LOG=ERROR stamp is gone: it pinned litellm's stderr handler at ERROR for the whole process (silencing _turn_on_debug for good) while the logger level was what actually gated the chatter — output verified identical with and without it across plain, logging-configured, and debug hosts.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* style: essential comments only

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* test: the utils-import stamp check lives with its package-surface sibling

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* refactor: drop the CLI's duplicate no-fetch stamp — utils' import sets it

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk

* fix: the thinking-budget ceiling imported litellm ahead of the stamp

_default_max_tokens is reached from anthropic_runner_config/messages(). On a managed cloud client the constructor never imports pageindex.utils, and local_chat deliberately keeps utils off its import path, so that lane imported litellm with LITELLM_LOCAL_MODEL_COST_MAP unset: the model cost map was fetched over the network (3517 entries vs the bundled 2982) and offline printed the very warning this branch removes. Verified end to end through the public API, before and after.

Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
2026-09-03 18:37:11 +08:00
Ray 85180eeacd fix: extra_body refuses the skeleton keys; sampling fields ride ModelSettings on LiteLLM-routed models (#472)
* fix: extra_body refuses the skeleton keys

extra_body merges last on every lane, so a caller's system / instructions / input / messages / tools replaced the managed prompt, the conversation or the doc tools wholesale. The run succeeded and answered without tool guidance or document scoping, silently. Those are the SDK's on every lane (the three-layer rule: skeleton, named knobs, caller extras); the door for the prompt is instructions=. Refused at the two seams where extra_body meets the wire, before the Anthropic transport exists on the Messages lane.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit 4d18854378)

* fix: extra_body sampling fields ride their ModelSettings field on LiteLLM-routed models

On the answer lane, extra_body for a non-OpenAI destination becomes ModelSettings.extra_args, the LiteLLM kwargs channel (LiteLLM would plant extra_body as a literal request field, which Anthropic rejects). openai-agents passes ModelSettings' own fields to litellm.acompletion by name beside **extra_args, so a caller's temperature / top_p / max_tokens / penalties / tool_choice in extra_args collided with the explicit keyword and died as a raw TypeError inside the framework. Split by the framework's own contract: keys that are ModelSettings fields ride their field (the caller's value winning over ours), the rest stay LiteLLM kwargs. The field list is the public dataclass, not a hand-kept name list; bad values now fail ModelSettings validation and surface as PageIndexAPIError. response_format is not a ModelSettings field and still has no door on this lane; documented.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit c1ad980505)

* docs: chat_completions extra_body names the skeleton carve-out

chat_completions shares the seam that now refuses the skeleton keys; its docstring still promised an unconditional merged-last win. Same sentence as chat().

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit f9072f9fff)
2026-09-03 18:27:16 +08:00
Ray c8c8293aa7 ChatStream moves to a light module so chat()'s type hints resolve at runtime (#471)
* fix: ChatStream lives in a light module, so chat()'s type hints resolve at runtime

ChatStream was importable in client.py only under TYPE_CHECKING (a real
import would have dragged local_chat's asyncio stack into `import
pageindex`), which left chat()'s return annotation a dangling string:
typing.get_type_hints(PageIndexClient.chat) raised NameError, and so
did anything that introspects signatures — agents' function_tool
(client.chat) died on it before looking at a single parameter. The
class touches neither asyncio nor the agent frameworks, so it moves to
pageindex/chat_stream.py, client.py imports it for real, and the
package exports it directly instead of lazily. `import pageindex` still
leaves local_chat unloaded.

Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY

* fix: the ChatStream move keeps its own invariants — future annotations, a guarded eager path, the old import path pinned

Review of #471 found the move's guards thinner than they look:

- chat_stream.py had no `from __future__ import annotations`, unlike
  every sibling module, which made its `-> "ChatStream"` quotes
  load-bearing: unquoting them — the very edit this move made in
  client.py, and what `ruff --select UP037 --fix` does — broke
  `import pageindex` outright.
- The import in client.py is the whole fix and reads like a
  typing-only one; a comment says why it must stay real.
- The lazy-import test's denylist named no framework, so agents,
  litellm, openai or anthropic could join the eager path with a green
  suite — the cost the module was split out to avoid.
- The type-hints walk had no floor, so it could silently stop covering
  anything, and nothing pinned `pageindex.local_chat.ChatStream`, the
  path the class shipped under in 0.2.11-0.2.14.
- local_chat's module docstring still claimed the class.

All four guards mutation-checked red.

* style: the walk-collapsed assertion message fits the 79-col convention

* style: no rationale comments — the guard is the test, the why is the commit
2026-09-03 18:19:15 +08:00
Ray 564bc97c0f Chat process follow-ups: a mid-stream error raises, .events survives partial reads, bad show_process chokes first (#470)
* fix: .events survives partial reads; hidden call lines still label results; bad show_process chokes first

- ChatStream.events delegated with `yield from`, so a dropped handle
  (next(stream.events), for ... break) closed the shared run on GC and
  the rest of the run silently vanished. A plain loop leaves it alone.
- _weave filled call_args only past the tool_call visibility guard, so
  with call lines hidden the standalone result lines never carried the
  arguments they promise.
- show_process is validated before the stream check: an invalid value
  is refused as such instead of being told to add stream=True and then
  refused again; the managed lane's duplicate choke goes with it.
- Docstring: show_process is not own-model-only.

Claude-Session: https://claude.ai/code/session_016M3qaQedSK7L4DwysFRmk2

* fix: mid-stream error chunk raises instead of ending as a short answer; stream docstring says show_process is on by default

- The managed endpoint reports a server-side failure as a final
  {"error": ...} chunk after the partial answer (api.py refunds the
  credits, then yields it). Neither chunk decoder looked at it, so
  chat(stream=True) and chat_completions(stream=True) in both modes
  ended as an apparently complete short answer with no exception. One
  guard in each decoder raises PageIndexAPIError; the partial answer
  is still delivered first.
- The `stream:` arg and the Returns block still described the pre-PR
  contract (bare text chunks); only the show_process paragraph said it
  is on by default.

Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY
2026-09-03 15:52:20 +08:00
Ray 1ccb31bf9d fix: Responses effort rides extra_body; envelope reports metadata; "" is unset on every lane (#469)
* fix: Responses effort rides extra_body; chat() docs say what the wire does

The rule chat() follows, written down: chat() names PageIndex's own
parameters plus a subset of LiteLLM's unified vocabulary; anything
vendor-specific rides extra_body under the lane's wire names. Three
layers meet on the wire — the skeleton (managed prompt, conversation,
tools) is the SDK's, the named knobs are translated per lane, and
extra_body is the caller's, merged last so it wins. A key the SDK writes
into a nested object must go in through extra_body too, so the caller's
other keys in that object survive — the SDK merge is shallow.

- Responses lane: reasoning_effort joins the caller's extra_body
  "reasoning" object (their keys win) instead of riding a separate
  reasoning= that extra_body's object replaced whole on the wire.
  Mirrors the Messages lane's output_config. The envelope's
  given.get("reasoning") now reports the merged object for free.
- reasoning_effort="" is unset on both protocol lanes, like model=""
  and instructions="".
- extra_body docstring: drop the invitation to override the SDK's
  system/input — that is the skeleton, and a fixed input breaks the
  tool loop; name the wire fields per lane instead of one mixed list.
- messages docstring: every system row joins the managed prompt on the
  answer lane, not only a leading one (_split_chat_messages hoists all).
- max_turns docstring: the OpenAI lanes raise at the cap, the Messages
  lane returns the truncated run — the divergence was documented only on
  the private door.
- Migration message: parameters go by keyword; the doors' sampling and
  thinking fields ride extra_body.
- chat_model docstring: responses is no longer a chat surface.
- _split_chat_messages refusals no longer name chat_completions, a
  method the chat() caller never typed.
- Tests: the Responses door equivalence takes the extra_body shape;
  effort/extra_body collision, key survival and "" pinned on both
  protocol lanes; the ModelSettings spy asserts the values reached the
  wire, not only the envelope.

Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE

* fix: Responses envelope reports the metadata sent

metadata is a Responses request field the caller sets through
extra_body; the envelope hard-coded None. Same source as the other
caller-set fields: given.get().

Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE

* fix: "" is unset on the answer lane; managed-cloud gate reads falsy as unset

reasoning_effort="" reached LiteLLM as a literal empty effort on the
answer lane while the protocol lanes already treated it as unset. And
chat_completions' managed-cloud own-model gate, which chat() routes
through, refused model="" / reasoning_effort="" / {} as knobs the caller
never set. Non-numeric knobs now read falsy as unset, as the local lane
always has (model or chat_model, if extra_body, backend or {}); numeric
ones keep `is not None`.

Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit b15858b156)
2026-09-03 15:52:16 +08:00
Ray a7995a4d22 Ignore the default results/ output directory (#468)
`run_pageindex.py` writes its trees to `./results` (`output_dir = './results'`,
lines 142 and 198), so running the CLI leaves untracked JSON at the repo root.
That output has already been committed by accident ten times — 51 files, 3.2 MB,
across commits from "first commit" through "improve tree optimization".

`examples/documents/results/` is unaffected: this only ignores the top-level
directory the CLI writes to.

Claude-Session: https://claude.ai/code/session_01KXmNuHrkQEkefNS12V7apV
2026-09-03 14:29:25 +08:00
Mingtian Zhang 892f11a297 Merge pull request #464 from VectifyAI/feat/notebook
delete toturials
2026-09-02 18:08:35 +08:00
张鸣天 0dd6982a66 delete toturials 2026-09-02 11:07:06 +01:00
Mingtian Zhang 23d62d6bc2 Merge pull request #463 from VectifyAI/feat/notebook
add pageindex flash notebook
2026-09-02 18:01:39 +08:00
张鸣天 8da1a18e96 add pageindex flash notebook 2026-09-02 10:56:23 +01:00
Mingtian Zhang b064ee7d0e Merge pull request #462 from VectifyAI/zmtomorrow-patch-1
Update PageIndex Flash entry in README
2026-09-02 17:37:06 +08:00
Mingtian Zhang cb34f11b30 Update PageIndex Flash entry in README 2026-09-02 10:36:33 +01:00
Ray 5189cf2892 chat(protocol=): the protocol doors move behind the front door (#460)
* feat: chat(protocol=) — the protocol doors move behind the front door

responses() and messages() become _responses()/_messages(): the same
engines, reachable as chat(protocol="responses"|"messages") with the
protocol's own input and output shapes (transcript items / content
blocks in, the envelope or native stream out). The old names are the
vendor SDKs' own, and an agent-written client.messages(...) now fails
fast with the way in — a runtime-only __getattr__, invisible to the
type checker so attribute typos on the client still get flagged.

chat() gains the knobs that lost their public home: instructions
(appended after the managed prompt — a string on every lane, Messages
system blocks with protocol="messages"), max_turns, backend,
extra_headers, extra_body. reasoning_effort lands natively on each
lane: LiteLLM's kwarg, Responses reasoning.effort, Anthropic
output_config.effort. show_process stays the answer lane's view.

Six overloads keep the return types narrow for py.typed consumers; the
answer lane's contract is unchanged. The managed cloud chat rejects the
own-model knobs as before.

Claude-Session: https://claude.ai/code/session_017uvMD9eatgdaupLGFUjuKM

* fix: chat(protocol=) honors extra_body thinking; honest envelope and remedies

- Messages lane: through the public door the thinking budget rides
  extra_body, so the default max_tokens lift reads it there (extra_body
  wins on the wire, so it wins in the lift); the documented
  extra_body={"thinking": ...} no longer 400s on max_tokens < budget.
- Responses lane: temperature/top_p/reasoning/max_output_tokens reach the
  wire via extra_body only; the envelope now reports what was sent instead
  of the never-set locals.
- One _require_own_chat refusal for chat(protocol=...), the doors behind
  it, and instructions. The doors' own copies had already drifted from
  chat()'s text, and every copy pointed a local client with chat_model
  blank at the managed chat it does not have; a door reached directly on
  such a client fell through to a raw AttributeError. The check is the
  one chat_completions already makes.
- Messages lane treats model="" as unset, like every other model check.
- The protocol refusal for show_process runs before the stream=True hint,
  so the first remedy offered is the right one.
- instructions="" configures nothing (no empty system row, cache key
  unchanged), matching the protocol lanes.
- Responses validation names messages, the public parameter, not input.
- Stale docstring pointer to messages() fixed.
- A protocol=None, stream: bool overload restores str | ChatStream for a
  runtime-variable stream; the catch-all had widened it to a 4-way union.
- Tests: door X refuses like chat(protocol=X) on managed and blank-local
  clients; the protocol gate and the non-stream show_process order are
  asserted by distinct messages; chat(stream=True) joins the max_turns
  matrix; each fix carries a red-verified assertion.

Claude-Session: https://claude.ai/code/session_01M9GdjnuDHHwDKqMzvPWCcj
v0.2.14
2026-09-02 06:59:55 +08:00
Ray 91b6c365e0 Chat process display: labels are the event type names (#459)
* style: label woven lines by their event type names

[tool_call] and [tool_result] replace the "[tool]" label and the "->"
arrow, so the text view's labels are exactly the .events type names
(config keys stay plural — they switch a class of lines; each line is
one instance).

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* style: tool_result lines flush left — no nesting indent

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* style: one vocabulary — show_process keys are the event type names

tool_call / tool_result (singular) everywhere: event types, text labels,
and now the config keys, which select event types by name.

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
2026-09-02 04:28:38 +08:00
Ray ad7956d3c3 Keep litellm's terminal noise out of answers: stdout banners, logger chatter, retry print (#455)
* fix: mute litellm's stdout Provider List banner in the litellm lanes

litellm's OpenRouter adapter probes supports_reasoning() with the
provider-stripped model name on every completion, so any model missing
from its static map (e.g. openrouter/z-ai/glm-5.3-flash) makes
get_llm_provider print a red "Provider List:" banner straight into
stdout — interleaved with the streamed answer, once per agent turn.
suppress_debug_info is litellm's own embedder switch (its Router sets
it too) and gates only this banner and the "Give Feedback / Get Help"
one; errors still raise with their full text. Applied at the same lazy
hook points as the existing litellm repairs, plus the background
preload, so merely importing pageindex still leaves the host's litellm
untouched.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL

* fix: retry notice rides logging, not the caller's stdout

llm_completion/llm_acompletion printed '* Retrying *' straight into
stdout on every retried request — the channel that belongs to answers
and CLI output. The notice moves to logging.warning beside the error
line that already accompanies it.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL

* fix: gate litellm's stderr WARNING chatter alongside the stdout banners

The banner mute grows into _quiet_litellm: litellm's own logger sprays
WARNING records (remote-map fetch fallbacks, cost hiccups) onto stderr
from inside requests — not actionable for SDK callers, whose real
failures raise as exceptions. The preload stamps LITELLM_LOG=ERROR
before litellm's import initializes its logger (setdefault, so an
explicit caller choice wins, and litellm honors a chosen level itself);
the hook's setLevel covers litellm imported before us, plus the dotted
litellm namespace its adapters log under.

Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL
v0.2.13
2026-09-01 22:16:23 +08:00
Ray 555c370f66 Show the chat run: show_process weaving and the ChatStream events view (#454)
* feat: show the chat run — show_process weaving and the ChatStream events view

chat(stream=True) now returns a ChatStream: iterating it yields the
answer text with the run woven in by default — "[thinking] " sections,
one "[tool] name arguments" line per call with its clipped result —
and .events yields the run as typed dicts (thinking/answer deltas,
tool_call with parsed arguments, tool_result with the full output).
One run serves one view; close() kills it like a closed generator.

- show_process: on by default ("on where available"); False for the
  bare answer stream; a dict (ChatProcessOptions: thinking, tool_calls,
  tool_results, max_chars) selects the parts. Explicit True without
  stream=True raises.
- Managed clients weave what the endpoint serves: tool-call lines
  parsed from its block_metadata chunk tags (that wire carries no
  thinking and no tool results); old-wire chunks stay plain answer
  text, and the bare answer view no longer leaks tool-argument JSON.
- Engine: one typed-event primitive (_chat_events_agen /
  _cloud_chunk_events) with _weave as a pure renderer over it; the
  chat lane's prologue is shared via _chat_agent, behavior unchanged
  on chat_completions/responses/messages.

453 tests green (16 new, red-verified), no-openai-agents leg simulated,
pyright flat vs main.

Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP

* fix: pair woven tool results with their calls; drop empty deltas; validate show_process before the managed request

Three review findings on the show_process weave, all red-verified:

- _weave nested every tool result under the most recent [tool] line,
  which misattributes results when a turn makes parallel calls (the SDK
  streams all calls, then all results). A result now nests only when the
  line above is its own call (by call_id); otherwise it stands alone
  with its call's clipped arguments echoed, so same-name parallel calls
  stay tellable apart. Hidden-call mode falls out unchanged (no stored
  arguments, no echo).

- Empty deltas now stop at the event source. chat_completions and the
  managed chunk lane both filter them; the new local lane did not, so a
  mid-stream "" (litellm forwards annotated/provider-field empties)
  leaked into the bare view and flipped _weave sections, splitting one
  thinking burst into repeated labels.

- The managed streaming lane sent the billed request before
  _process_options ran, so a config typo cost a real chat call. chat()
  now chokes on bad show_process before dispatching, matching the local
  lane's validate-first order.

Two docstring truths: close()'s "a run never consumed never starts"
holds only for own-model chat (the managed request is already on the
wire), and the module docstring now covers the managed chunk weave.

456 tests green; the three new ones red-verified; managed-lane paths
re-run with the agents package blocked; pyright adds nothing on touched
lines.

Claude-Session: https://claude.ai/code/session_01XeeD2214z6Vd6qKcAi9ZvJ

* fix: seal open tool blocks from the answer; typed chokes; chat() stream overloads

Four review fixes plus two coverage gaps, each red- or
mutation-verified:

- _cloud_chunk_events treated any non-tool_use tag inside an open tool
  block as answer text, so argument JSON leaked into the
  show_process=False answer — the one meant to be appended back as
  conversation history. Inside an open block nothing is answer:
  argument chunks now accumulate under any tag. And non-string
  argument pieces stringify at the join instead of killing the whole
  stream with a raw TypeError.

- _process_options sorted unknown keys before repr-ing them, so
  mixed-type keys ({1: True, "foo": 1}) raised a bare TypeError past
  the caller's `except PageIndexAPIError`; sorting the reprs keeps the
  single error type.

- chat() gains @overload on stream, so the docstring's own `.events`
  usage type-checks for py.typed consumers (previously pyright ruled
  `Cannot access attribute "events" for class "str"` on the exact
  documented snippet). pageindex/ error count unchanged (234).

- Coverage: the streaming lane's whole `finally` could be deleted with
  the suite still green — the new abandonment test pins the teardown
  (pump exits, turn 2 emits nothing, _aclose_backend closes the
  per-call client). And FakeModel emitted only the reasoning event
  production never sends (litellm folds reasoning into
  reasoning_content, which arrives as summary deltas); it now
  alternates variants, so dropping either from the isinstance tuple
  goes red.

459 tests green; managed-path tests re-run with the agents package
blocked; flake8 parity on every touched file.

Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP

* fix: reading .events is inert — consuming claims the view; honest falsy show_process message

ChatStream.events was a property whose getter latched the stream's one
view on mere attribute access: a debugger variable pane, hasattr, or
getattr(stream, "events", None) — which PageIndexAPIError escapes, as
getattr only swallows AttributeError — was enough to make a later
`for chunk in stream:` refuse, with nothing consumed. The getter now
returns a lazy generator: the managed refusal, the view claim and the
run start all happen on first consumption, so introspection is
side-effect free and the text view stays usable after a probe.

And the stream=False guard's message told falsy-but-not-False values
("show_process=0", "") that they passed show_process=True; the check
itself is the ruled falsy-{} trap and stands, but the message now
names the off values and echoes what was got.

461 tests green (2 red-verified new: inert read on both lanes, plus
the falsy-message case); changed managed-path tests re-run with the
agents package blocked; pyright pageindex/ 234 -> 234.

Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
2026-09-01 22:16:14 +08:00
Ray 7d18cc033c docs: drop stray blank lines in the cloud quickstart snippet
Claude-Session: https://claude.ai/code/session_01Wh75QagUhr5DPVATjmGrCe
2026-08-31 19:00:00 +08:00
Mingtian Zhang 9fdbc7d6ad Merge pull request #451 from VectifyAI/docs/query-cost-chart
improve model name
2026-08-31 18:46:54 +08:00
张鸣天 2d02e8b6e8 improve model name 2026-08-31 11:45:09 +01:00
Mingtian Zhang 72469fb9e2 Merge pull request #449 from VectifyAI/docs/query-cost-chart
fix agent integration
2026-08-31 16:28:32 +08:00
张鸣天 aa153ab1e7 fix agent integration 2026-08-31 09:27:19 +01:00
Mingtian Zhang 635bca4e65 Merge pull request #448 from VectifyAI/docs/query-cost-chart
fix agent integration
2026-08-31 16:25:24 +08:00
张鸣天 d1b9fbd1b9 fix agent integration 2026-08-31 09:24:27 +01:00
Mingtian Zhang 0a3b563d69 Merge pull request #447 from VectifyAI/docs/query-cost-chart
Docs/query cost chart
2026-08-31 16:07:25 +08:00
张鸣天 869c76ad17 fix agent integration 2026-08-31 09:06:13 +01:00
张鸣天 70cc912300 fix agent integration 2026-08-31 09:04:16 +01:00
Mingtian Zhang 4f8414c328 Merge pull request #446 from VectifyAI/docs/query-cost-chart
fix chat
2026-08-31 14:16:35 +08:00
张鸣天 7cc46303d6 fix chat 2026-08-31 07:15:51 +01:00
Mingtian Zhang 75bca66e9b Merge pull request #445 from VectifyAI/docs/query-cost-chart
add image to readme
2026-08-31 09:43:50 +08:00
张鸣天 10e7433ec6 add image to readme 2026-08-31 02:41:09 +01:00
Ray 9fee239b17 docs: correct what the index model does (#441)
* docs: correct what the index model does

The index model does not build the tree structure — Flash extracts it
from the document layout without an LLM. The model only summarizes and
refines the tree.

Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD

* docs: name PageIndex Flash in the submit_document note

Claude-Session: https://claude.ai/code/session_01EtDZekHStmxXNexn95aAeD
v0.2.12
2026-08-28 01:57:30 +08:00
Ray 21f2b1018e docs: FinanceBench chart as local light/dark assets (#440)
* docs: swap the FinanceBench chart for local light/dark assets

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: replace the FinanceBench chart with light/dark assets

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: narrow the FinanceBench chart to 60%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: set the FinanceBench chart to 65%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF

* docs: set the FinanceBench chart to 70%

Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF
2026-08-28 00:56:37 +08:00
Ray 688ce200fa docs: hide the FinanceBench accuracy chart
Claude-Session: https://claude.ai/code/session_012edjBdTjAM24UAZGyq6TfF
2026-08-28 00:39:20 +08:00
Ray d947ab6c93 docs: README structure and layout pass (#438)
* docs: the benchmark charts render at 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render centered at 80% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 85% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 90% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the collapsed sections drop the spacer <br>

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the header link row drops Discord

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Updates entry comments out the MCP/API pointer

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the one-line summary stands without the Why it works heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: TEMP four punchline variants side by side

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the one-line summary reads as a centered pull quote

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: TEMP alert-style variants next to the pull quote

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the pull quote sits under an In one sentence heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the pull quote sits under a TL;DR heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote bolds the claims and drops the vector DB and chunking

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote reads no vector DBs or chunking

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote leaves retrieval unbolded

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Retrieve step leads with agentically

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote reads left-aligned

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the TLDR quote wraps on its own

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example uses a plain system-message string

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example keeps the system prompt inside the code block

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example inlines the system prompt

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Cloud example drops the doubled blank line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench case study returns to Benchmarks

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench paragraph speaks of PageIndex directly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence names the corpus and the margin plainly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence drops the corpus aside

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench sentence keeps the benchmark description

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Benchmarks splits into the local open-source run and FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: top-level sections use single-hash headings again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed Local mode

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed PageIndex Local

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the benchmark charts render at 80% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench chart stays at 90% and the benchmark gloss goes in parentheses

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench chart returns to its original 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark charts render at 70% width

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the local benchmark part is headed Running PageIndex locally, charts at 75%

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations example indents the prompt continuation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations prompt breaks before the cite tag

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the citations prompt breaks before using

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the collapsed usage guides follow the Quickstart

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a Usage section holds the two collapsed guides

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h3 with their steps one level down

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h2 so their steps keep their levels

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage guides sit at h3, inner levels untouched

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: each usage step collapses on its own under Usage

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Usage keeps its two guides as headings with collapsed items inside

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the *_config note folds into Other agent frameworks

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Usage and the Detailed Usage Guide open with a line of orientation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in drops the happy-path idiom

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in names the two guides directly

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the guide and integration lead-ins say what each block holds

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in contrasts direct use with integration

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in lists its two guides; their intros stay general

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in is one line with (i) and (ii)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in says SDK client and your own agent

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in uses parallel verbs

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in says access rather than call

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in drops the verb in (i)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in reads as one plain sentence pair

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the client guide is headed Use PageIndex through the SDK client

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in lists its two ways as bullets

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage lead-in returns to one (i)/(ii) line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Ready to Try It links cover only the noun

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Usage section is headed Usage Guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the summary heading reads tl;dr

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items use bold summaries so they sit close together

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: expanded items get a line of air under their summary

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items try h4 summaries again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: collapsed items settle on bold summaries with a spacer

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a line of air between neighbouring collapsed items

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Step 2 anchor lives inside its summary so item gaps match

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the two guide headings rely on their generated anchors

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide opens with a plainer line

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in drops workflow

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in names its three steps

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in, polished

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the SDK client guide lead-in points below

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: a line of air before the integration heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: no spacer before the integration heading

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the two usage guides are labelled (a) and (b)

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench closing line names the evaluation

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench closing line links only the results

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench claim links straight to its benchmark subsection

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the indexing-cost line names the model in prose

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench subsection is headed Leading on FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the FinanceBench heading reads Leading accuracy on FinanceBench

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Updates lists the PageIndex File System again

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the File System update entry links in plain weight

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the header link row gains Blog

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the summary heading reads TL;DR

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Model Recommendations points at the Usage Guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Model Recommendations names the SDK client usage guide

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the usage-guide pointer promises more than models

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the usage-guide pointer, one word shorter

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in says each example covers one framework

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the integration lead-in says a different framework

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: each framework fold shows the one-call and explicit forms end to end

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the streaming example uses a literal and the quickstart drops trailing spaces

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry says where the structure comes from

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry, tighter

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry says built by an LLM

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry, Ray's wording

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: the Flash update entry keeps extracted heuristically

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k

* docs: Updates dates read [Aug '26]

Claude-Session: https://claude.ai/code/session_01BtbxfFP1wdkFou1FMCMA6k
2026-08-28 00:14:31 +08:00
akong afb5e11976 docs: reduce README banner image size 2026-08-27 18:49:25 +08:00
Mingtian Zhang c346056ae9 Replace PageIndex banner image
Updated the banner image in the README file.
2026-08-27 17:24:56 +08:00
Ray d4ce6ee65f docs: the run_messages comment keeps the principle, drops the version numbers
The rationale to preserve is that execution policy belongs to the
vendor and only the runner's history is ground truth; the exact 1.0/1.1
switch point lives in the previous commit's message.
2026-08-27 03:37:47 +08:00
Ray 070293bdca test: the messages max_tokens test follows anthropic 1.1's stop policy
anthropic 1.1.0 replaced the runner's refusal special case with an
explicit stop-reason table: every non-tool_use stop is terminal and its
tool_use blocks are never executed (1.0 executed a max_tokens turn's
complete blocks). The envelope already keys on the runner's history, so
the product adapts by design; only the test had the 1.0 behavior baked
in. It now asserts the version-appropriate shape on both sides, and the
run_messages comment describing the old behavior is reworded to name
the policy split.

Verified: 434 green on anthropic 1.1.0 (CI's failing config) and on
0.120.2 (the pre-1.1 branch); no-frameworks collection stays clean.
2026-08-27 03:37:47 +08:00
Ray f47f27b7b3 docs: README restructure with quickstart, benchmarks, and a collapsed usage guide (#431)
Quickstart on the v0.2.11 index=/chat= spellings, a citations example, indexing-time benchmarks with new charts, the own-agent integration in its own collapsed block, a rewritten PageIndex Cloud section, the reasoning positioning restored, and prose em dashes replaced with plainer punctuation. Earlier snapshots of this branch landed via #418/#419; this squash carries the tail.
2026-08-27 03:28:17 +08:00
Ray 31f01910cd docs: API-key URLs follow the dashboard move to developer.pageindex.ai
dash.pageindex.ai/api-keys now 307-redirects to
developer.pageindex.ai/api-keys, and the README (PR #431) already
standardized on the developer domain — the client docstring and the
CloudClient keyless error message were the last SDK references to the
old host.
2026-08-27 02:54:46 +08:00
Ray 57252e2686 fix: the LiteLLM lanes hide litellm's bridge usage warning (#430)
Streaming chat() on an OpenAI gpt-5.4+ model with function tools makes
litellm re-route chat.completions through the Responses API. When the
stream ends, litellm's logging stores a chat-shaped usage dict inside a
ResponseAPIUsage field and model_dump()s it, so pydantic prints
"Expected `ResponseAPIUsage` - serialized value may not be as expected"
once per streamed turn. litellm does this on purpose
(litellm_logging._get_assembled_streaming_response, 1.97 and 1.98) and
the answer is unaffected.

Both seams that hand a LiteLLM-routed model to openai-agents — the chat
lanes' model builder and openai_agent_config()'s litellm/ lane —
register one warnings filter matching exactly that message; every other
warning still surfaces. A process-wide filter is the only placement
that works: litellm emits the warning from its async success handler
on the worker thread, out of catch_warnings' reach.

Claude-Session: https://claude.ai/code/session_01APJbbp3jpRxJhvchcU2Aed
2026-08-26 20:03:34 +08:00
Ray 174f95f35b fix: the slots accept any Mapping at runtime; eight docstrings stop calling cloud tool scoping server-side (#429)
fix: the slots accept any Mapping at runtime, as their annotation admits; eight docstrings stop calling cloud tool scoping server-side

4e9c56c widened index=/chat= to Mapping[str, Any] so the exported
TypedDicts pass a checker, but _resolve_index_slot/_resolve_chat_slot
still dispatched on isinstance(..., dict): a MappingProxyType or ChainMap
was pyright-clean and raised "must be a string or a dict" at construction.
The resolvers now narrow on Mapping — the comprehension already copies,
so a read-only proxy proves the caller's mapping is never mutated.

4e9c56c corrected three of eleven "scoping is server-side" sites; the
remaining eight said the same untrue thing about doc_id on cloud (its
tools carry no allowlist — targeting is prompt-level, as the runtime
error already explains). Deleted rather than reworded.

local_chat.py's module docstring predates own-model chat over the cloud
bridge; storage_path's prose now names the PathLike 5e2dc9b typed.

Claude-Session: https://claude.ai/code/session_01VQ6mruXZBgw9Hjii8KPbQP
2026-08-26 17:00:41 +08:00
Ray b9a9a3b5aa fix: the .env search ends at the cwd tree; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike (#428)
* fix: .env stays unset when the cwd tree has none; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike

find_dotenv(usecwd=True) returns '' when nothing is reachable from the
cwd, and `or None` turned that into load_dotenv's own upward walk from
utils.py — the install-dir leak the cwd search was added to replace. A
pip-installed SDK could load another project's .env from above
site-packages, silently.

_local_chat treats a blank chat_model as "managed chat", which a client
without an api_key does not have: chat_completions() then reached for
LocalAPI.chat_completions and raised a bare AttributeError. The managed
branch now refuses as a PageIndexAPIError naming chat_model.

py.typed made the annotations authoritative while storage_path was typed
str; _ARG_TYPES accepts os.PathLike, so Path(...) ran fine and failed the
user's type check. Both signatures and LocalIndexConfig now say so.

Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb

* fix: the exported config shapes pass into index=/chat=; a comment and two docstrings stop overclaiming

The slots were annotated dict[str, Any]. A TypedDict is consistent with
Mapping[str, object], never with dict (PEP 589: a dict-typed receiver
could write arbitrary keys through it), so the four shapes types.py
exports — and py.typed advertises to installed callers' checkers —
could not be passed to the one place they describe. pyright on a probe
that does exactly that: 9 errors before, 0 after. The constructor only
reads the slot (items(), then a fresh conf dict), so Mapping is the
honest bound; a plain dict is a Mapping, and TypedDict instances are
plain dicts at runtime, so nothing moves at runtime.

The _ARG_TYPES comment said "every value" is shape-checked; api_key is
not in the table (its empty check is separate, its type check stays
unchecked by ruling), so the comment now speaks for the table only.

_local_doc_scope and _require_local_scope still explained the cloud
drop as "scoping is server-side" — true of the managed chat, which
never reaches either function. What reaches them on a cloud client is
own-model chat and the config helpers, whose cloud tools take no
allowlist: targeting there is prompt-level only, as the error message
between them already said.

434 passed; pyright on pageindex/ unchanged at 235 (0 in the touched
files, before and after).

Claude-Session: https://claude.ai/code/session_01TxG8u8x29XRnK4yscZVCch

* test: the install-dir .env test is named for what it asserts

Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb
2026-08-26 15:08:34 +08:00
Ray 920db2b1b1 feat: the client grows two sides — documents and chat each pick their home (#424)
* feat: the client grows two sides — documents and chat each pick their home

One client, two independent switches: api_key decides where documents
live (the PageIndex cloud, or the local store); a configured chat model
decides who answers (your own model in your process, or the managed
cloud chat). Their free combination opens the bridge — cloud documents,
your model — and the fourth cell stays unspellable.

- index=/chat= slots: string shorthand or grouped dict, 1:1 with the
  flat arguments; one spelling per side, sides mix freely
- optional "type" everywhere (top-level and in either dict): always
  omittable, checked against the content, meaningful alone —
  type="cloud" is a keyless cloud spelling
- PAGEINDEX_API_KEY is read only when the code explicitly says cloud
  (PageIndexCloudClient(), type="cloud", "pageindex-cloud",
  {"type": "cloud"}); a bare PageIndexClient() stays local
- bare mode words ("cloud", "local", …) are reserved: they error
  with the real spellings instead of silently parsing as model names
- bridge chat runs the in-process agent over the live cloud MCP tools
  and instructions; doc_id targets at the prompt level; citations stay
  managed-only; an auth-shaped backend failure explains whose
  credentials run the model
- typed shapes (IndexConfig, ChatConfig) ship as optional annotations

Every previously working program is byte-for-byte unchanged: the only
behavioral deltas are error paths — reworded guidance, and the
api_key+chat_model combination graduating from an error into the
bridge.

* fix: the constructor refuses empty and mistyped values on every spelling

- .env keys reach all four keyless-cloud spellings: utils' import-time
  load_dotenv now runs before every PAGEINDEX_API_KEY read
- an empty chat-side value ("", {}) errors instead of silently selecting
  own-model chat on the default model; None-valued slot keys mean absent,
  exactly like the flat arguments
- _local_chat derives from chat_model, so a post-construction assignment
  switches the whole client, never half of it
- model= beside a slot gets the split guidance (index_model=/chat_model=)
  instead of "two spellings of the same thing"
- the messages door wraps provider failures through _model_backend_error,
  and 401s count as auth-shaped even without "api key" in the text
- keyless-cloud hints name the spelling that actually combines; slot
  strings are stripped; wrong-typed values raise PageIndexAPIError
- retrieve_model/chat_backend docs drop the stale "Local mode only";
  the local-scope refusal no longer claims bridge tools are server-scoped

* fix: type= cross-checks the index slot; the cloud pinned class frees its chat side

- type= beside index= now does what the docstring promises: agreement
  passes, disagreement errors, and a mistyped value reports the
  vocabulary error instead of a spelling collision
- PageIndexCloudClient grows the chat-side arguments (chat=, chat_model,
  retrieve_model, chat_backend), so "pin the index side" is literally
  true and the chat surfaces' construct-with-chat_model guidance is
  followable on it
- the four chat doors' doc_id entries carry the enforcement split the
  config helpers already state (local: tool-layer allowlist; cloud:
  prompt-level / server-side)
- types.py stops claiming slot keys share the flat names — the side
  prefix is factored out, index={"model"} is index_model=

* docs: bridge-reachable wording — dependency errors say own-model chat, hints name a chat= model

- the three framework-missing errors said "in local mode", which is
  wrong on a bridge client (cloud documents + own model) — they now
  explain the dependency the way the surfaces do: your own chat model
- the construct-with guidance reads "(or a chat= model)": a bare
  chat="pageindex-cloud" is also chat= but selects the managed side
- the mechanical Local-only → Own-model-chat-only substitution left
  orphan fragments and two overlong lines; those paragraphs re-flowed

* fix: managed chat reads None; the bridge stops paying per-turn tool lists

- a managed-chat cloud client stores chat_model/chat_backend as None, so
  the documented attribute reads instead of raising AttributeError;
  _local_chat derives from "is a chat model configured"
- McpBridge caches tools/list per session — every chat turn rebuilds the
  tool set, and the round trip was pure latency; the 404 session-expiry
  reset drops the cache with the session
- run_messages builds tools before the transport: on a bridge client
  that build is network I/O, and a failure there stranded a per-call
  anthropic client ahead of the try/finally

* refactor: the side declaration is spelled mode=, not type=

"type" is Python's own word — a builtin, and "data type" beside the
TypedDict shapes; "mode" is what the SDK already calls the two sides
("local mode", "cloud mode"). Same grammar everywhere the declaration
appears: the top-level argument, the index dict, the chat dict, the
typed shapes. The rename also frees the builtin inside the constructor,
so the shape-check error names the offending class through type() again.
"type" in a slot dict is now an ordinary unknown key.

* fix: the reserved-word errors stop calling "cloud" not a mode word

With the declaration key spelled mode=, 'index="cloud" is not a mode
word' contradicted its own remedy, index={"mode": "cloud"} — "cloud" is
exactly a mode value. The four bare strings are reserved words; the
message now says so.

* fix: a blank tools/list is not cached; the auth note's managed exit is chat-lane only

- McpBridge.list_tools caches only a non-empty list — a transient blank
  (a deploy blip, a gate misconfiguration) would otherwise run every later
  turn with zero tools while the instructions still name them, and only a
  404 session reset could clear it
- the 401 architecture note appends "drop the chat model configuration"
  only on the chat lane: responses() and messages() refuse a client
  without an own model, so on those lanes the exit sent the caller in a
  circle
- CloudIndexConfig says api_key is omittable only while mode: "cloud"
  stays — index={} refuses as an empty dict rather than reading the env

* fix: the bridge fetches tools/list per call again; .env resolves from the cwd

- McpBridge.list_tools no longer caches: the tool set is built once per
  SDK call (Agent(tools=...) ahead of Runner.run; build_anthropic_tools
  ahead of tool_runner), not per model turn, so the cache saved one round
  trip per later call while a mid-pagination 404 replayed a dead cursor
  into a duplicated (and cached) list, and the list went out by reference
  across a lock dropped between miss and store
- utils.load_dotenv searches upward from the cwd: a bare load_dotenv()
  walked up from utils.py, which is site-packages for an installed SDK,
  so the four keyless-cloud spellings never saw a project-root .env; the
  package-relative walk stays as the fallback
- the emptiness guard strips strings: chat_model=" " selected own-model
  chat, the silent flip the guard's own comment rules out
- the _local_chat comment stops advertising post-construction assignment
  as a full mode switch

* fix: the pinned classes take index=/chat=; "cloud"/"local" are mode words; a blank chat_model stays managed

- PageIndexLocalClient takes index= and chat=, PageIndexCloudClient takes
  index= — the grouped spelling of the flat vocabulary each already took;
  their refusals name the class and an exit that class can take, and the
  mode cross-check runs before any environment read
- "cloud" and "local" are accepted wherever "pageindex-cloud" was (index=,
  chat=, mode=, {"mode": ...}), case- and whitespace-insensitive; "hosted"
  and "managed" still refuse, pointing at the real word
- every spelling strips its strings, and the slot spellings' type/empty
  errors name the slot key (index["model"]), not the flat argument
- _local_chat treats a blank chat_model as managed: the constructor
  refuses "", so assignment agrees instead of opening the bridge on a
  nameless model; openai_agent_config carries no model then either
- an empty MCP tools/list raises like empty instructions does — a
  zero-tool agent would answer from the model's own knowledge silently
- enable_citations names the real gate (managed vs own chat), not
  "cloud-only", on a cloud own-model client
- pageindex/py.typed: the exported config TypedDicts reach installed
  type-checked callers

* test: the two framework-door tests skip without openai-agents

as_openai_tools() and openai_agent_config() need the agents package, which
the "without frameworks" CI legs do not install — the same importorskip
every other test on those doors already carries.
v0.2.11
2026-08-26 04:13:33 +08:00
Ray 416e304f51 perf: expand schedules dependency-exact at thirty-two concurrent proposals (#422)
* perf: expand proposes a wave of nodes concurrently

The expand loop awaited one propose_children at a time — 20-30 nodes at
~3s each put 1-3 minutes of pure round-trip latency on every default
local submit. Nodes waiting in a wave are all frontier leaves whose
decisions cannot affect each other, so the model half now runs
concurrently (EXPAND_CONCURRENCY = 8) while the apply half stays serial
in wave order: decisions, log entries, and child ids land exactly as
before, and children attach into the next wave. A fatal classification
still aborts the run right after the wave's gather.

Benchmarked on real PDFs with a fixed-latency fake model: 408 pages
21.1s -> 3.0s, 758 pages 28.2s -> 3.5s (7-8x); final trees byte-identical
to the serial pass on both. The cap stays low on purpose: expand treats
an exhausted retry ladder as fatal, and a wide burst on a rate-limited
account would trip exactly that — 8 already collapses minutes to seconds.

* perf: expand schedules dependency-exact instead of in waves

A child's only prerequisite is its own parent's apply, so each kept
node gathers its children directly rather than waiting for its whole
generation to finish. Same recursive shape as summarize_tree; the
semaphore still caps in-flight proposals at 8; trees are unchanged.

* perf: expand admits thirty-two concurrent proposals

Cap sweeps on six real documents put the speed plateau at 32: the
ready frontier tops out at 21-28 nodes on few-hundred-page PDFs, so
64 buys nothing while doubling the burst. Live runs at 32 cut the
expand phase 24-30% on the two documents wide enough to feel it,
with zero ladder retries anywhere - and summaries already burst
twice as wide through the same ladder.
2026-08-22 14:49:01 +08:00
Ray 8289729aff fix: indexing failures surface loud, chat envelopes append verbatim (#421)
Indexing: dead credentials or a missing model fail the run instead of
storing a document with blank summaries; a 400 (context_length_exceeded)
skips the retry ladder — the prompt will not shrink — and stays a
per-prompt failure the run absorbs; all-empty model replies can no longer
store a retrieval-ready document; the one-sentence doc description
absorbs its own context overflow instead of discarding a fully indexed
document; the heading-less flash refusal points at mode='standard'.

Chat: messages() output is append-verbatim clean — unset response-only
defaults are dropped (no "caller": null the request schema rejects);
Claude cache marks follow the wire routing; model_settings and name are
openai_agent_config parameters; one Anthropic client per backend; lifted
thinking defaults are clamped to the model's output ceiling from
LiteLLM's capability map.

Store and inputs: lone surrogates are scrubbed from page text and the
stored basename, so the returned name is byte-for-byte the stored name
and the rename warning fires; NaN/Infinity metadata is rejected at the
gate; every cloud error now carries its HTTP status.

CLI: the flash lane resolves the summary model through ConfigLoader like
the standard and markdown lanes; an empty flash structure errors like the
SDK instead of writing "structure": [] with exit 0; --summary-model
reaches the markdown lane; the SDK page-spec surface keeps 0.2.10's
whitespace tolerance while the tool layer stays strict.

pypdfium2 stays on the 5.x line for every install; the 4.x code paths are
tested compatibility insurance with their own CI leg; process-pool
construction failure falls back to the sequential parse; a py3.10 GC
flake in text extraction is fixed.


Port of feat/local-chat 0667e3b..1993740 (28 commits); README and assets untouched.
2026-08-21 21:35:45 +08:00