* Tool results reach the frameworks as MCP content; the frameworks render it
get_document_image returns an MCP image block, and the SDK flattened every
tool result to text on its way to a framework, so an own-model chat received
"[image/png content omitted: ~256 KB]" where the hosted lanes (HostedMCPTool,
the http MCP config, the managed chat) received the page.
The rule now: the SDK carries MCP types and renders nothing. Each framework
gets the tool set the way it consumes MCP, and converts the content itself.
- McpBridge.call_tool returns the content blocks untouched; the invoke
contract behind _tool_specs is (content blocks, is_error) on the cloud and
local paths alike (local tools send one text block).
- openai-agents: an in-process MCPServer over the tool specs, and the
FunctionTools are MCPUtil.to_function_tool over it. The framework's own
MCP conversion builds the tools (schema verbatim, strict off) and renders
results (text items, image data URLs); its failure pipeline answers
malformed arguments. The hand-built FunctionTools are gone.
- Anthropic tool runner: results validated as CallToolResult and rendered by
the Anthropic SDK's mcp_content (text blocks, base64 image blocks), on the
error channel too.
- Claude Agent SDK: the content passes through, it is MCP already.
- Plain functions (agent_tools): render_text, the former flattening, keeps
the size stub for binary content; a string cannot carry an image.
Wire consequences: tool results on the openai-agents lanes are the
framework's structured items (a single text item for a text result) rather
than a bare string, and the Anthropic tool_result content is a block list.
ChatStream tool_result events carry that structured output; the woven
display shows its text.
mcp becomes a declared dependency: it already arrives with openai-agents,
and both adapters use its types.
Verified live on a 15-page paper, page 1, across chat() on the LiteLLM lane
(gpt-5.6-luna), chat(protocol="responses"), and chat(protocol="messages",
claude-sonnet-4-6): each answer described the red attribution text and the
arXiv stamp rotated along the left margin.
Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
* Relay test asserts the wire aliases: mcp 1.x and 2.x differ in attribute names
CI installs mcp 2.x, where CallToolResult exposes is_error and content
blocks exposes mime_type as attributes, with the wire names (isError,
mimeType) as aliases; mcp 1.x uses the wire names as attributes. The
adapters only ever validate from and dump to the wire shape, so they run on
both; the test now does the same.
Claude-Session: https://claude.ai/code/session_01F8QMbFKngRT4TpBAKfuNVW
pyproject publishes README.md as the PyPI description, and PyPI does not
resolve a relative path against the GitHub repo — the five charts render
as broken images on the project page (verified on 0.2.14). Every chart
now loads from raw.githubusercontent.com; the ten other images in the
file already did.
PyPI's sanitizer drops <source>, so only the <img> fallbacks matter
there; the srcset URLs change too, to keep GitHub's dark variants
spelled the same way.
* fix: flip litellm's banner switch at the provider lookups, not after them
_openai_agent asks litellm which provider serves a model twice
(_openai_protocol, _cache_extra_args) before _openai_model flips
suppress_debug_info, and openai_agent_config's litellm/ branch never
flipped it at all; a managed cloud client runs no preload, so each
failed lookup still print()ed the "Provider List:" banner. Quiet at the
two get_llm_provider call sites instead, which covers every lane.
Also: LITELLM_LOG at ERROR or above now clamps to that level (CRITICAL
used to skip the clamp and end up noisier than unset); the retry notice
is logged only when a retry follows; test_preload_stamps_litellm_log_level
no longer leaks LITELLM_LOG=ERROR into the session; the config-time
repair test covers _quiet_litellm with the preload thread pinned off.
Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9
* refactor: the preload thread only imports — every litellm entry quiets itself now
Since 30652ca each get_llm_provider caller and every completion helper
flips the banner switch and clamps litellm's logger before touching
litellm, so the background import's own _quiet_litellm() had no path
left that depended on it; it was also the one place that mutated
process-global logging at a moment the caller could not predict. The
env stamps stay: import-time records still need them. The config-time
repair test no longer needs the preload guard pinned.
Claude-Session: https://claude.ai/code/session_01Gq5Kwi7wEgQNVT4k1XUhq9
* fix: litellm quiet is a default, not a clamp — and the no-fetch stamp covers every lane
_quiet_litellm() set litellm's loggers to ERROR on every entry, overriding whatever the host, litellm._turn_on_debug(), or LITELLM_LOG=DEBUG had set. It now sets a level only while a logger is still NOTSET: the default stays ERROR, anything set explicitly wins, and there is no per-call setLevel churn.
LITELLM_LOCAL_MODEL_COST_MAP moves from the client's preload to utils' import. The CLI pipeline and openai_agent_config on a cloud client import litellm without the preload and fetched the remote model map on first use (offline: a WARNING). Every litellm lane imports utils first.
The LITELLM_LOG=ERROR stamp is gone: it pinned litellm's stderr handler at ERROR for the whole process (silencing _turn_on_debug for good) while the logger level was what actually gated the chatter — output verified identical with and without it across plain, logging-configured, and debug hosts.
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* style: essential comments only
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* test: the utils-import stamp check lives with its package-surface sibling
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* refactor: drop the CLI's duplicate no-fetch stamp — utils' import sets it
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* fix: the thinking-budget ceiling imported litellm ahead of the stamp
_default_max_tokens is reached from anthropic_runner_config/messages(). On a managed cloud client the constructor never imports pageindex.utils, and local_chat deliberately keeps utils off its import path, so that lane imported litellm with LITELLM_LOCAL_MODEL_COST_MAP unset: the model cost map was fetched over the network (3517 entries vs the bundled 2982) and offline printed the very warning this branch removes. Verified end to end through the public API, before and after.
Claude-Session: https://claude.ai/code/session_01CrSg9uRBvzVujNyU25Wmgk
* fix: extra_body refuses the skeleton keys
extra_body merges last on every lane, so a caller's system / instructions / input / messages / tools replaced the managed prompt, the conversation or the doc tools wholesale. The run succeeded and answered without tool guidance or document scoping, silently. Those are the SDK's on every lane (the three-layer rule: skeleton, named knobs, caller extras); the door for the prompt is instructions=. Refused at the two seams where extra_body meets the wire, before the Anthropic transport exists on the Messages lane.
Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit 4d18854378)
* fix: extra_body sampling fields ride their ModelSettings field on LiteLLM-routed models
On the answer lane, extra_body for a non-OpenAI destination becomes ModelSettings.extra_args, the LiteLLM kwargs channel (LiteLLM would plant extra_body as a literal request field, which Anthropic rejects). openai-agents passes ModelSettings' own fields to litellm.acompletion by name beside **extra_args, so a caller's temperature / top_p / max_tokens / penalties / tool_choice in extra_args collided with the explicit keyword and died as a raw TypeError inside the framework. Split by the framework's own contract: keys that are ModelSettings fields ride their field (the caller's value winning over ours), the rest stay LiteLLM kwargs. The field list is the public dataclass, not a hand-kept name list; bad values now fail ModelSettings validation and surface as PageIndexAPIError. response_format is not a ModelSettings field and still has no door on this lane; documented.
Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit c1ad980505)
* docs: chat_completions extra_body names the skeleton carve-out
chat_completions shares the seam that now refuses the skeleton keys; its docstring still promised an unconditional merged-last win. Same sentence as chat().
Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit f9072f9fff)
* fix: ChatStream lives in a light module, so chat()'s type hints resolve at runtime
ChatStream was importable in client.py only under TYPE_CHECKING (a real
import would have dragged local_chat's asyncio stack into `import
pageindex`), which left chat()'s return annotation a dangling string:
typing.get_type_hints(PageIndexClient.chat) raised NameError, and so
did anything that introspects signatures — agents' function_tool
(client.chat) died on it before looking at a single parameter. The
class touches neither asyncio nor the agent frameworks, so it moves to
pageindex/chat_stream.py, client.py imports it for real, and the
package exports it directly instead of lazily. `import pageindex` still
leaves local_chat unloaded.
Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY
* fix: the ChatStream move keeps its own invariants — future annotations, a guarded eager path, the old import path pinned
Review of #471 found the move's guards thinner than they look:
- chat_stream.py had no `from __future__ import annotations`, unlike
every sibling module, which made its `-> "ChatStream"` quotes
load-bearing: unquoting them — the very edit this move made in
client.py, and what `ruff --select UP037 --fix` does — broke
`import pageindex` outright.
- The import in client.py is the whole fix and reads like a
typing-only one; a comment says why it must stay real.
- The lazy-import test's denylist named no framework, so agents,
litellm, openai or anthropic could join the eager path with a green
suite — the cost the module was split out to avoid.
- The type-hints walk had no floor, so it could silently stop covering
anything, and nothing pinned `pageindex.local_chat.ChatStream`, the
path the class shipped under in 0.2.11-0.2.14.
- local_chat's module docstring still claimed the class.
All four guards mutation-checked red.
* style: the walk-collapsed assertion message fits the 79-col convention
* style: no rationale comments — the guard is the test, the why is the commit
* fix: .events survives partial reads; hidden call lines still label results; bad show_process chokes first
- ChatStream.events delegated with `yield from`, so a dropped handle
(next(stream.events), for ... break) closed the shared run on GC and
the rest of the run silently vanished. A plain loop leaves it alone.
- _weave filled call_args only past the tool_call visibility guard, so
with call lines hidden the standalone result lines never carried the
arguments they promise.
- show_process is validated before the stream check: an invalid value
is refused as such instead of being told to add stream=True and then
refused again; the managed lane's duplicate choke goes with it.
- Docstring: show_process is not own-model-only.
Claude-Session: https://claude.ai/code/session_016M3qaQedSK7L4DwysFRmk2
* fix: mid-stream error chunk raises instead of ending as a short answer; stream docstring says show_process is on by default
- The managed endpoint reports a server-side failure as a final
{"error": ...} chunk after the partial answer (api.py refunds the
credits, then yields it). Neither chunk decoder looked at it, so
chat(stream=True) and chat_completions(stream=True) in both modes
ended as an apparently complete short answer with no exception. One
guard in each decoder raises PageIndexAPIError; the partial answer
is still delivered first.
- The `stream:` arg and the Returns block still described the pre-PR
contract (bare text chunks); only the show_process paragraph said it
is on by default.
Claude-Session: https://claude.ai/code/session_01PYr9yG1FPQxKCA9m7ECQWY
* fix: Responses effort rides extra_body; chat() docs say what the wire does
The rule chat() follows, written down: chat() names PageIndex's own
parameters plus a subset of LiteLLM's unified vocabulary; anything
vendor-specific rides extra_body under the lane's wire names. Three
layers meet on the wire — the skeleton (managed prompt, conversation,
tools) is the SDK's, the named knobs are translated per lane, and
extra_body is the caller's, merged last so it wins. A key the SDK writes
into a nested object must go in through extra_body too, so the caller's
other keys in that object survive — the SDK merge is shallow.
- Responses lane: reasoning_effort joins the caller's extra_body
"reasoning" object (their keys win) instead of riding a separate
reasoning= that extra_body's object replaced whole on the wire.
Mirrors the Messages lane's output_config. The envelope's
given.get("reasoning") now reports the merged object for free.
- reasoning_effort="" is unset on both protocol lanes, like model=""
and instructions="".
- extra_body docstring: drop the invitation to override the SDK's
system/input — that is the skeleton, and a fixed input breaks the
tool loop; name the wire fields per lane instead of one mixed list.
- messages docstring: every system row joins the managed prompt on the
answer lane, not only a leading one (_split_chat_messages hoists all).
- max_turns docstring: the OpenAI lanes raise at the cap, the Messages
lane returns the truncated run — the divergence was documented only on
the private door.
- Migration message: parameters go by keyword; the doors' sampling and
thinking fields ride extra_body.
- chat_model docstring: responses is no longer a chat surface.
- _split_chat_messages refusals no longer name chat_completions, a
method the chat() caller never typed.
- Tests: the Responses door equivalence takes the extra_body shape;
effort/extra_body collision, key survival and "" pinned on both
protocol lanes; the ModelSettings spy asserts the values reached the
wire, not only the envelope.
Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE
* fix: Responses envelope reports the metadata sent
metadata is a Responses request field the caller sets through
extra_body; the envelope hard-coded None. Same source as the other
caller-set fields: given.get().
Claude-Session: https://claude.ai/code/session_01Hj6t26s7thUjkcho6sn4zE
* fix: "" is unset on the answer lane; managed-cloud gate reads falsy as unset
reasoning_effort="" reached LiteLLM as a literal empty effort on the
answer lane while the protocol lanes already treated it as unset. And
chat_completions' managed-cloud own-model gate, which chat() routes
through, refused model="" / reasoning_effort="" / {} as knobs the caller
never set. Non-numeric knobs now read falsy as unset, as the local lane
always has (model or chat_model, if extra_body, backend or {}); numeric
ones keep `is not None`.
Claude-Session: https://claude.ai/code/session_01BAmVWYKoSnFEydbMjuZjHc
(cherry picked from commit b15858b156)
`run_pageindex.py` writes its trees to `./results` (`output_dir = './results'`,
lines 142 and 198), so running the CLI leaves untracked JSON at the repo root.
That output has already been committed by accident ten times — 51 files, 3.2 MB,
across commits from "first commit" through "improve tree optimization".
`examples/documents/results/` is unaffected: this only ignores the top-level
directory the CLI writes to.
Claude-Session: https://claude.ai/code/session_01KXmNuHrkQEkefNS12V7apV
* feat: chat(protocol=) — the protocol doors move behind the front door
responses() and messages() become _responses()/_messages(): the same
engines, reachable as chat(protocol="responses"|"messages") with the
protocol's own input and output shapes (transcript items / content
blocks in, the envelope or native stream out). The old names are the
vendor SDKs' own, and an agent-written client.messages(...) now fails
fast with the way in — a runtime-only __getattr__, invisible to the
type checker so attribute typos on the client still get flagged.
chat() gains the knobs that lost their public home: instructions
(appended after the managed prompt — a string on every lane, Messages
system blocks with protocol="messages"), max_turns, backend,
extra_headers, extra_body. reasoning_effort lands natively on each
lane: LiteLLM's kwarg, Responses reasoning.effort, Anthropic
output_config.effort. show_process stays the answer lane's view.
Six overloads keep the return types narrow for py.typed consumers; the
answer lane's contract is unchanged. The managed cloud chat rejects the
own-model knobs as before.
Claude-Session: https://claude.ai/code/session_017uvMD9eatgdaupLGFUjuKM
* fix: chat(protocol=) honors extra_body thinking; honest envelope and remedies
- Messages lane: through the public door the thinking budget rides
extra_body, so the default max_tokens lift reads it there (extra_body
wins on the wire, so it wins in the lift); the documented
extra_body={"thinking": ...} no longer 400s on max_tokens < budget.
- Responses lane: temperature/top_p/reasoning/max_output_tokens reach the
wire via extra_body only; the envelope now reports what was sent instead
of the never-set locals.
- One _require_own_chat refusal for chat(protocol=...), the doors behind
it, and instructions. The doors' own copies had already drifted from
chat()'s text, and every copy pointed a local client with chat_model
blank at the managed chat it does not have; a door reached directly on
such a client fell through to a raw AttributeError. The check is the
one chat_completions already makes.
- Messages lane treats model="" as unset, like every other model check.
- The protocol refusal for show_process runs before the stream=True hint,
so the first remedy offered is the right one.
- instructions="" configures nothing (no empty system row, cache key
unchanged), matching the protocol lanes.
- Responses validation names messages, the public parameter, not input.
- Stale docstring pointer to messages() fixed.
- A protocol=None, stream: bool overload restores str | ChatStream for a
runtime-variable stream; the catch-all had widened it to a 4-way union.
- Tests: door X refuses like chat(protocol=X) on managed and blank-local
clients; the protocol gate and the non-stream show_process order are
asserted by distinct messages; chat(stream=True) joins the max_turns
matrix; each fix carries a red-verified assertion.
Claude-Session: https://claude.ai/code/session_01M9GdjnuDHHwDKqMzvPWCcj
* style: label woven lines by their event type names
[tool_call] and [tool_result] replace the "[tool]" label and the "->"
arrow, so the text view's labels are exactly the .events type names
(config keys stay plural — they switch a class of lines; each line is
one instance).
Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
* style: tool_result lines flush left — no nesting indent
Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
* style: one vocabulary — show_process keys are the event type names
tool_call / tool_result (singular) everywhere: event types, text labels,
and now the config keys, which select event types by name.
Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
* fix: mute litellm's stdout Provider List banner in the litellm lanes
litellm's OpenRouter adapter probes supports_reasoning() with the
provider-stripped model name on every completion, so any model missing
from its static map (e.g. openrouter/z-ai/glm-5.3-flash) makes
get_llm_provider print a red "Provider List:" banner straight into
stdout — interleaved with the streamed answer, once per agent turn.
suppress_debug_info is litellm's own embedder switch (its Router sets
it too) and gates only this banner and the "Give Feedback / Get Help"
one; errors still raise with their full text. Applied at the same lazy
hook points as the existing litellm repairs, plus the background
preload, so merely importing pageindex still leaves the host's litellm
untouched.
Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL
* fix: retry notice rides logging, not the caller's stdout
llm_completion/llm_acompletion printed '* Retrying *' straight into
stdout on every retried request — the channel that belongs to answers
and CLI output. The notice moves to logging.warning beside the error
line that already accompanies it.
Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL
* fix: gate litellm's stderr WARNING chatter alongside the stdout banners
The banner mute grows into _quiet_litellm: litellm's own logger sprays
WARNING records (remote-map fetch fallbacks, cost hiccups) onto stderr
from inside requests — not actionable for SDK callers, whose real
failures raise as exceptions. The preload stamps LITELLM_LOG=ERROR
before litellm's import initializes its logger (setdefault, so an
explicit caller choice wins, and litellm honors a chosen level itself);
the hook's setLevel covers litellm imported before us, plus the dotted
litellm namespace its adapters log under.
Claude-Session: https://claude.ai/code/session_018psbiPrxdiCqFsS7Tk3eFL
* feat: show the chat run — show_process weaving and the ChatStream events view
chat(stream=True) now returns a ChatStream: iterating it yields the
answer text with the run woven in by default — "[thinking] " sections,
one "[tool] name arguments" line per call with its clipped result —
and .events yields the run as typed dicts (thinking/answer deltas,
tool_call with parsed arguments, tool_result with the full output).
One run serves one view; close() kills it like a closed generator.
- show_process: on by default ("on where available"); False for the
bare answer stream; a dict (ChatProcessOptions: thinking, tool_calls,
tool_results, max_chars) selects the parts. Explicit True without
stream=True raises.
- Managed clients weave what the endpoint serves: tool-call lines
parsed from its block_metadata chunk tags (that wire carries no
thinking and no tool results); old-wire chunks stay plain answer
text, and the bare answer view no longer leaks tool-argument JSON.
- Engine: one typed-event primitive (_chat_events_agen /
_cloud_chunk_events) with _weave as a pure renderer over it; the
chat lane's prologue is shared via _chat_agent, behavior unchanged
on chat_completions/responses/messages.
453 tests green (16 new, red-verified), no-openai-agents leg simulated,
pyright flat vs main.
Claude-Session: https://claude.ai/code/session_014GN8u3zdH3RpeftHZChavP
* fix: pair woven tool results with their calls; drop empty deltas; validate show_process before the managed request
Three review findings on the show_process weave, all red-verified:
- _weave nested every tool result under the most recent [tool] line,
which misattributes results when a turn makes parallel calls (the SDK
streams all calls, then all results). A result now nests only when the
line above is its own call (by call_id); otherwise it stands alone
with its call's clipped arguments echoed, so same-name parallel calls
stay tellable apart. Hidden-call mode falls out unchanged (no stored
arguments, no echo).
- Empty deltas now stop at the event source. chat_completions and the
managed chunk lane both filter them; the new local lane did not, so a
mid-stream "" (litellm forwards annotated/provider-field empties)
leaked into the bare view and flipped _weave sections, splitting one
thinking burst into repeated labels.
- The managed streaming lane sent the billed request before
_process_options ran, so a config typo cost a real chat call. chat()
now chokes on bad show_process before dispatching, matching the local
lane's validate-first order.
Two docstring truths: close()'s "a run never consumed never starts"
holds only for own-model chat (the managed request is already on the
wire), and the module docstring now covers the managed chunk weave.
456 tests green; the three new ones red-verified; managed-lane paths
re-run with the agents package blocked; pyright adds nothing on touched
lines.
Claude-Session: https://claude.ai/code/session_01XeeD2214z6Vd6qKcAi9ZvJ
* fix: seal open tool blocks from the answer; typed chokes; chat() stream overloads
Four review fixes plus two coverage gaps, each red- or
mutation-verified:
- _cloud_chunk_events treated any non-tool_use tag inside an open tool
block as answer text, so argument JSON leaked into the
show_process=False answer — the one meant to be appended back as
conversation history. Inside an open block nothing is answer:
argument chunks now accumulate under any tag. And non-string
argument pieces stringify at the join instead of killing the whole
stream with a raw TypeError.
- _process_options sorted unknown keys before repr-ing them, so
mixed-type keys ({1: True, "foo": 1}) raised a bare TypeError past
the caller's `except PageIndexAPIError`; sorting the reprs keeps the
single error type.
- chat() gains @overload on stream, so the docstring's own `.events`
usage type-checks for py.typed consumers (previously pyright ruled
`Cannot access attribute "events" for class "str"` on the exact
documented snippet). pageindex/ error count unchanged (234).
- Coverage: the streaming lane's whole `finally` could be deleted with
the suite still green — the new abandonment test pins the teardown
(pump exits, turn 2 emits nothing, _aclose_backend closes the
per-call client). And FakeModel emitted only the reasoning event
production never sends (litellm folds reasoning into
reasoning_content, which arrives as summary deltas); it now
alternates variants, so dropping either from the isinstance tuple
goes red.
459 tests green; managed-path tests re-run with the agents package
blocked; flake8 parity on every touched file.
Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
* fix: reading .events is inert — consuming claims the view; honest falsy show_process message
ChatStream.events was a property whose getter latched the stream's one
view on mere attribute access: a debugger variable pane, hasattr, or
getattr(stream, "events", None) — which PageIndexAPIError escapes, as
getattr only swallows AttributeError — was enough to make a later
`for chunk in stream:` refuse, with nothing consumed. The getter now
returns a lazy generator: the managed refusal, the view claim and the
run start all happen on first consumption, so introspection is
side-effect free and the text view stays usable after a probe.
And the stream=False guard's message told falsy-but-not-False values
("show_process=0", "") that they passed show_process=True; the check
itself is the ruled falsy-{} trap and stands, but the message now
names the off values and echoes what was got.
461 tests green (2 red-verified new: inert read on both lanes, plus
the falsy-message case); changed managed-path tests re-run with the
agents package blocked; pyright pageindex/ 234 -> 234.
Claude-Session: https://claude.ai/code/session_01DWBCCTDzwuVamBf5MvQ4eP
The rationale to preserve is that execution policy belongs to the
vendor and only the runner's history is ground truth; the exact 1.0/1.1
switch point lives in the previous commit's message.
anthropic 1.1.0 replaced the runner's refusal special case with an
explicit stop-reason table: every non-tool_use stop is terminal and its
tool_use blocks are never executed (1.0 executed a max_tokens turn's
complete blocks). The envelope already keys on the runner's history, so
the product adapts by design; only the test had the 1.0 behavior baked
in. It now asserts the version-appropriate shape on both sides, and the
run_messages comment describing the old behavior is reworded to name
the policy split.
Verified: 434 green on anthropic 1.1.0 (CI's failing config) and on
0.120.2 (the pre-1.1 branch); no-frameworks collection stays clean.
Quickstart on the v0.2.11 index=/chat= spellings, a citations example, indexing-time benchmarks with new charts, the own-agent integration in its own collapsed block, a rewritten PageIndex Cloud section, the reasoning positioning restored, and prose em dashes replaced with plainer punctuation. Earlier snapshots of this branch landed via #418/#419; this squash carries the tail.
dash.pageindex.ai/api-keys now 307-redirects to
developer.pageindex.ai/api-keys, and the README (PR #431) already
standardized on the developer domain — the client docstring and the
CloudClient keyless error message were the last SDK references to the
old host.
Streaming chat() on an OpenAI gpt-5.4+ model with function tools makes
litellm re-route chat.completions through the Responses API. When the
stream ends, litellm's logging stores a chat-shaped usage dict inside a
ResponseAPIUsage field and model_dump()s it, so pydantic prints
"Expected `ResponseAPIUsage` - serialized value may not be as expected"
once per streamed turn. litellm does this on purpose
(litellm_logging._get_assembled_streaming_response, 1.97 and 1.98) and
the answer is unaffected.
Both seams that hand a LiteLLM-routed model to openai-agents — the chat
lanes' model builder and openai_agent_config()'s litellm/ lane —
register one warnings filter matching exactly that message; every other
warning still surfaces. A process-wide filter is the only placement
that works: litellm emits the warning from its async success handler
on the worker thread, out of catch_warnings' reach.
Claude-Session: https://claude.ai/code/session_01APJbbp3jpRxJhvchcU2Aed
fix: the slots accept any Mapping at runtime, as their annotation admits; eight docstrings stop calling cloud tool scoping server-side
4e9c56c widened index=/chat= to Mapping[str, Any] so the exported
TypedDicts pass a checker, but _resolve_index_slot/_resolve_chat_slot
still dispatched on isinstance(..., dict): a MappingProxyType or ChainMap
was pyright-clean and raised "must be a string or a dict" at construction.
The resolvers now narrow on Mapping — the comprehension already copies,
so a read-only proxy proves the caller's mapping is never mutated.
4e9c56c corrected three of eleven "scoping is server-side" sites; the
remaining eight said the same untrue thing about doc_id on cloud (its
tools carry no allowlist — targeting is prompt-level, as the runtime
error already explains). Deleted rather than reworded.
local_chat.py's module docstring predates own-model chat over the cloud
bridge; storage_path's prose now names the PathLike 5e2dc9b typed.
Claude-Session: https://claude.ai/code/session_01VQ6mruXZBgw9Hjii8KPbQP
* fix: .env stays unset when the cwd tree has none; a local client with a blank chat_model refuses at the chat door; storage_path is typed PathLike
find_dotenv(usecwd=True) returns '' when nothing is reachable from the
cwd, and `or None` turned that into load_dotenv's own upward walk from
utils.py — the install-dir leak the cwd search was added to replace. A
pip-installed SDK could load another project's .env from above
site-packages, silently.
_local_chat treats a blank chat_model as "managed chat", which a client
without an api_key does not have: chat_completions() then reached for
LocalAPI.chat_completions and raised a bare AttributeError. The managed
branch now refuses as a PageIndexAPIError naming chat_model.
py.typed made the annotations authoritative while storage_path was typed
str; _ARG_TYPES accepts os.PathLike, so Path(...) ran fine and failed the
user's type check. Both signatures and LocalIndexConfig now say so.
Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb
* fix: the exported config shapes pass into index=/chat=; a comment and two docstrings stop overclaiming
The slots were annotated dict[str, Any]. A TypedDict is consistent with
Mapping[str, object], never with dict (PEP 589: a dict-typed receiver
could write arbitrary keys through it), so the four shapes types.py
exports — and py.typed advertises to installed callers' checkers —
could not be passed to the one place they describe. pyright on a probe
that does exactly that: 9 errors before, 0 after. The constructor only
reads the slot (items(), then a fresh conf dict), so Mapping is the
honest bound; a plain dict is a Mapping, and TypedDict instances are
plain dicts at runtime, so nothing moves at runtime.
The _ARG_TYPES comment said "every value" is shape-checked; api_key is
not in the table (its empty check is separate, its type check stays
unchecked by ruling), so the comment now speaks for the table only.
_local_doc_scope and _require_local_scope still explained the cloud
drop as "scoping is server-side" — true of the managed chat, which
never reaches either function. What reaches them on a cloud client is
own-model chat and the config helpers, whose cloud tools take no
allowlist: targeting there is prompt-level only, as the error message
between them already said.
434 passed; pyright on pageindex/ unchanged at 235 (0 in the touched
files, before and after).
Claude-Session: https://claude.ai/code/session_01TxG8u8x29XRnK4yscZVCch
* test: the install-dir .env test is named for what it asserts
Claude-Session: https://claude.ai/code/session_017Fd7jVm366S2Xamzhxv6yb
* feat: the client grows two sides — documents and chat each pick their home
One client, two independent switches: api_key decides where documents
live (the PageIndex cloud, or the local store); a configured chat model
decides who answers (your own model in your process, or the managed
cloud chat). Their free combination opens the bridge — cloud documents,
your model — and the fourth cell stays unspellable.
- index=/chat= slots: string shorthand or grouped dict, 1:1 with the
flat arguments; one spelling per side, sides mix freely
- optional "type" everywhere (top-level and in either dict): always
omittable, checked against the content, meaningful alone —
type="cloud" is a keyless cloud spelling
- PAGEINDEX_API_KEY is read only when the code explicitly says cloud
(PageIndexCloudClient(), type="cloud", "pageindex-cloud",
{"type": "cloud"}); a bare PageIndexClient() stays local
- bare mode words ("cloud", "local", …) are reserved: they error
with the real spellings instead of silently parsing as model names
- bridge chat runs the in-process agent over the live cloud MCP tools
and instructions; doc_id targets at the prompt level; citations stay
managed-only; an auth-shaped backend failure explains whose
credentials run the model
- typed shapes (IndexConfig, ChatConfig) ship as optional annotations
Every previously working program is byte-for-byte unchanged: the only
behavioral deltas are error paths — reworded guidance, and the
api_key+chat_model combination graduating from an error into the
bridge.
* fix: the constructor refuses empty and mistyped values on every spelling
- .env keys reach all four keyless-cloud spellings: utils' import-time
load_dotenv now runs before every PAGEINDEX_API_KEY read
- an empty chat-side value ("", {}) errors instead of silently selecting
own-model chat on the default model; None-valued slot keys mean absent,
exactly like the flat arguments
- _local_chat derives from chat_model, so a post-construction assignment
switches the whole client, never half of it
- model= beside a slot gets the split guidance (index_model=/chat_model=)
instead of "two spellings of the same thing"
- the messages door wraps provider failures through _model_backend_error,
and 401s count as auth-shaped even without "api key" in the text
- keyless-cloud hints name the spelling that actually combines; slot
strings are stripped; wrong-typed values raise PageIndexAPIError
- retrieve_model/chat_backend docs drop the stale "Local mode only";
the local-scope refusal no longer claims bridge tools are server-scoped
* fix: type= cross-checks the index slot; the cloud pinned class frees its chat side
- type= beside index= now does what the docstring promises: agreement
passes, disagreement errors, and a mistyped value reports the
vocabulary error instead of a spelling collision
- PageIndexCloudClient grows the chat-side arguments (chat=, chat_model,
retrieve_model, chat_backend), so "pin the index side" is literally
true and the chat surfaces' construct-with-chat_model guidance is
followable on it
- the four chat doors' doc_id entries carry the enforcement split the
config helpers already state (local: tool-layer allowlist; cloud:
prompt-level / server-side)
- types.py stops claiming slot keys share the flat names — the side
prefix is factored out, index={"model"} is index_model=
* docs: bridge-reachable wording — dependency errors say own-model chat, hints name a chat= model
- the three framework-missing errors said "in local mode", which is
wrong on a bridge client (cloud documents + own model) — they now
explain the dependency the way the surfaces do: your own chat model
- the construct-with guidance reads "(or a chat= model)": a bare
chat="pageindex-cloud" is also chat= but selects the managed side
- the mechanical Local-only → Own-model-chat-only substitution left
orphan fragments and two overlong lines; those paragraphs re-flowed
* fix: managed chat reads None; the bridge stops paying per-turn tool lists
- a managed-chat cloud client stores chat_model/chat_backend as None, so
the documented attribute reads instead of raising AttributeError;
_local_chat derives from "is a chat model configured"
- McpBridge caches tools/list per session — every chat turn rebuilds the
tool set, and the round trip was pure latency; the 404 session-expiry
reset drops the cache with the session
- run_messages builds tools before the transport: on a bridge client
that build is network I/O, and a failure there stranded a per-call
anthropic client ahead of the try/finally
* refactor: the side declaration is spelled mode=, not type=
"type" is Python's own word — a builtin, and "data type" beside the
TypedDict shapes; "mode" is what the SDK already calls the two sides
("local mode", "cloud mode"). Same grammar everywhere the declaration
appears: the top-level argument, the index dict, the chat dict, the
typed shapes. The rename also frees the builtin inside the constructor,
so the shape-check error names the offending class through type() again.
"type" in a slot dict is now an ordinary unknown key.
* fix: the reserved-word errors stop calling "cloud" not a mode word
With the declaration key spelled mode=, 'index="cloud" is not a mode
word' contradicted its own remedy, index={"mode": "cloud"} — "cloud" is
exactly a mode value. The four bare strings are reserved words; the
message now says so.
* fix: a blank tools/list is not cached; the auth note's managed exit is chat-lane only
- McpBridge.list_tools caches only a non-empty list — a transient blank
(a deploy blip, a gate misconfiguration) would otherwise run every later
turn with zero tools while the instructions still name them, and only a
404 session reset could clear it
- the 401 architecture note appends "drop the chat model configuration"
only on the chat lane: responses() and messages() refuse a client
without an own model, so on those lanes the exit sent the caller in a
circle
- CloudIndexConfig says api_key is omittable only while mode: "cloud"
stays — index={} refuses as an empty dict rather than reading the env
* fix: the bridge fetches tools/list per call again; .env resolves from the cwd
- McpBridge.list_tools no longer caches: the tool set is built once per
SDK call (Agent(tools=...) ahead of Runner.run; build_anthropic_tools
ahead of tool_runner), not per model turn, so the cache saved one round
trip per later call while a mid-pagination 404 replayed a dead cursor
into a duplicated (and cached) list, and the list went out by reference
across a lock dropped between miss and store
- utils.load_dotenv searches upward from the cwd: a bare load_dotenv()
walked up from utils.py, which is site-packages for an installed SDK,
so the four keyless-cloud spellings never saw a project-root .env; the
package-relative walk stays as the fallback
- the emptiness guard strips strings: chat_model=" " selected own-model
chat, the silent flip the guard's own comment rules out
- the _local_chat comment stops advertising post-construction assignment
as a full mode switch
* fix: the pinned classes take index=/chat=; "cloud"/"local" are mode words; a blank chat_model stays managed
- PageIndexLocalClient takes index= and chat=, PageIndexCloudClient takes
index= — the grouped spelling of the flat vocabulary each already took;
their refusals name the class and an exit that class can take, and the
mode cross-check runs before any environment read
- "cloud" and "local" are accepted wherever "pageindex-cloud" was (index=,
chat=, mode=, {"mode": ...}), case- and whitespace-insensitive; "hosted"
and "managed" still refuse, pointing at the real word
- every spelling strips its strings, and the slot spellings' type/empty
errors name the slot key (index["model"]), not the flat argument
- _local_chat treats a blank chat_model as managed: the constructor
refuses "", so assignment agrees instead of opening the bridge on a
nameless model; openai_agent_config carries no model then either
- an empty MCP tools/list raises like empty instructions does — a
zero-tool agent would answer from the model's own knowledge silently
- enable_citations names the real gate (managed vs own chat), not
"cloud-only", on a cloud own-model client
- pageindex/py.typed: the exported config TypedDicts reach installed
type-checked callers
* test: the two framework-door tests skip without openai-agents
as_openai_tools() and openai_agent_config() need the agents package, which
the "without frameworks" CI legs do not install — the same importorskip
every other test on those doors already carries.
* perf: expand proposes a wave of nodes concurrently
The expand loop awaited one propose_children at a time — 20-30 nodes at
~3s each put 1-3 minutes of pure round-trip latency on every default
local submit. Nodes waiting in a wave are all frontier leaves whose
decisions cannot affect each other, so the model half now runs
concurrently (EXPAND_CONCURRENCY = 8) while the apply half stays serial
in wave order: decisions, log entries, and child ids land exactly as
before, and children attach into the next wave. A fatal classification
still aborts the run right after the wave's gather.
Benchmarked on real PDFs with a fixed-latency fake model: 408 pages
21.1s -> 3.0s, 758 pages 28.2s -> 3.5s (7-8x); final trees byte-identical
to the serial pass on both. The cap stays low on purpose: expand treats
an exhausted retry ladder as fatal, and a wide burst on a rate-limited
account would trip exactly that — 8 already collapses minutes to seconds.
* perf: expand schedules dependency-exact instead of in waves
A child's only prerequisite is its own parent's apply, so each kept
node gathers its children directly rather than waiting for its whole
generation to finish. Same recursive shape as summarize_tree; the
semaphore still caps in-flight proposals at 8; trees are unchanged.
* perf: expand admits thirty-two concurrent proposals
Cap sweeps on six real documents put the speed plateau at 32: the
ready frontier tops out at 21-28 nodes on few-hundred-page PDFs, so
64 buys nothing while doubling the burst. Live runs at 32 cut the
expand phase 24-30% on the two documents wide enough to feel it,
with zero ladder retries anywhere - and summaries already burst
twice as wide through the same ladder.
Indexing: dead credentials or a missing model fail the run instead of
storing a document with blank summaries; a 400 (context_length_exceeded)
skips the retry ladder — the prompt will not shrink — and stays a
per-prompt failure the run absorbs; all-empty model replies can no longer
store a retrieval-ready document; the one-sentence doc description
absorbs its own context overflow instead of discarding a fully indexed
document; the heading-less flash refusal points at mode='standard'.
Chat: messages() output is append-verbatim clean — unset response-only
defaults are dropped (no "caller": null the request schema rejects);
Claude cache marks follow the wire routing; model_settings and name are
openai_agent_config parameters; one Anthropic client per backend; lifted
thinking defaults are clamped to the model's output ceiling from
LiteLLM's capability map.
Store and inputs: lone surrogates are scrubbed from page text and the
stored basename, so the returned name is byte-for-byte the stored name
and the rename warning fires; NaN/Infinity metadata is rejected at the
gate; every cloud error now carries its HTTP status.
CLI: the flash lane resolves the summary model through ConfigLoader like
the standard and markdown lanes; an empty flash structure errors like the
SDK instead of writing "structure": [] with exit 0; --summary-model
reaches the markdown lane; the SDK page-spec surface keeps 0.2.10's
whitespace tolerance while the tool layer stays strict.
pypdfium2 stays on the 5.x line for every install; the 4.x code paths are
tested compatibility insurance with their own CI leg; process-pool
construction failure falls back to the sequential parse; a py3.10 GC
flake in text extraction is fixed.
Port of feat/local-chat 0667e3b..1993740 (28 commits); README and assets untouched.