* Document targeting is conversation content: document_context() replaces doc_id on the agent surfaces
The doc-targeting block ("The user has specified document: ...") was
appended to the system prompt by agent_instructions() and the three
*_agent_config bundles, and placed as its own system block by the Messages
chat lane. A per-request target in the system prompt breaks the cached
prefix, sticks across turns, and on cloud splices SDK text onto the
server-served instructions. The cloud's managed chat puts the same block at
the head of the conversation, so every surface now does the same.
- chat(): all three lanes prepend the block as the first user message
(the Messages lane moved off its system block).
- New client.document_context(doc_id): the block text for callers who own
the conversation (the framework routes), to lead their first message.
- BREAKING: doc_id removed from agent_instructions(), openai_agent_config(),
anthropic_runner_config(), claude_agent_config(), agent_tools(),
as_openai_tools(), as_anthropic_tools() and as_claude_mcp(). The
tool-layer allowlist was local-only and raised on cloud; it stays
internal to local chat(doc_id=).
- The block is one get_document per id: names are unique per library
(uploads suffix a taken name), so the shadow check and its listing sweep
went. The cloud endpoint carries metadata, so local get_document() gained
the key for parity. Rendering matches the cloud's: one document is an
object, several a list.
Live-verified on a local store and the cloud library across chat() (answer
lane, chat_completions, responses, messages) and the OpenAI Agents,
Anthropic tool_runner and Claude Agent SDK routes.
Claude-Session: https://claude.ai/code/session_01VxguPoTv2BmzS9d3erzTrS
* Agent surfaces: keyword-only tails so a stale positional doc_id raises
doc_id sat first in the positional list on agent_instructions,
openai_agent_config and claude_agent_config, and second on
anthropic_runner_config and as_claude_mcp. With it removed, a v0.2.15 call
such as openai_agent_config("pi-a") no longer failed: the id landed on
include_management (truthy, so the management tools came along and
targeting silently vanished) or on server_name. A bare * after the
surviving leading positional turns those calls into an immediate
TypeError; every in-repo caller already passes keywords.
Claude-Session: https://claude.ai/code/session_017FumozBm2xbT2SG6WBxjMe
* Fix the delegate build_claude_mcp missed by the keyword-only tails
5451e4e put a `*` on the client methods that dropped doc_id, so a stale
positional call raises instead of landing on the next parameter.
build_claude_mcp, which as_claude_mcp delegates to, dropped doc_ids the
same way and let server_name move into its slot: a stale
build_claude_mcp(client, False, ['pi-a']) built an in-process server
named ['pi-a'], or on cloud returned the http config with the argument
dropped. Now keyword-only, covered by the same test.
Doc truths: document_context no longer claims a later turn can move on
to another document (chat's docstrings say to keep doc_id constant, and
the block is prepended at the head of the conversation, so a retarget
lands behind the history); submit_document's metadata inventory and
CloudAPI.get_document's key list both name the metadata and folderId
keys get_document returns.
Docstring trims: the "names are unique per library" rationale sentence
and the multi-line test docstrings, one of which still described the
listing backfill this PR removed.
Claude-Session: https://claude.ai/code/session_018FoZJ97cuM9hAjrk3goZ6u
* Folders on chat: chat(folder_id=) and folder_context()
The managed /chat/completions already takes folder_id and renders a
folder targeting block ahead of the document block (pageindex-compute
api.py:491, :3182); the SDK had no way to send it, and own-model chat
had no folder analogue of document_context(). Both land here.
folder_context(folder_id) renders that block as the managed chat
renders its own: the name, an id/name/description metadata row (the
server's FolderContext), and the directive to discover the folder's
documents with browse_documents / search_documents(folder_id=...,
recursive=true). One list_folders() call picks the folder by id: the
endpoint has no LIMIT and lists the same userId + api_surface
partition the managed chat validates against, so what the helper finds
is what the server accepts, and it carries the same plan gate, so a
non-Max key fails here rather than mid-run. A missing folder raises.
Cloud-only: local libraries have no folders.
chat(folder_id=) places it. The managed lane sends the field and lets
the server render and gate it; the three own-model lanes lead the
conversation with one user message, folder block then document block,
joined as the server joins them. folder_id joins the prompt-cache key
for the reason doc_id does: the block is byte-identical for every
conversation about a folder.
"root" (and "") targets nothing on every lane and in every mode, as
the server treats folder_id="root": folder_context returns "" and chat
places nothing. folder_id is keyword-only on chat() and trailing on
every other signature, so no positional call shifts.
Tests: the wire field on the managed lane; block order and root/"" on
the shared renderer; folder_id threaded through the chat, responses
and messages engines; the local cloud-only refusal.
Claude-Session: https://claude.ai/code/session_01NsQjN5HMmmX4n5NVsZWYaw
* Keep folder-less cache keys stable; folder_id keyword-only everywhere
_conversation_cache_key splices folder_id into the seed only when one
is set, so a conversation without a folder keeps the prompt_cache_key it
had before folders existed instead of cold-starting on upgrade; the
released key is pinned.
folder_id gets the bare * on chat_completions, _responses and _messages,
keyword-only on every chat surface as on chat(); every caller already
passes it by keyword.
The three *_agent_config docstrings name the context helpers in the
order the lanes place them: folder, then documents.
test_doc_targeting_is_one_lookup_per_document now counts get_document
calls; it accepted a second lookup per id before.
Claude-Session: https://claude.ai/code/session_01T2i8viYf3Cb6wZcjToLEDD
* Rebase onto the chat_completions protocol lane: folder_id joins its exact payload
The lane's payload assertion is exact, and folder_id is a managed chat
request field, so it appears there too — unset on a folder-less call.
* folder_id rides the chat_completions protocol lane too
chat(protocol="chat_completions") reached main while this branch was
open, and it is a managed chat_completions call site of its own: a
folder_id passed there was dropped on the way to the endpoint, scoping
nothing with no error. It carries it now, asserted on the wire beside
the answer lane's.
PageIndex: Vectorless, Reasoning-based RAG
Reasoning-based RAG ◦ No Vector DB, No Chunking ◦ Context-Aware Retrieval ◦ Reads Like a Human
🌐 Website • ☁️ Cloud • 📖 Docs • 📝 Blog • ✉️ Contact
Updates
- [Aug '26] 🔥 PageIndex SDK:
pip install -U pageindexnow ships local mode: index, retrieve, and chat entirely on your machine with your own LLM key, or point the same client at PageIndex Cloud with an API key. - [Aug '26] ⚡ PageIndex Flash: fast tree index generation for text-based PDFs, now the default indexing method in PageIndex SDK local mode.
- Scale PageIndex to Millions of Documents: PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document.
- PageIndex App: a human-like document analysis agent for long professional documents.
What is PageIndex?
Are you frustrated with vector database retrieval accuracy for long and complex documents? Vector-based RAG retrieves by semantic similarity. But similarity ≠ relevance — what retrieval actually needs is relevance, and relevance requires reasoning. On professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search misses what is relevant but not similar, and returns what is similar but not relevant.
Inspired by AlphaGo, PageIndex replaces the vector index with a hierarchical tree index and lets an LLM reason its way through it, the way a human expert turns to and reads the right section of a long report. Retrieval happens in two steps:
- Index: generate a tree-structure index for each document
- Retrieve: agentically search that tree with LLM reasoning
TL;DR
PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, with no vector DBs or chunking.
Compare with Vector RAG
| Vector RAG | PageIndex | |
|---|---|---|
| Index | vector index | tree index |
| Retrieval | semantic similarity search | LLM reasoning over the tree |
| Result | opaque, “vibe retrieval” | traceable to explicit references |
| Context | query embedding only | full context: conversation history, domain knowledge, etc. |
It is ideal for financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any other long, complex professional document.
Quickstart
pip install -U pageindex
import os
from pageindex import PageIndexClient
os.environ["OPENAI_API_KEY"] = "your-openai-key"
client = PageIndexClient(
index="gpt-5.6-luna", # model to build the tree index
chat="gpt-5.6-sol", # model to search the tree
)
doc_id = client.submit_document("report.pdf")["doc_id"]
answer = client.chat("What was the 2023 operating margin?", doc_id=doc_id)
print(answer)
Model Recommendations
index=: a basic model is sufficient. The tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it, which a basic model does well.chat=: use the best model you can afford. The chat model searches the tree to retrieve information. See Query cost and accuracy.
Use PageIndex through the SDK client →
Configure other models, streaming, multi-document search, citations, and more.
Integrate PageIndex with your own agent →
Drop PageIndex tools into the OpenAI Agents SDK, the Claude Agent SDK, or any other framework.
Benchmarks
Local indexing cost and time
Building a tree locally runs about $0.001 per page with gpt-5.6-luna as the index model, so a 1,000-page textbook costs a little over a dollar and a few minutes, once, and every later question reuses it. PageIndex is designed not to rely heavily on the model used at index time, so in our experiments a basic model does not hurt quality.
Indexing time also scales predictably with document length. In the same local setup, the benchmark documents (9 to 1,098 pages) finished in roughly 13 seconds to 4.5 minutes.
Query cost and accuracy
PageIndex-OSS-Benchmark measures exactly the setup in the quickstart above (PageIndexClient() in local mode, flash indexing, no OCR) on 62 lookup questions over 34 PDFs (1,945 pages) drawn from MMLongBench-Doc-V2. Every question's answer is a fact stated in running text, so a wrong answer is a retrieval or reading failure, not a reasoning one.
Full results, data, and the runner are in the benchmark repo.
Cost per query vs. native PDF input
The alternative to retrieval is handing the model the whole PDF on every question. That cost grows with the document; PageIndex's does not, because it reads only the nodes its reasoning reaches. On documents where both routes return the same answer, native PDF input costs 2.1× more at 52 pages and 16.6× more at 420 (gpt-5.6-sol, prompt caching excluded) — and at 805 pages the document no longer fits in the context window at all.
Leading accuracy on FinanceBench
PageIndex reached a state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector-based RAG.
Explore the full FinanceBench evaluation results and the blog post.
PageIndex Cloud
The open-source version is ideal for text-heavy PDFs and local workflows. With PageIndex Cloud, document indexing and storage run in the cloud: PageIndex handles parsing, OCR, image understanding, tree-index construction, and managed storage for you. The chat and retrieval layer remains compatible with your model, so you can search the cloud-hosted index using the model provider your application already uses.
Moving indexing and storage from Local to Cloud only requires a PageIndex API key:
import os
from pageindex import PageIndexClient
os.environ["PAGEINDEX_API_KEY"] = "your-pageindex-key"
os.environ["OPENAI_API_KEY"] = "your-openai-key"
client = PageIndexClient(
index="cloud", # build and store the index in PageIndex Cloud
chat="gpt-5.6-sol", # use your preferred compatible model for chat
)
doc_id = client.submit_document("report.pdf", wait=True)["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))
| Capability | Local (this repo) | Cloud (get an API key) |
|---|---|---|
| Best for | text-heavy PDFs and local workflows | scanned, image-heavy, and large document collections |
| Indexing | runs locally | runs in PageIndex Cloud, with production OCR and image understanding |
| Storage | local | managed in PageIndex Cloud |
| Chat model | your model | your model, or the managed chat included with your key |
| Citations | page-level | line-level |
| Image understanding | — | ✅ |
| Multi-document scale | manual | PageIndex File System |
| MCP server | — | ✅ |
More About PageIndex Cloud
- Scale PageIndex to Millions of Documents: PageIndex File System is a Cloud-only, file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document.
Ready to Try It?
- Get a PageIndex API key
- Read the PageIndex Cloud documentation
For dedicated deployment (VPC or on-premises), contact us or book a demo.
⭐ Support Us
Leave us a star 🌟 if you like our project. Thank you!
Please cite this work as:
Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
Or use the BibTeX citation.
@article{zhang2025pageindex,
author = {Mingtian Zhang and Yu Tang and PageIndex Team},
title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
journal = {PageIndex Blog},
year = {2025},
month = {September},
note = {https://pageindex.ai/blog/pageindex-intro},
}
© 2026 PageIndex AI



