* Block citations resolve to bounding boxes: get_block() and resolve_citations()
resolve_citations(answer, doc_ids=None) reads the <cite doc= page= block=/>
and <doc=…;page=…;block=…> tags of a cited answer and returns one entry per
distinct citation with the document id and, for block-level citations on
cloud documents, the block's page, bbox, type and text from the new
GET /doc/{doc_id}/block/{block_id}/ route, wrapped as get_block(). Local
mode stays page-level; a name shared by two documents raises instead of
guessing; a block the document lacks keeps its entry without a bbox.
* resolve_citations: paired quotes, tag-edged names, tolerant block lookups
_CITE_ATTR_RE pairs its quotes, so an apostrophe inside a quoted name
(Moody's Outlook.pdf) stays in the name instead of ending it.
_OLD_CITATION_RE stops a name at < or >, so a <doc=name> tag without its
own ';' no longer swallows the text up to the next tag's ';' and that
citation with it.
The block lookup keeps its bbox-less entry on a status-less raise (local
mode, where pages have no blocks) and on 403, as doc_targeting_block
does; a transport failure still propagates with its status.
The library listing is _all_documents(), which tolerates an absent
'total' and ends on the empty page; ids are deduplicated before the
collision check, so a document re-served after an upload shifted the
window is one document, not a collision with itself.
doc_ids=[] fails loud like chat(doc_id=[]) instead of resolving nothing.
Docstrings: the bbox coordinate space on both get_block(); the
shared-document promise dropped, since the metadata route is owner-only.
* Citation tag scan is linear
_CITE_TAG_RE's second alternative (the paired form) was unreachable: the
first alternative matches the opening tag whenever a '>' follows, and
without one neither alternative can match. Dropping it, stopping the tag
body at '<' and letting the group absorb the whitespace makes an
unterminated '<cite ' scan linear: '<cite ' plus 4000 spaces went from
2 minutes to 0.1 ms, 99 KB of unterminated tags from 27 minutes to
about a millisecond. Output is identical on every shape tried, the
paired form included.
* Finish the tag-edge and listing-entry sweep
Two holes of the kinds already closed, left open one line and one loop
away from their siblings.
The old format's block group still crossed tag boundaries the way its
name group used to: "<doc=a.pdf;page=3;block=p3_text_5 <cite doc=..." put
the whole following citation inside block_id and, since a block lookup
that 404s is now kept without a bbox, lost it silently. It stops at the
tag's edge like the name beside it; a block id has never contained < or
>, the server writes p{page}_{type}_{seq}.
The library listing reads entries with .get() like the rest of the
codebase does: _all_documents() tolerates a listing whose shape varies,
so reading its result with raw subscripts put the KeyError back one
level down. An entry without a name or an id cannot answer a citation
either way, so it is skipped.
* resolve_citations: block ids trimmed, linear attribute scan, doc_id= like chat()
add() trims both values itself: the name, which each caller used to
trim, and the block id, which neither did. "block=p3_text_5 >" and
block=" p3_text_5 " now look up p3_text_5. A padded id 404s, and a 404
is kept without a bbox, so the padding lost the bbox silently and handed
the padded id back to the caller.
_CITE_ATTR_RE anchors the attribute name at a word boundary. Without
\b the greedy (\w+) restarted at every character of a long word inside
a tag body: '<cite ' + 64 KB of x + '>' cost 19 s on the caller's
thread. Anchored, 64 KB is 1.6 ms and 1 MB 27 ms. Output is identical:
the leftmost match of (\w+)= always begins at a word boundary (0 diffs
on 30,000 random inputs). The linear-time test now reaches this
expression; its inputs had no '>' and never did.
The scope parameter is doc_id, the name chat(), chat_completions() and
document_context() give the same str-or-list argument, and the one the
docstring already used to describe it. The integration builders keep
their doc_ids.
PageIndex: Vectorless, Reasoning-based RAG
Reasoning-based RAG ◦ No Vector DB, No Chunking ◦ Context-Aware Retrieval ◦ Reads Like a Human
🌐 Website • ☁️ Cloud • 📖 Docs • 📝 Blog • ✉️ Contact
Updates
- [Aug '26] 🔥 PageIndex SDK:
pip install -U pageindexnow ships local mode: index, retrieve, and chat entirely on your machine with your own LLM key, or point the same client at PageIndex Cloud with an API key. - [Aug '26] ⚡ PageIndex Flash: fast tree index generation for text-based PDFs, now the default indexing method in PageIndex SDK local mode.
- Scale PageIndex to Millions of Documents: PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document.
- PageIndex App: a human-like document analysis agent for long professional documents.
What is PageIndex?
Are you frustrated with vector database retrieval accuracy for long and complex documents? Vector-based RAG retrieves by semantic similarity. But similarity ≠ relevance — what retrieval actually needs is relevance, and relevance requires reasoning. On professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search misses what is relevant but not similar, and returns what is similar but not relevant.
Inspired by AlphaGo, PageIndex replaces the vector index with a hierarchical tree index and lets an LLM reason its way through it, the way a human expert turns to and reads the right section of a long report. Retrieval happens in two steps:
- Index: generate a tree-structure index for each document
- Retrieve: agentically search that tree with LLM reasoning
TL;DR
PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, with no vector DBs or chunking.
Compare with Vector RAG
| Vector RAG | PageIndex | |
|---|---|---|
| Index | vector index | tree index |
| Retrieval | semantic similarity search | LLM reasoning over the tree |
| Result | opaque, “vibe retrieval” | traceable to explicit references |
| Context | query embedding only | full context: conversation history, domain knowledge, etc. |
It is ideal for financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any other long, complex professional document.
Quickstart
pip install -U pageindex
import os
from pageindex import PageIndexClient
os.environ["OPENAI_API_KEY"] = "your-openai-key"
client = PageIndexClient(
index="gpt-5.6-luna", # model to build the tree index
chat="gpt-5.6-sol", # model to search the tree
)
doc_id = client.submit_document("report.pdf")["doc_id"]
answer = client.chat("What was the 2023 operating margin?", doc_id=doc_id)
print(answer)
Model Recommendations
index=: a basic model is sufficient. The tree structure itself is extracted from the document layout without an LLM; the index model only summarizes and refines it, which a basic model does well.chat=: use the best model you can afford. The chat model searches the tree to retrieve information. See Query cost and accuracy.
Use PageIndex through the SDK client →
Configure other models, streaming, multi-document search, citations, and more.
Integrate PageIndex with your own agent →
Drop PageIndex tools into the OpenAI Agents SDK, the Claude Agent SDK, or any other framework.
Benchmarks
Local indexing cost and time
Building a tree locally runs about $0.001 per page with gpt-5.6-luna as the index model, so a 1,000-page textbook costs a little over a dollar and a few minutes, once, and every later question reuses it. PageIndex is designed not to rely heavily on the model used at index time, so in our experiments a basic model does not hurt quality.
Indexing time also scales predictably with document length. In the same local setup, the benchmark documents (9 to 1,098 pages) finished in roughly 13 seconds to 4.5 minutes.
Query cost and accuracy
PageIndex-OSS-Benchmark measures exactly the setup in the quickstart above (PageIndexClient() in local mode, flash indexing, no OCR) on 62 lookup questions over 34 PDFs (1,945 pages) drawn from MMLongBench-Doc-V2. Every question's answer is a fact stated in running text, so a wrong answer is a retrieval or reading failure, not a reasoning one.
Full results, data, and the runner are in the benchmark repo.
Cost per query vs. native PDF input
The alternative to retrieval is handing the model the whole PDF on every question. That cost grows with the document; PageIndex's does not, because it reads only the nodes its reasoning reaches. On documents where both routes return the same answer, native PDF input costs 2.1× more at 52 pages and 16.6× more at 420 (gpt-5.6-sol, prompt caching excluded) — and at 805 pages the document no longer fits in the context window at all.
Leading accuracy on FinanceBench
PageIndex reached a state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector-based RAG.
Explore the full FinanceBench evaluation results and the blog post.
PageIndex Cloud
The open-source version is ideal for text-heavy PDFs and local workflows. With PageIndex Cloud, document indexing and storage run in the cloud: PageIndex handles parsing, OCR, image understanding, tree-index construction, and managed storage for you. The chat and retrieval layer remains compatible with your model, so you can search the cloud-hosted index using the model provider your application already uses.
Moving indexing and storage from Local to Cloud only requires a PageIndex API key:
import os
from pageindex import PageIndexClient
os.environ["PAGEINDEX_API_KEY"] = "your-pageindex-key"
os.environ["OPENAI_API_KEY"] = "your-openai-key"
client = PageIndexClient(
index="cloud", # build and store the index in PageIndex Cloud
chat="gpt-5.6-sol", # use your preferred compatible model for chat
)
doc_id = client.submit_document("report.pdf", wait=True)["doc_id"]
print(client.chat("What was the 2023 operating margin?", doc_id=doc_id))
| Capability | Local (this repo) | Cloud (get an API key) |
|---|---|---|
| Best for | text-heavy PDFs and local workflows | scanned, image-heavy, and large document collections |
| Indexing | runs locally | runs in PageIndex Cloud, with production OCR and image understanding |
| Storage | local | managed in PageIndex Cloud |
| Chat model | your model | your model, or the managed chat included with your key |
| Citations | page-level | line-level |
| Image understanding | — | ✅ |
| Multi-document scale | manual | PageIndex File System |
| MCP server | — | ✅ |
More About PageIndex Cloud
- Scale PageIndex to Millions of Documents: PageIndex File System is a Cloud-only, file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document.
Ready to Try It?
- Get a PageIndex API key
- Read the PageIndex Cloud documentation
For dedicated deployment (VPC or on-premises), contact us or book a demo.
⭐ Support Us
Leave us a star 🌟 if you like our project. Thank you!
Please cite this work as:
Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
Or use the BibTeX citation.
@article{zhang2025pageindex,
author = {Mingtian Zhang and Yu Tang and PageIndex Team},
title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
journal = {PageIndex Blog},
year = {2025},
month = {September},
note = {https://pageindex.ai/blog/pageindex-intro},
}
© 2026 PageIndex AI



