Ray d375c00a5a feat: local mode for the PageIndex SDK (v0.2.9) (#389)
* feat: add local mode to the PageIndex SDK client

One PageIndexClient, two backends. With api_key: the 0.2.x cloud SDK,
request for request (with the reviewed fixes: bounded timeouts on JSON
endpoints, none on uploads, URL-encoded ids, 401 key hint, empty DELETE
body tolerated). Without api_key: the same methods run locally —
page_index builds the tree in submit_document (mode="flash" uses
PageIndex Flash), documents are stored as plain JSON per doc under
storage_path, submit_query is LLM tree search with retrieve_model, and
chat_completions answers over the retrieved nodes with OpenAI-style
responses and streaming.

Local responses mirror the cloud wire shapes verified against the server
source: tree nodes rename start_index to page_index and drop end_index,
a non-leaf summary becomes prefix_summary, and the metadata/list/delete/
retrieval envelopes match key for key. Cloud-only features (folders,
beta_headers, enable_citations) raise instead of pretending.

Replaces the demo-only workspace client (index/get_document_structure/
get_page_content had no real users) and its retrieve.py helpers.
page_index_main gains an optional logger param so the SDK can keep
./logs out of the caller's working directory; pymupdf import is now lazy
(only the optional PyMuPDF parser path needs it).

* chore: package pageindex 0.3.0.dev4 for PyPI

Poetry packaging for the combined SDK + local pipeline: every production
import is a declared dependency (openai and requests join requirements.txt
for the same reason), config.yaml and the flash data tables ship in the
wheel, the benchmark PNG does not. pymupdf drops to an optional note now
that its import is lazy. dev4 follows the already-published 0.3.0.dev1-3;
pip still resolves plain 'pip install pageindex' to 0.2.8 until a final
0.3.0 — install with --pre.

* docs: add SDK section to README; move the agentic demo onto the SDK

The demo keeps its flow and the post-cutoff demo paper, swapping the
removed workspace client for PageIndexClient local mode (list_documents
for the doc-id cache, get_tree/get_ocr behind the agent tools). The old
examples/workspace JSONs demoed the removed format and go with it.

* refactor: rebuild cloud_api on the 0.2.8 client text

Ray's rule for the cloud half: his 0.2.8 code is the base; a Kylin-lineage
change survives only when strictly better — invisible on healthy traffic
while fixing a real failure mode. Kept under that bar: request timeouts
(dead connections hung forever; uploads still pass none, exactly like
0.2.8), the upload handle closed via with (leaked on request errors),
URL-encoded path ids (a crafted id could reroute the URL), the
empty-DELETE-body guard, and stream hardening (choices guard,
response.close in finally). Reverted as not strictly better: the
lowercase summary param (the server accepts both spellings), the 401
message hint (visible text change; AUTH_HINT dropped from errors.py), and
the _request/requests.request reorganization — every method body,
docstring, and section comment is 0.2.8's text again.

diff -w against ../pageindex_sdk/pageindex/client.py now reads as that
surgical patch plus plumbing: CloudAPI reads BASE_URL/api_key through the
owning client, PageIndexAPIError comes from errors.py, and
is_retrieval_ready lives verbatim on PageIndexClient shared by both
modes. A mocked-requests harness driving 0.2.8 and this file through 19
identical calls shows the only remaining request-level difference is
timeout.

* refactor: drop the local retrieval endpoints — cloud-only, deprecated

Ray: the cloud already marks POST /retrieval/ and GET /retrieval/{id}/
deprecated in favor of chat completions, so local mode should not grow a
fresh implementation of a retiring surface. submit_query/get_retrieval
now raise in local mode with a pointer to chat_completions; cloud mode is
untouched (the endpoint still works there and 0.2.8 code keeps running).
The tree search that backed them stays as chat_completions' retrieval
engine; the retrievals/ storage goes away. This also closes the one real
cross-mode parity gap — the retrieved_nodes inner shape — by removing
its local half.

* feat: manifest.json — one-file document listings for the local store

Ray wanted a metafile that shows every document in one place instead of
per-directory reads. It is a cache, never a second source of truth:
writers update it best-effort after save/delete (atomic replace, no
locks), and list_metas trusts it only while its id set matches the docs/
directory names — documents are immutable, so matching names imply valid
content. Any mismatch (lost concurrent update, crash, corrupt or deleted
manifest) rebuilds it from the doc.json files, reading only the missing
entries. Incomplete dirs (no doc.json) stay invisible and are never
recorded, so a save that completes later is still picked up.

1000-doc listing: 37ms of per-dir reads -> 2.3ms warm (scandir names +
one manifest read); one-time rebuild 160ms.

* fix: align local doc_id prefix and createdAt format with the cloud

Local doc ids now carry the cloud's pi- namespace prefix (random token
stays uuid4 hex — nothing parses cuid internals, the prefix is the
contract; chat ids already mirrored chatcmpl-). createdAt now matches
the server byte for byte: the cloud emits the DB datetime's bare
isoformat — naive UTC, second precision — while we emitted microseconds
plus +00:00.

* feat: optional metadata tags on submit_document (both modes)

Ray's call after the alignment review: the metadata key in tree/OCR
envelopes and list entries should carry real data, and the only honest
way is exposing the field the server already accepts. Cloud mode
forwards it as the existing metadata form field; local mode validates it
early (a JSON-serializable dict, checked before any LLM spend), stores
it in doc.json, and returns it from the same three places. get_document
still omits it, mirroring the server, whose metadata-endpoint SQL never
selects that column. Scope stays deliberately narrow: set at submit and
read back — no metadata_filter, no update API.

* docs: state that createdAt is UTC and show how to localize it

The value is naive UTC in both modes (the cloud column is timestamp
DEFAULT CURRENT_TIMESTAMP on a UTC server, emitted via bare isoformat).
Wall-clock display is the consumer's layer: emitting local time under
the same format would silently change meaning per machine, and adding an
offset marker would break both the byte-format parity and string-order
sorting.

* fix: createdAt carries milliseconds, matching the cloud's datetime(3)

The earlier second-precision alignment was reasoned from the postgres
schema file, but production is MySQL (DATABASE_BACKEND defaults to
mysql) and its FilePageIndex.createdAt is datetime(3) DEFAULT
CURRENT_TIMESTAMP(3) — so the server isoformat()s a millisecond-
precision naive-UTC datetime, emitting .XXX000 fractions (bare seconds
only when the millisecond happens to be zero). Local now generates
through the same mechanism. The docs' 2024-01-15T10:30:00.000Z sample is
a JS-style string Python isoformat cannot produce — not evidence.

* fix: createdAt at millisecond precision, matching the timestamp(3) column

The cloud column is timestamp(3); its datetimes render through bare
isoformat() as six fractional digits ending in 000 (or no fraction when
the millisecond is exactly zero). Truncate to the millisecond and render
the same way, replacing the second-precision guess from cd44ef0. Worth
one confirmation against a live cloud response when a key is around.

* chore: trim non-essential comments

Alignment narration (mirrors the cloud, matches 0.2.8, trailing-period
notes) moves out of the code — that rationale lives in the commit
history. Kept only constraints the code can't show: the storage commit
protocol and manifest trust rule, the id-escape guard, the silent
logger's reason to exist, the client-reference indirection, and the
tree-node reshape spec.

* fix: close the local_store crash and corruption holes found in review

Adversarial review reproduced three edge holes in the no-lock design:
a torn delete (doc.json unlinked, rmtree unfinished) left a ghost the
manifest kept listing forever; one truncated doc.json crashed every
listing with a raw JSONDecodeError; and two concurrent deletes of the
same id could crash the loser's rmtree.

doc.json is now the existence marker in both directions — written last
on save, unlinked first on delete — and list_metas serves an entry only
after confirming it still exists, instead of trusting the manifest/dir
name-set proxy. Atomic writes fsync before replace, closing the
power-loss truncation window. An unreadable JSON file is logged and
treated as absent; a document meta additionally falls back to the
manifest copy (immutable docs make the cache a valid replica). rmtree
runs with ignore_errors: the commit point has passed and cleanup is
best-effort, which also absorbs the double-delete race.

Also restores the reversed-page-range error in the demo's page parser
(silently empty since the retrieve.py removal). Warm 1000-doc listing
goes 2.3ms -> 8.3ms for the per-doc existence check; full parse remains
37ms.

* docs: correct two docstring claims and the pymupdf note

is_retrieval_ready reports only API errors as False — transport errors
propagate; delete_document may return {} on an empty cloud body; pymupdf
is also used by tree_optimize's page loading, not just get_page_tokens.

* chore: target 0.2.9 for the local-mode release

Ray's call: 0.2.x stays the no-collections line, so local mode ships as
0.2.9 and 0.3.0 stays reserved; plain pip installs never see the
0.3.0.devN pre-releases, and the cookbooks' existing 'pip install
--upgrade pageindex' will deliver the new SDK without any --pre
instructions.

* ci: publish to PyPI on version tags

Tag-driven releases: the pushed v-tag is the single source of truth for
the version — validated as PEP 440, injected into pyproject, built, and
published via OIDC trusted publishing with no stored credentials; a
GitHub Release with the artifacts is created alongside.

* chore: tighten the store docstring to essentials

* fix: contain invalid-UTF-8 corruption; fail loud on unreadable data files

The corruption guard caught JSONDecodeError but not the
UnicodeDecodeError a torn multi-byte write produces — the exact scenario
the guard targets whenever names or descriptions carry non-ASCII text —
so that flavor crashed listings and gets raw, and regressed the old
manifest read's broader ValueError guard. _read_json now catches
ValueError, which covers both.

Unreadable tree.json/pages.json under an intact doc.json previously
served an empty tree with retrieval_ready true — a silent lie; those
paths now raise 'stored document data is unreadable' and
is_retrieval_ready honestly reports False. delete_document survives a
doc.json tampered into a directory (cleans it, reports not-found)
while real unlink failures such as permissions stay loud.

* feat: LocalClient and CloudClient for explicit mode selection

PageIndexClient(api_key=os.getenv(...)) with an unset variable gets None
and silently falls into local mode — methods keep working against local
storage on the caller's own LLM bill, the silent mode flip the
empty-string guard can't see. The explicit classes close it at the type
level: CloudClient raises on a missing key, LocalClient has no api_key
parameter at all. Names per Ray.

* feat: explicit-mode clients PageIndexCloudClient and PageIndexLocalClient

PageIndexClient(api_key=os.getenv(...)) with an unset variable yields
None and silently lands in local mode — with both modes fully working,
that's a silent mode flip onto the user's own LLM bill. The explicit
classes pin the mode at construction: the cloud one refuses a missing or
empty key, the local one has no key parameter at all. Names follow the
package's PageIndex- prefix convention.

* fix local mode edge cases

* fix: keep the litellm/ prefix normalization the demo depends on

a108c02 added _normalize_retrieve_model because the agentic demo hands
client.retrieve_model straight to the OpenAI Agents SDK, which routes a
non-OpenAI provider only when the name carries a litellm/ prefix. The
normalization sat in client.py, so the demo line never had to change --
and this rewrite dropped the helper while leaving that lone consumer
untouched. A retrieve_model like anthropic/claude-sonnet-4-6, the form
config.yaml documents, then died at Agent() with "Unknown prefix".

Local mode is indifferent to which form it gets: llm_completion and
_chat_llm both removeprefix("litellm/"), count_tokens returns the same
count either way, and _is_openai_model classifies both as LiteLLM, so
_require_llm_key still asks for no OPENAI_API_KEY. Models without a
provider path (the packaged gpt-5.4 default) pass through untouched.

* fix: restore the published 0.2.8 helper signatures the cookbooks call

pyproject.toml makes this repo the source of the PyPI pageindex package,
so this utils.py replaces the published 0.2.8 one -- whose helpers are
the documented surface of the cookbook notebooks. Three had drifted:
remove_fields lost max_len, create_node_mapping lost
include_page_ranges/max_page, print_tree lost exclude_fields. Both
README-linked notebooks open with `pip install --upgrade pageindex` and
pass exactly those kwargs, so tagging v0.2.9 as-is would TypeError
every Colab run of vision_RAG_pageindex.ipynb.

remove_fields and create_node_mapping readopt the 0.2.8 bodies, strict
supersets of the current ones (no in-repo caller passes the new params).
print_tree keeps the outline view as its default and routes an explicit
exclude_fields= to the 0.2.8 pprint view -- the two versions disagree
on what the second positional means (indent vs exclude_fields), and the
notebooks pass it by keyword. call_llm, the fifth published name, stays
out: nothing imports it -- the one notebook using a call_llm defines
its own, with a different signature.

* perf: resolve the indexing stack lazily from pageindex/__init__

`import pageindex` eagerly pulled page_index, flash, and tree_optimize
-- 0.73s warm, numpy and pypdfium2 in-process -- while the published
0.2.8 package imported in 0.08s on requests+openai alone. A cloud-only
SDK user upgrading to 0.2.9 would pay for an indexing stack they never
call, on every interpreter start.

The eager surface shrinks to client and errors (2ms warm); everything
else resolves on first attribute access via PEP 562 and is cached in
the module namespace. `pageindex.page_index` stays the function, still
shadowing its submodule as the old star-import had it. __all__ now
names the public surface, so `from pageindex import *` binds the same
working set as before instead of 124 names including stdlib modules.
A TYPE_CHECKING block keeps real signatures visible to IDEs.

The test suite's sys.modules lookup assumed the eager import chain;
it now imports pageindex.page_index explicitly.

* fix: wrap PDF read failures in PageIndexAPIError on local submit

_extract_page_texts sat one line outside the try that wraps everything
else in submit_document, so a corrupt PDF surfaced as a raw
PyPDF2.errors.PdfReadError and a password-protected one as
FileNotDecryptedError -- while a blank PDF, checked on the very next
line, got a clean PageIndexAPIError. Callers handling the SDK error
type crashed on exactly the malformed downloads and encrypted files
an ingest loop sees most.

The extraction gets its own wrap rather than joining the indexer try
below, whose except would re-prefix the blank-PDF error into "Failed
to submit document: Failed to submit document: ...". FileNotFoundError
stays native, asserted by test_submit_rejections as cloud parity.

* fix: address code review findings across local mode and publish workflow

- _parse_json_reply: switch from extract_json to _reply_json (dead
  try/except, global None→null substitution, wrong error messages)
- client.py: config override filter uses `is not None` instead of
  truthiness, so empty-string model args are no longer silently dropped
- llm_completion/llm_acompletion: raise RuntimeError after retries
  exhausted instead of returning empty string
- _require_llm_key: extend to anthropic/gemini/mistral providers
- _chat_llm: reuse _openai_sync_client singleton from utils
- _tree_search: drop redundant deepcopy before non-mutating remove_fields
- _stream_chunks: move final chunk inside try so usage data is reachable;
  close stays in finally
- _validate_chat_messages: accept system messages, merge them into the
  internal system prompt for cloud/local parity
- _build_chat_context: accept pre-read metas and pass structure through
  to _tree_search, eliminating double get_meta and double tree.json reads
- _index_standard: reuse ConfigLoader from construction; reject empty
  structure (parity with flash mode)
- extract_json: bare except: → except Exception: (no longer swallows
  KeyboardInterrupt)
- Replace all str.removeprefix() with _strip_prefix() helper to restore
  Python 3.7 compatibility; pyproject.toml back to python >= 3.7
- publish.yml: add test job (py3.10 + py3.13) gating the publish job
- Remove unused bare `import pageindex` from tests

* fix: tighten CI permissions, timestamp format, import style, and docstring

- publish.yml: add permissions: {contents: read} to the test job
- local_api: emit 3-digit ms timestamps via isoformat(timespec="milliseconds")
- local_api: use relative import for pageindex.utils in _chat_llm
- local_api: note the node_summary gate in _format_tree_node docstring
- test_client: update createdAt regex to match the new 3-digit format

* fix: accept GOOGLE_API_KEY for Gemini; mark pre-releases in GitHub

- _require_llm_key: Gemini provider now accepts either GEMINI_API_KEY
  or GOOGLE_API_KEY (LiteLLM supports both)
- publish.yml: set prerelease flag on rc/dev/alpha/beta tags so they
  are not shown as regular releases on GitHub

* fix: preserve print_tree backward compat with 0.2.8 positional call

Detect list passed as second positional arg (old 0.2.8 signature) and
treat it as exclude_fields instead of indent.

* refactor: print_tree param order — exclude_fields second for 0.2.8 compat

Move exclude_fields back to the second position (matching 0.2.8) instead
of detecting list-as-indent. Recursive call uses indent= keyword arg.

* refactor: remove _require_llm_key pre-check entirely

Let OpenAI SDK and LiteLLM report their own missing-key errors instead
of maintaining a parallel provider-to-env-var map. The except Exception
wrapper in chat_completions already converts these to PageIndexAPIError.

* fix: let LLM provider errors propagate instead of wrapping them

OpenAI/litellm auth, rate-limit, and other provider errors now reach the
caller as their original type (e.g. openai.AuthenticationError) instead
of being wrapped in PageIndexAPIError. PageIndexAPIError stays reserved
for PageIndex's own errors (bad doc_id, invalid params, indexing failures).

* refactor: catch only RuntimeError instead of isinstance check on openai

Only our own _tree_search logic raises RuntimeError (bad JSON, missing
node_list). Provider errors and unexpected bugs propagate naturally.

* fix: restore createdAt to 6-digit .177000 format matching cloud DATETIME(3)

Revert the timespec="milliseconds" that a linter introduced — it output
.177 (3 digits) while the cloud server's isoformat() outputs .177000
(6 digits). Verified against pageindex-compute server/lib/db/planet.py:
DATETIME(3) column + bare isoformat() = .177000.

* fix: resolve 15 review findings from PR #389

- Rename page_index.py → page_index_classic.py to fix __getattr__
  shadowing (function permanently replaced by submodule after import)
- Add return_exceptions=True to verify_toc, generate_summaries, and
  summarize_tree gathers so one LLM failure doesn't abort the batch
- Gracefully degrade generate_doc_description to "" on failure instead
  of discarding the entire completed index
- Normalize file_path with str()/expanduser/abspath to support
  pathlib.Path and tilde paths
- Strip text from tree.json at save time; reconstruct from pages.json
  on read via _load_tree_with_text
- Filter empty-string model overrides in PageIndexClient constructor
- Guard empty choices list before indexing response.choices[0]
- Wrap streaming iteration errors as PageIndexAPIError inside the
  generator
- Catch PermissionError in _read_json alongside FileNotFoundError
- Set max_retries=3 for OpenAI and num_retries=3 for litellm in
  _chat_llm to match the retry behavior of the indexing path
- Replace _SilentLogger with logging.getLogger(__name__)
- Add .github/workflows/tests.yml for PR and push-to-main test runs

* fix: close publish workflow injection and detect flash silent summary failure

1. publish.yml: pass VERSION via os.environ instead of shell interpolation
   into python -c, eliminating the command injection vector.

2. summarize_tree: raise RuntimeError when every node's summary generation
   fails (e.g. missing LLM credentials), instead of silently saving a
   document with all-empty summaries marked as completed.

* refactor: make chat_completions cloud-only until agent-based local chat lands

The local implementation was a tree-search RAG engine (retrieve prompt +
context stuffing) that diverged from the cloud /chat/completions design,
where an agent navigates documents through MCP tools. Rather than ship
the divergent engine in 0.2.9, remove it; local chat returns as an agent
loop built on the agent-tools layer in the 0.2.10 line.

Also from the PR #389 review:
- get_tree fails loud on unreadable pages.json instead of silently
  serving textless nodes
- drop the wasted deepcopy in get_tree (_format_tree_node builds new dicts)
- generate_doc_description catches only RuntimeError so provider errors
  (bad key, unknown model) propagate instead of storing ""
- _read_json treats IsADirectoryError as unreadable
- list_metas skips directory names that fail _is_safe_id
- constructing with api_key plus local-only args raises PageIndexAPIError
  (was ValueError) to match the empty-api_key path

* fix: verification follow-ups for the chat removal commit

- retrieve_model docstring no longer promises the deleted tree-search
  machinery; the param is reserved for the coming agent-based local chat
- empty pages.json ([]) fails loud like unreadable pages: a stored
  document can never legitimately have zero pages, and get_tree's
  "nodes always carry text" promise held only for the None case
- README: chat example moved to its own cloud-only block so the local
  quickstart no longer ends in a raise
- tests: pin IsADirectoryError loud-fail, list_metas unsafe-name skip,
  empty-pages loud-fail, and generate_doc_description's swallow-vs-
  propagate boundary

* fix: detect classic-path silent summary failure like flash already does

generate_summaries_for_structure absorbed every per-node error to
summary: "" with no systemic check, so a bad key or model name produced
an all-empty-summary tree with zero errors on the default CLI config
(if_add_node_summary: yes, if_add_doc_description: no). Mirror flash's
summarize_tree guard: partial failures still absorb, all-failed raises.

* fix: accurate wording for retrieve_model docstring and empty-pages error

- retrieve_model is not "unused": the agent demo drives its model from
  client.retrieve_model today; say so instead
- an empty pages.json is invalid, not unreadable — dedicated
  _require_pages raises "stored document has no page content" so the
  operator reindexes instead of hunting for file corruption

* chore: trim non-essential comments, fix three docstring issues

Review fixes:
- client.py: remove unverified "pending" from get_document status enum
- cloud_api.py: update stale "kept line-for-line" module docstring
- cloud_api.py: add node_summary to get_tree Args

Comment trimming across __init__.py, client.py, cloud_api.py, errors.py,
local_api.py, local_store.py — module docstrings shortened, explanatory
inline comments removed, multi-line class docstrings collapsed to one line.

* feat: add get_page_content and get_tree include_text parameter

- client.get_page_content(doc_id, pages): convenience method wrapping
  get_ocr + page filtering; _parse_pages moved from the demo into client.py
- client.get_tree(..., include_text=False): skips node text for
  structure-only views; local reads the stored tree directly (no text
  to begin with), cloud strips client-side via remove_fields
- demo simplified: tools now call the new client methods directly

* simplify: drop redundant try/except in demo get_page_content tool

* feat: add get_document_structure convenience method

* test: cover get_page_content, get_document_structure, include_text=False

* fix: guard against four edge-case crashes found in PR #389 review

- Demo: use getattr for retrieve_model so cloud clients don't AttributeError
- utils: return empty string when start/end page index is None instead of TypeError
- local_store: clean up temp file on _write_json_atomic failure
- local_api: re-raise PageIndexAPIError before the catch-all Exception block

* chore: trim verbose optional-dep comments in requirements.txt

* revert: restore README.md to main — SDK section deferred to next version

* fix: eliminate double PDF parse in local standard indexing

_extract_page_texts already reads the PDF via PyPDF2; pass the
pre-extracted texts as page_list to page_index_main so it skips
its own get_page_tokens call. Token counts are computed once via
litellm.token_counter in _index_standard.

* test: assert page_list is passed and correctly shaped

* fix: wire include_text to cloud API and guard get_page_content on processing docs

cloud_api.py now sends include_text as a query param so the server can
omit node text from tree responses, saving bandwidth.  Backward-compatible:
old servers ignore the param, client-side remove_fields still strips.

get_page_content raises PageIndexAPIError instead of TypeError when the
document is still processing (get_ocr returns result: null).
2026-08-11 21:52:03 +08:00
2026-03-27 03:30:13 +08:00
…
2026-08-04 18:38:29 +08:00

PageIndex Banner

VectifyAI%2FPageIndex | Trendshift

PageIndex: Vectorless, Reasoning-based RAG

Reasoning-based RAG  ◦  No Vector DB, No Chunking  ◦  Context-Aware Retrieval  ◦  Reads Like a Human

🌐 Website  •   🖥️ Chat Platform  •   🔌 MCP & API  •   📖 Docs  •   💬 Discord  •   ✉️ Contact 

📢 Updates

  • 🔥 Agentic Vectorless RAG — A simple agentic, vectorless RAG example with self-hosted PageIndex, using OpenAI Agents SDK.
  • Scale PageIndex to Millions of Documents — PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document, enabling massive-scale document search.
  • PageIndex Chat — Human-like document analysis agent platform for professional long documents. Also available via MCP or API.
  • PageIndex Framework — Deep dive into PageIndex: an agentic, in-context tree index that enables LLMs to perform reasoning-based, context-aware retrieval over long documents.

📑 Introduction to PageIndex

Are you frustrated with vector database retrieval accuracy for long professional documents? Traditional vector-based RAG relies on semantic similarity rather than true relevance. But similarity ≠ relevance — what we truly need in retrieval is relevance, and that requires reasoning. When working with professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search often falls short — missing what's relevant but not similar, and returning what's similar yet not relevant.

Inspired by AlphaGo, we propose PageIndex — a vectorless, reasoning-based RAG system that builds a hierarchical tree index from long documents, and uses LLMs to reason over that index for agentic, context-aware retrieval. The retrieval is traceable and explainable, with no vector DBs or chunking. PageIndex simulates how human experts navigate and extract knowledge from complex documents through tree search, enabling LLMs to think and reason their way to the most relevant document sections. It performs retrieval in two steps:

  1. Generate a “Table-of-Contents” tree structure index of documents
  2. Perform (agentic) reasoning-based retrieval through tree search

🎯 Core Features

PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, without vector databases or chunking.

Compared to traditional vector-based RAG, PageIndex features:

  • No Vector DB: Uses document structure and LLM reasoning for retrieval, instead of vector similarity search.
  • No Chunking: Documents are organized into natural sections, not artificial chunks.
  • Better Traceability & Explainability: Retrieval is reasoning-driven and grounded in explicit page and section references, making every result traceable and interpretable — no more “vibe retrieval” with opaque, approximate vector search.
  • Context-Aware Retrieval: Retrieval depends on your full context (e.g., conversation history and domain knowledge), and easily incorporates new context.
  • Human-like Retrieval: Mirrors how human experts navigate and extract knowledge from complex documents.

PageIndex achieved state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector RAG solutions on professional document analysis (blog post).

📍 Explore PageIndex

To learn more, please see a detailed introduction to the PageIndex framework. Check out our GitHub for open-source code, and the cookbooks, tutorials, and blog for more usage guides and examples.

The PageIndex service is available as a ChatGPT-style chat platform, or can be integrated via MCP or API, with enterprise deployment available.

🛠️ Deployment Options

  • Self-host — run locally with this open-source repo (using standard PDF parsing).
  • Cloud Service — production-grade pipeline with enhanced OCR, tree building, and retrieval for best results. Try instantly on our Chat Platform, or integrate via MCP or API.
  • Enterprise — dedicated or private deployment (VPC, on-prem). Contact us or book a demo to learn more.

🧪 Quick Hands-on

  • ⚡ PageIndex Flash (preview) — ultra fast PageIndex tree structure generation from PDFs.
  • 🔥 Agentic Vectorless RAG (latest) — a simple but complete agentic vectorless RAG example with self-hosted PageIndex, using OpenAI Agents SDK.
  • Try the Vectorless RAG notebook — a minimal, hands-on example of reasoning-based RAG using PageIndex.
  • Check out Vision-based Vectorless RAG — no OCR; a minimal, vision-based & reasoning-native RAG pipeline that works directly over page images.
View on GitHub: Agentic Vectorless RAG
Open in Colab: Vectorless RAG    Open in Colab: Vision RAG

🌲 PageIndex Tree Structure

PageIndex can transform lengthy PDF documents into a semantic tree structure, similar to a “table of contents” but optimized for use with LLMs and AI agents. It's ideal for: financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any long, complex professional documents.

Below is an example PageIndex tree structure. Also see more example documents and generated tree structures.

...
{
  "title": "Financial Stability",
  "node_id": "0006",
  "start_index": 21,
  "end_index": 22,
  "summary": "The Federal Reserve ...",
  "nodes": [
    {
      "title": "Monitoring Financial Vulnerabilities",
      "node_id": "0007",
      "start_index": 22,
      "end_index": 28,
      "summary": "The Federal Reserve's monitoring ..."
    },
    {
      "title": "Domestic and International Cooperation and Coordination",
      "node_id": "0008",
      "start_index": 28,
      "end_index": 31,
      "summary": "In 2023, the Federal Reserve collaborated ..."
    }
  ]
}
...

You can generate PageIndex tree structures with this open-source repo. Or use our API for higher-quality results powered by our enhanced OCR and tree building pipeline.


⚙️ Package Usage

Note: This package uses standard PDF parsing. For use cases with complex PDFs, our cloud service (via MCP and API) offers enhanced OCR, tree building, and retrieval.

You can follow these steps to generate a PageIndex tree from a PDF document.

1. Install dependencies

pip3 install --upgrade -r requirements.txt

2. Set your LLM API key

Create a .env file in the root directory with your LLM API key. Multi-LLM is supported via LiteLLM:

OPENAI_API_KEY=your_openai_key_here

3. Generate PageIndex structure for your PDF

python3 run_pageindex.py --pdf_path /path/to/your/document.pdf
Optional parameters
You can customize the processing with additional optional arguments:
--model                 LLM model to use (default: gpt-4o-2024-11-20)
--toc-check-pages       Pages to check for table of contents (default: 20)
--max-pages-per-node    Max pages per node (default: 10)
--max-tokens-per-node   Max tokens per node (default: 20000)
--if-add-node-id        Add node ID (yes/no, default: yes)
--if-add-node-summary   Add node summary (yes/no, default: yes)
--if-add-doc-description Add doc description (yes/no, default: yes)
Markdown support
We also provide markdown support for PageIndex. You can use the `--md_path` flag to generate a tree structure for a markdown file.
python3 run_pageindex.py --md_path /path/to/your/document.md

Note: in this mode, we use "#" to determine node headings and their levels. For example, "##" is level 2, "###" is level 3, etc. Make sure your markdown file is formatted correctly. If your Markdown file was converted from a PDF or HTML, we don't recommend using this mode, since most existing conversion tools cannot preserve the original hierarchy. Instead, use our PageIndex OCR, which is designed to preserve it, to convert the PDF to a markdown file and then use this mode.

⚡ PageIndex Flash (preview)

PageIndex Flash (pageindex/flash) generates tree structures from PDFs in seconds. Structure extraction is purely heuristic-based, no LLM needed. LLM is only used to generate node summaries.

python3 run_pageindex.py --flash --pdf_path /path/to/your/document.pdf

Add --optimize to refine the tree structure for more efficient retrieval (with an LLM expansion pass).

🚀 Agentic Vectorless RAG: An Example

For a simple, end-to-end agentic vectorless RAG example using self-hosted PageIndex (with OpenAI Agents SDK), see examples/agentic_vectorless_rag_demo.py.

# Install optional dependency
pip3 install openai-agents

# Run the demo
python3 examples/agentic_vectorless_rag_demo.py

📈 Case Study: PageIndex Leads Finance QA Benchmark

Mafin 2.5 is a reasoning-based RAG system for financial document analysis, powered by PageIndex. It achieved a state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), significantly outperforming traditional vector-based RAG systems.

PageIndex's hierarchical indexing and reasoning-driven retrieval enable precise navigation and extraction of relevant context from complex financial reports, such as SEC filings and earnings disclosures.

Explore the full benchmark results and our blog post for detailed comparisons and performance metrics.


🧭 Resources

  • 📝 Blog: technical articles, research insights, and product updates.
  • 🔧 Developer: MCP setup, API docs, and integration guides.
  • 🧪 Cookbooks: hands-on, runnable examples and advanced use cases.
  • 📖 Tutorials: practical guides and strategies, including Document Search and Tree Search.

⭐ Support Us

Leave us a star 🌟 if you like our project. Thank you!

Please cite this work as:

Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
Or use the BibTeX citation.
@article{zhang2025pageindex,
  author = {Mingtian Zhang and Yu Tang and PageIndex Team},
  title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
  journal = {PageIndex Blog},
  year = {2025},
  month = {September},
  note = {https://pageindex.ai/blog/pageindex-intro},
}

🌐 Open-Source Ecosystem

PageIndex anchors a growing open-source ecosystem of long-context AI infra — OpenKB is an LLM knowledge base that compiles documents into an interlinked wiki. ChatIndex provides tree indexing and retrieval for long conversational histories and memory. ConDB is a KV-cache native context database for tree-based retrieval at scale. PageIndex MCP is PageIndex's MCP server.

Connect with Us

Website  Twitter  LinkedIn  Discord  Book a Demo  Contact Us


© 2026 Vectify AI

Languages
Python 99.5%
JavaScript 0.4%
Shell 0.1%