Ray bc1c1740ff feat: model knobs, LiteLLM-verbatim chat lane, per-door passthrough params (#409)
* docs: drop the demo's install step — openai-agents ships with the SDK now

* feat: the chat lane routes every model through LiteLLM — bare names included

The direct-OpenAI special case existed to dodge LiteLLM's import cost,
and it made OpenAI's own Responses-first models fail on the front door:
gpt-5.6-sol 400s on chatcmpl+tools while reasoning is on (server-side
policy — wire-captured with no reasoning_effort in our request).
LiteLLM 1.97 translates such calls onto /v1/responses; 1.84 does not,
so the sol-class 400 now carries its two exits (upgrade litellm /
responses()).

Routing after the flip: chat protocol — bare names are OpenAI-compatible
shorthand (wire form openai/<name>; OPENAI_API_KEY / OPENAI_BASE_URL
still select the backend, and the missing key stays a build-time
failure), litellm/ strips, openai/ opts out to the OpenAI SDK directly;
responses protocol unchanged (OpenAI-SDK native, LiteLLM refused).

The import cost is handled instead of dodged: local clients preload
litellm on a background thread (first call then perceives 0.0s), and
pageindex sets LITELLM_LOCAL_MODEL_COST_MAP=True via setdefault —
LiteLLM's import otherwise blocks on a network fetch of its price map
(fresh venv: 5.6s -> 1.3s; offline it hangs to the timeout).

Also restores prompt_cache_key delivery, found dead during the flip's
gating verification: openai-agents 0.20 no longer derives it from
RunConfig.group_id, so both lanes sent nothing. ModelSettings.extra_body
is the one channel all three model classes put on the wire (the bare
kwarg is dropped by LiteLLM; extra_args[extra_body] collides with the
responses model's own parameter — both wire-verified), and it is scoped
to OpenAI destinations: LiteLLM plants extra_body as a literal field in
other providers' bodies, and Anthropic rejects unknown fields — the
anthropic wire test now pins the absence.

Verified before landing: mock-server matrix (OPENAI_BASE_URL + bare
name works through LiteLLM; gpt-named self-hosted models are NOT
bridged off a custom base_url; prompt_cache_key on the wire in every
OpenAI lane with distinct per-conversation keys; anthropic body clean)
and live (sol answers through chat(), gpt-5.4 unchanged, responses()
bare unchanged with the key on its wire).

* refactor: no prefix-triggered direct lane — chat model names are LiteLLM's, verbatim

Ray's ruling on the flip's remaining carve-out: a routing decision must
never hide in a model-name prefix. openai/ now means what LiteLLM says
it means (its openai provider), like every other name on the chat lane —
the grammar is LiteLLM's with zero exceptions.

The two defenses for keeping a direct carve-out had no concrete victim:
debugging isolation (litellm is unavoidable in indexing anyway, and
responses() IS the OpenAI-SDK-native door), and endpoint determinism
(litellm sends chatcmpl for openai-provider models except the gpt-5
bridge, which never fires against a custom base_url — wire-verified).
If a direct escape is ever needed, it will be a declared parameter,
never name grammar.

openai/-prefixed names keep the build-time OPENAI_API_KEY check for
parity with bare names; responses() is untouched (bare and openai/
still drive the OpenAI SDK — LiteLLM cannot speak that protocol).

* fix: third-audit findings — extra_body naming drift, litellm floor hint

The prompt_cache_key delivery channel went through three iterations and
settled on ModelSettings.extra_body; the _conversation_cache_key
docstring still named extra_args from the middle iteration.

The litellm install hint said >=1.30, below both our own pyproject floor
(>=1.84.0) and the floor openai-agents' litellm extra declares (>=1.83).
A user in a broken environment following it would land on a version the
package itself rules out. The hint now matches the declared floor.

* test: pin the bundle door's Agents-SDK model grammar; rename the cache-key test to extra_body

The config bundle hands its model string to the Agents SDK's own
MultiProvider grammar, which refuses unknown prefixes (probe on 0.20:
'anthropic/x' -> UserError: Unknown prefix). _normalize_retrieve_model's
litellm/ spelling is what keeps that door working — a link the existing
self-referential assert (config["model"] == client.retrieve_model)
could not catch. Pinned with a provider-slashed name.

Also renames the cache-key delivery test to its real channel,
extra_body — the extra_args name survived from the superseded delivery
attempt.

* refactor: name the retrieve_model helper for its reason — the Agents SDK's grammar

_normalize_retrieve_model said what it does, not why. The litellm/
spelling exists because the Agents SDK resolves raw model strings with
its own prefix grammar and refuses unknown prefixes — the name now
points at that constraint.

* test: the bundle-grammar test skips without openai-agents, like its file's siblings

Every agents-dependent test in this file importorskips; without the
guard this one errors where the others skip.

* feat: index_model + chat_model — two-knob model surface with full legacy fallback

The documented surface becomes two role knobs: index_model builds the
index, chat_model answers on the chat surfaces. model turns into the
set-both umbrella (its 0.2.8 indexing semantics are a strict subset, so
old configs run unchanged); summary_model and retrieve_model stay
accepted as legacy role names.

Resolution lives in ConfigLoader.load(), the one seam every consumer
already passes through (client, CLI standard/md paths, flash's
summary fallback, tree_optimize's default_model): new names win over
old, specific over general, model sets every role, and code constants
close each chain. The packaged yaml no longer ships model keys — key
presence is what separates a user's explicit choice from a built-in
default, and _validate_keys accepts the five model names explicitly.

Consequences: with no config at all, classic-mode structure extraction
now uses DEFAULT_INDEX_MODEL (gpt-5.6-luna) instead of the yaml's old
gpt-4o-2024-11-20 line (ratified; flash-default users see no change).
client.retrieve_model becomes a read-only alias for client.chat_model.

The resolution matrix test pins one row per released generation:
0.2.8 (model), 0.3.0.dev (model+retrieve_model), 0.2.10.dev (all three
legacy names), the new pair, umbrella-only, and mixed.

* feat: --index-model on the CLI; README flag docs follow

The CLI leads with --index-model; --model stays as its legacy synonym
(the CLI only indexes, so the umbrella and the index role coincide).
The flash branch's summary fallback gains the index position, and the
standard branch now forwards --summary-model, which it had silently
ignored — the flag's help always claimed it worked there. The md
branch's unfiltered model=None no longer clobbers the default: the
resolver treats None as unset.

* feat: default chat model becomes gpt-5.6-sol

Ray's pick for the out-of-box QA default; indexing stays on luna. sol
runs tools on the Responses lane — current litellm bridges chat()
there automatically; older litellm gets the guided 400 naming both
exits.

* chore: trim the model-keys comment to the constraint

The per-generation history lives in 7244ee4's message.

* feat: per-door reasoning passthrough — reasoning_effort / reasoning / thinking

Each chat door gains its own protocol's native thinking control,
forwarded verbatim with no invented vocabulary and no default of ours:
chat_completions(reasoning_effort=...), responses(reasoning={...}),
messages(thinking={...}). Unset sends nothing, so backend defaults
(sol: medium, adaptive) are untouched. chat() stays answer-only.

Delivery channels, each verified: the chat door rides
extra_args["reasoning_effort"] — LiteLLM's own top-level kwarg on
every supported openai-agents version, admitting non-enum values
("none"); newer openai-agents promotes it to the top-level argument
and pops the duplicate. Wire-captured on a mock backend
(/v1/chat/completions body carries it) and coexists with the Claude
cache marker in one dict. The responses door rides
ModelSettings.reasoning — coerced to the typed openai Reasoning object
and forwarded verbatim by the Responses model; the envelope echoes the
caller's dict. The messages door joins the existing anthropic
passthrough dict, asserted through the real tool runner.

LiteLLM semantics observed and accepted as-is: unknown models refuse
the param loudly with LiteLLM's own remedies, and gpt-5.4+ names with
an explicit effort route to /v1/responses even against a custom
api_base (its documented pre-existing arm). The sol-class 400 guidance
now names the third exit — an explicit effort routes on older litellm
releases too.

Cloud chat_completions rejects the new parameter like model/max_turns;
responses()/messages() are local-only already.

* feat: extra_body escape hatch on the three protocol doors

The industry-standard per-request extension channel (openai/anthropic
SDK trio): a dict merged verbatim into the backend request, last, so
caller keys win over SDK-set ones. Routing per door: OpenAI-compatible
destinations get a true body merge (ModelSettings.extra_body / the
anthropic SDK's native extra_body); LiteLLM-routed providers take the
keys as LiteLLM's own top-level kwargs instead, since LiteLLM plants
extra_body as literal fields other providers reject. Cloud mode rejects
it like the other local-only knobs. chat() stays answer-only.

* chore: litellm floor 1.84 -> 1.97.0

1.97.0 is where the unset-effort chatcmpl->responses bridge landed
(responses_api_bridge_check's on_constraint_enforcing_endpoint arm,
A/B-verified against 1.96.2), so sol-class models work through the
chat lane out of the box instead of 400ing until a manual upgrade.
Three spots move together: the pyproject floor, the requirements.txt
CI pin, and the install hint.

* feat: named top_p/max_tokens on the chat door, max_output_tokens on responses

extra_body could not carry these: openai-agents' LitellmModel passes
every ModelSettings sampling field as an explicit keyword and unpacks
extra_args into the same call, so the common knobs collided with a bare
TypeError on the LiteLLM lane (reproduced against a stub — Python call
semantics, callee-independent). Named params ride ModelSettings fields,
the one channel clean on every lane; responses() uses the protocol's
own name (openai_responses maps ModelSettings.max_tokens to
max_output_tokens on the wire) and the envelope now echoes the real
value instead of a constant None. Both caps bound each backend call in
the agent loop, not the whole run — documented. Cloud rejects them like
the other local-only knobs; the long tail (frequency_penalty etc.)
stays extra_body-blocked-loudly on that lane by choice.

* chore: the missing-key error names chat_model as the other exit

The default chat_model is what put keyless users on the OpenAI lane,
so the error now points at the knob that picks a different backend.
2026-08-17 00:42:17 +08:00
2026-03-27 03:30:13 +08:00
2025-04-01 18:54:08 +08:00

PageIndex Banner

VectifyAI%2FPageIndex | Trendshift

PageIndex: Vectorless, Reasoning-based RAG

Reasoning-based RAG  ◦  No Vector DB, No Chunking  ◦  Context-Aware Retrieval  ◦  Reads Like a Human

🌐 Website  •   🖥️ Chat Platform  •   🔌 MCP & API  •   📖 Docs  •   💬 Discord  •   ✉️ Contact 

📢 Updates

  • 🔥 Agentic Vectorless RAG — A simple agentic, vectorless RAG example with self-hosted PageIndex, using OpenAI Agents SDK.
  • Scale PageIndex to Millions of Documents — PageIndex File System is a file-level tree indexing layer that lets PageIndex reason over an entire corpus, not just a single document, enabling massive-scale document search.
  • PageIndex Chat — Human-like document analysis agent platform for professional long documents. Also available via MCP or API.
  • PageIndex Framework — Deep dive into PageIndex: an agentic, in-context tree index that enables LLMs to perform reasoning-based, context-aware retrieval over long documents.

📑 Introduction to PageIndex

Are you frustrated with vector database retrieval accuracy for long professional documents? Traditional vector-based RAG relies on semantic similarity rather than true relevance. But similarity ≠ relevance — what we truly need in retrieval is relevance, and that requires reasoning. When working with professional documents that demand contextual understanding, domain expertise, and multi-step reasoning, similarity search often falls short — missing what's relevant but not similar, and returning what's similar yet not relevant.

Inspired by AlphaGo, we propose PageIndex — a vectorless, reasoning-based RAG system that builds a hierarchical tree index from long documents, and uses LLMs to reason over that index for agentic, context-aware retrieval. The retrieval is traceable and explainable, with no vector DBs or chunking. PageIndex simulates how human experts navigate and extract knowledge from complex documents through tree search, enabling LLMs to think and reason their way to the most relevant document sections. It performs retrieval in two steps:

  1. Generate a “Table-of-Contents” tree structure index of documents
  2. Perform (agentic) reasoning-based retrieval through tree search

🎯 Core Features

PageIndex is a vectorless, reasoning-based RAG engine that mirrors how humans read, delivering traceable, explainable, and context-aware retrieval, without vector databases or chunking.

Compared to traditional vector-based RAG, PageIndex features:

  • No Vector DB: Uses document structure and LLM reasoning for retrieval, instead of vector similarity search.
  • No Chunking: Documents are organized into natural sections, not artificial chunks.
  • Better Traceability & Explainability: Retrieval is reasoning-driven and grounded in explicit page and section references, making every result traceable and interpretable — no more “vibe retrieval” with opaque, approximate vector search.
  • Context-Aware Retrieval: Retrieval depends on your full context (e.g., conversation history and domain knowledge), and easily incorporates new context.
  • Human-like Retrieval: Mirrors how human experts navigate and extract knowledge from complex documents.

PageIndex achieved state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), vastly outperforming vector RAG solutions on professional document analysis (blog post).

📍 Explore PageIndex

To learn more, please see a detailed introduction to the PageIndex framework. Check out our GitHub for open-source code, and the cookbooks, tutorials, and blog for more usage guides and examples.

The PageIndex service is available as a ChatGPT-style chat platform, or can be integrated via MCP or API, with enterprise deployment available.

🛠️ Deployment Options

  • Self-host — run locally with this open-source repo (using standard PDF parsing).
  • Cloud Service — production-grade pipeline with enhanced OCR, tree building, and retrieval for best results. Try instantly on our Chat Platform, or integrate via MCP or API.
  • Enterprise — dedicated or private deployment (VPC, on-prem). Contact us or book a demo to learn more.

🧪 Quick Hands-on

  • ⚡ PageIndex Flash (preview) — ultra fast PageIndex tree structure generation from PDFs.
  • 🔥 Agentic Vectorless RAG (latest) — a simple but complete agentic vectorless RAG example with self-hosted PageIndex, using OpenAI Agents SDK.
  • Try the Vectorless RAG notebook — a minimal, hands-on example of reasoning-based RAG using PageIndex.
  • Check out Vision-based Vectorless RAG — no OCR; a minimal, vision-based & reasoning-native RAG pipeline that works directly over page images.
View on GitHub: Agentic Vectorless RAG
Open in Colab: Vectorless RAG    Open in Colab: Vision RAG

🌲 PageIndex Tree Structure

PageIndex can transform lengthy PDF documents into a semantic tree structure, similar to a “table of contents” but optimized for use with LLMs and AI agents. It's ideal for: financial reports, legal documents, regulatory filings, technical manuals, medical literature, academic textbooks, and any long, complex professional documents.

Below is an example PageIndex tree structure. Also see more example documents and generated tree structures.

...
{
  "title": "Financial Stability",
  "node_id": "0006",
  "start_index": 21,
  "end_index": 22,
  "summary": "The Federal Reserve ...",
  "nodes": [
    {
      "title": "Monitoring Financial Vulnerabilities",
      "node_id": "0007",
      "start_index": 22,
      "end_index": 28,
      "summary": "The Federal Reserve's monitoring ..."
    },
    {
      "title": "Domestic and International Cooperation and Coordination",
      "node_id": "0008",
      "start_index": 28,
      "end_index": 31,
      "summary": "In 2023, the Federal Reserve collaborated ..."
    }
  ]
}
...

You can generate PageIndex tree structures with this open-source repo. Or use our API for higher-quality results powered by our enhanced OCR and tree building pipeline.


⚙️ Package Usage

Note: This package uses standard PDF parsing. For use cases with complex PDFs, our cloud service (via MCP and API) offers enhanced OCR, tree building, and retrieval.

You can follow these steps to generate a PageIndex tree from a PDF document.

1. Install dependencies

pip3 install --upgrade -r requirements.txt

2. Set your LLM API key

Create a .env file in the root directory with your LLM API key. Multi-LLM is supported via LiteLLM:

OPENAI_API_KEY=your_openai_key_here

3. Generate PageIndex structure for your PDF

python3 run_pageindex.py --pdf_path /path/to/your/document.pdf
Optional parameters
You can customize the processing with additional optional arguments (the structure-tuning flags below require --mode standard):
--mode                  Processing mode: flash (default) or standard
--index-model           LLM model used to index the document (default: gpt-5.6-luna)
--toc-check-pages       Pages to check for table of contents (default: 20)
--max-pages-per-node    Max pages per node (default: 10)
--max-tokens-per-node   Max tokens per node (default: 20000)
--if-add-node-id        Add node ID (yes/no, default: yes)
--if-add-node-summary   Add node summary (yes/no, default: yes)
--if-add-doc-description Add doc description (yes/no, default: yes)
Markdown support
We also provide markdown support for PageIndex. You can use the `--md_path` flag to generate a tree structure for a markdown file.
python3 run_pageindex.py --md_path /path/to/your/document.md

Note: in this mode, we use "#" to determine node headings and their levels. For example, "##" is level 2, "###" is level 3, etc. Make sure your markdown file is formatted correctly. If your Markdown file was converted from a PDF or HTML, we don't recommend using this mode, since most existing conversion tools cannot preserve the original hierarchy. Instead, use our PageIndex OCR, which is designed to preserve it, to convert the PDF to a markdown file and then use this mode.

⚡ PageIndex Flash (preview)

PageIndex Flash (pageindex/flash) generates tree structures from PDFs in seconds. Structure extraction is purely heuristic-based, no LLM needed. An LLM is used only for node summaries and the optimization's expansion pass.

python3 run_pageindex.py --mode flash --pdf_path /path/to/your/document.pdf

Tree optimization for retrieval (a deterministic merge, then an LLM expansion pass) is on by default; pass --optimize off to disable.

🚀 Agentic Vectorless RAG: An Example

For a simple, end-to-end agentic vectorless RAG example using self-hosted PageIndex (with OpenAI Agents SDK), see examples/agentic_vectorless_rag_demo.py.

python3 examples/agentic_vectorless_rag_demo.py

📈 Case Study: PageIndex Leads Finance QA Benchmark

Mafin 2.5 is a reasoning-based RAG system for financial document analysis, powered by PageIndex. It achieved a state-of-the-art 98.7% accuracy on FinanceBench (financial document QA benchmark), significantly outperforming traditional vector-based RAG systems.

PageIndex's hierarchical indexing and reasoning-driven retrieval enable precise navigation and extraction of relevant context from complex financial reports, such as SEC filings and earnings disclosures.

Explore the full benchmark results and our blog post for detailed comparisons and performance metrics.


🧭 Resources

  • 📝 Blog: technical articles, research insights, and product updates.
  • 🔧 Developer: MCP setup, API docs, and integration guides.
  • 🧪 Cookbooks: hands-on, runnable examples and advanced use cases.
  • 📖 Tutorials: practical guides and strategies, including Document Search and Tree Search.

⭐ Support Us

Leave us a star 🌟 if you like our project. Thank you!

Please cite this work as:

Mingtian Zhang, Yu Tang and PageIndex Team,
"PageIndex: Next-Generation Vectorless, Reasoning-based RAG",
PageIndex Blog, Sep 2025.
Or use the BibTeX citation.
@article{zhang2025pageindex,
  author = {Mingtian Zhang and Yu Tang and PageIndex Team},
  title = {PageIndex: Next-Generation Vectorless, Reasoning-based RAG},
  journal = {PageIndex Blog},
  year = {2025},
  month = {September},
  note = {https://pageindex.ai/blog/pageindex-intro},
}

🌐 Open-Source Ecosystem

PageIndex anchors a growing open-source ecosystem of long-context AI infra — OpenKB is an LLM knowledge base that compiles documents into an interlinked wiki. ChatIndex provides tree indexing and retrieval for long conversational histories and memory. ConDB is a KV-cache native context database for tree-based retrieval at scale. PageIndex MCP is PageIndex's MCP server.

Connect with Us

Website  Twitter  LinkedIn  Discord  Book a Demo  Contact Us


© 2026 Vectify AI

Languages
Python 99.5%
JavaScript 0.4%
Shell 0.1%