Ask a vector database “what is the exception for returns?” and it hands back the five paragraphs that sound most like your question. The paragraph that actually answers it often sounds like nothing, because it says “see section 4.3” and lives under a heading three levels up. Vectorless RAG goes after exactly that failure. PageIndex is the MIT-licensed project doing it, it shipped v0.2.21 on October 1, and the idea is simple enough to explain in one breath: build a table of contents for the document, then let the LLM read it and decide where to open the book.
How vectorless RAG works in PageIndex
The normal recipe is chunk, embed, store, then take the top-K nearest vectors for the question. PageIndex’s README names the problem in one line: similarity is not relevance. A chunk can be very similar to your question and still be the wrong place, and the right place can look nothing like it.
So PageIndex does two things instead. At index time it builds a tree of the document: sections, subsections, page ranges, and a short summary on each node. At question time the LLM reads that tree, picks a branch, reads the pages under it, and keeps going until it has the answer. The README’s phrase is “reads like a human”, which is a bit much, but the mechanism is real. You can print every node it visited, so when the answer is wrong you can see where it took the wrong turn. Try doing that with a cosine score of 0.81.
The setup is small. pip install -U pageindex, then PageIndexClient(), submit_document("report.pdf") and chat("What does the report conclude?"). According to the v0.2.21 release notes, local mode needs no server, no vector DB and no PageIndex API key, only your own LLM key. The default engine, called Flash, takes the tree from the PDF’s layout and bookmarks, so no LLM is involved in the structure. The model only writes the node summaries.
That is the whole trick behind vectorless RAG. There is no embedding model in the loop, and no chunk size to tune or overlap setting to argue about on a Tuesday.

What the numbers say, and who produced them
Every number in this section comes from the project itself, so read them as claims with public code, not as a verdict.
Indexing is cheap. The README says about $0.001 per page with gpt-5.6-luna as the index model, so a 1,000-page book is a bit over a dollar, once. The 0.2.20 release notes say a 222-page report indexes in 73 seconds and a 758-page book in 137.
Answering is where the money goes. The project’s PageIndex-OSS-Benchmark runs 62 questions over 34 PDFs, 1,945 pages in total. With the cheap chat model at high reasoning effort it got 60 of 62 right, 96.8%, at $0.0036 a question. The top model got 62 of 62 at roughly $0.08 a question. Mettons you ask 1,000 questions a day: that is $3.60 on the cheap setup and about $82 on the top one. My arithmetic, their prices.
Two things that benchmark does not do. It has no vector-database baseline, so it cannot tell you PageIndex beats anything. And the authors kept only questions whose answer is a plain fact in running text: no charts, no tables, no arithmetic. For a financial filing, tables are half the job.
The famous number is 98.7% on FinanceBench. The source is Vectify’s Mafin 2.5 write-up from February 2025, and Mafin 2.5 is the company’s own product built on PageIndex, not the repo you install. The accompanying chart says plain vector RAG scored 50%, but I could not find the baseline’s setup, so I would not repeat that comparison to anyone. What I can say is that the evaluation code is public and the benchmark itself, FinanceBench, is a real academic dataset of questions over SEC filings.
Where it loses to a vector database
Latency first. A vector lookup is one embedding call and one index read. Vectorless RAG makes the model take several steps down the tree, and each step is a round trip to an LLM. I have no latency table from the project, which is itself a finding, so measure it before you promise anyone a chat box that answers in two seconds.
Cost per query, same story, and it scales with how much you ask. If I add the numbers up the way I did in my post about where an agent pipeline’s API bill actually went, the retrieval step is the part that stops being free.
Many small documents are the third loss. A tree over a two-page FAQ is a tree with one branch. Embeddings were built for thousands of short passages, and for those they are fine. PageIndex’s own answer for huge corpora is the PageIndex File System, announced May 3, 2026, but that ships in the enterprise and cloud product, not in the MIT repo.
Then the churn. In ten days the project went from 0.2.19 to 0.2.21, and two of those releases carry BREAKING in the notes: resolve_citations renamed to get_citations in 0.2.19, and the tree node shape changed in 0.2.21, with page_index and prefix_summary gone and start_index, end_index and summary in their place. Pin the version.
Last one: the open-source path is built for text-heavy PDFs. The README says so itself, and scanned documents or heavy OCR work point you to the cloud product.
Who should try it, who should keep the vector DB
My seat is Trellis, so my example is Amazon sellers. The documents they live in are long, nested and full of “except as described in” sentences: policy pages, fee schedules, category rules. That is exactly the shape where a chunk that matches your question is the wrong chunk. A seller asking whether a rule applies to their category needs the general rule, the exception, and the category table, and those sit in three places with three different vocabularies.
Me, I have not run PageIndex on our documents yet, so I will not tell you it works there. The test I would run first is small: take 30 real questions from support tickets, pick the 10 long documents they point to, and answer them twice, once with your current chunk-and-embed setup and once with PageIndex on the cheapest model. Count wrong answers, not vibes, and write down the cost per question for both.
Try vectorless RAG if your corpus is a few dozen long, structured PDFs, each one queried many times, and a wrong answer costs you more than a few cents. Contracts, filings, manuals, compliance docs.
Keep your vector database if you have thousands of short documents, if you need answers in under a second, or if your retrieval bill has to be a rounding error. A support-article search box is the standard example.
Both can live in one product, and vectorless RAG does not have to replace anything to be useful. Embeddings pick the right document out of a thousand, and the tree finds the right page inside it. That is my guess, not something I have measured, and I will know in a week what the 30-question test says.