Benchmark

Measured, not promised.

Grape against Vector RAG on the same documents, the same questions and the same LLM.

Summary

The short version.

Correct answers, English

28 / 28Grape

vs 24 / 28 Vector RAG

4 more right, 5% cheaper
Ready after a change

0.1 sGrape

vs 97.7 s Vector RAG

About 1,000x sooner
Gujarati questions, English documents

28 / 28Grape

vs 5 / 28 Vector RAG

Against an English embedder
Right passage ranked first

78.2%Grape

vs 69.8% Vector RAG

SQuAD 2.0, no LLM
Wins
Accuracy and cost in English, questions in another language, and time to take in a change.
Loses
1 to 2 fewer right answers on Gujarati documents and the 18 MB corpus; Gujarati questions cost more.
Method
Wikipedia articles, Claude Haiku with thinking off, Grape's server reranker, every suite run three times.

Wikipedia articles, questions written before each run, Claude Haiku with thinking off, every suite run three times. Grape used its server reranker and sent 5 passages; clearly off-topic English questions were refused with no LLM call. Vector RAG used bge-small embeddings (multilingual-e5-large for Gujarati) and the top 5 chunks. Anthropic list prices. Every question and answer is in the full report, and the benchmark is open source.

Correct answers

More right in English, and in more languages.

English documents, 28 questions Higher is better

Grape28 / 28

Vector RAG24 / 28

Large corpus, 18 MB, 29 questions Higher is better

Grape27.7 / 29

Vector RAG29 / 29

Gujarati documents and questions Higher is better

Grape24.7 / 28

Vector RAG27 / 28

Off-topic questions refused (English) Higher is better

Grape5 / 5

Vector RAG5 / 5

Cost and speed

Ready a thousand times sooner.

Vector RAG re-embeds a document before it can answer from it; Grape re-indexes it in a tenth of a second. Per question, Grape reads 7% fewer tokens in English and 59% fewer on Gujarati documents.

Time to take in a changed document Lower is better

Grape0.1 s

Vector RAG97.7 s

Input tokens per question Lower is better

Grape1,217

Vector RAG1,307

Cost per question Lower is better

Grape$0.00148

Vector RAG$0.00155

LLM calls per question Lower is better

Grape1.12

Vector RAG1.00

Input tokens per question, Gujarati documents Lower is better

Grape1,684

Vector RAG4,137

Retrieval alone

Finds the right passage first, more often.

Right passage ranked first Higher is better

Grape78.2%

Vector RAG69.8%

Right passage in the top 8 Higher is better

Grape96.5%

Vector RAG96.0%

Ranked first, with a reranker Higher is better

Grape90.0%

Vector RAG89.5%

Search time per question (CPU) Lower is better

Grape10 ms

Vector RAG28 ms

Large corpus

Eighteen megabytes, indexed in a second.

300 long Wikipedia articles. Both answer almost everything and cost about the same per question; Vector RAG spends 46 minutes of CPU embedding them first, and again after every change.

Correct answers Higher is better

Grape27.7 / 29

Vector RAG29 / 29

Time to index 18 MB Lower is better

Grape1.1 s

Vector RAG46 min

Input tokens per question Lower is better

Grape1,413

Vector RAG1,391

Cost per question Lower is better

Grape$0.00167

Vector RAG$0.00163

Gujarati

Asked in Gujarati, answered from English documents.

The same twelve English articles, every question written in Gujarati. Grape rewrites the question in English and searches again, which costs one small extra LLM call. An English embedding model finds almost nothing; a large multilingual one needs 12 minutes of embedding first.

Against an English embedding model (bge-small) Higher is better

Grape28 / 28

RAG, English5 / 28

Against a multilingual embedding model (e5-large) Higher is better

Grape26.7 / 28

RAG, multilingual26 / 28

Cost per question, against e5-large Lower is better

Grape$0.00234

RAG, multilingual$0.00199

Time to index the documents Lower is better

Grape0.1 s

RAG, multilingual12.5 min

Where Grape loses

Two-part questions, and one extra call.

On Gujarati documents Grape answered 24.7 of 28 against 27, and on the 18 MB corpus 27.7 of 29 against 29: with 5 passages, a question with two parts in two articles sometimes misses one half. Gujarati questions on English documents cost about 18% more than a multilingual Vector RAG, because Grape spends a small call rewriting the question in English. On the 18 MB corpus the cost is about even (2% more tokens). Each Grape answer also takes under a second longer.