Measured, not promised.
Grape against Vector RAG on the same documents, the same questions and the same LLM.
The short version.
28 / 28Grape
vs 24 / 28 Vector RAG
4 more right, 5% cheaper0.1 sGrape
vs 97.7 s Vector RAG
About 1,000x sooner28 / 28Grape
vs 5 / 28 Vector RAG
Against an English embedder78.2%Grape
vs 69.8% Vector RAG
SQuAD 2.0, no LLM- Wins
- Accuracy and cost in English, questions in another language, and time to take in a change.
- Loses
- 1 to 2 fewer right answers on Gujarati documents and the 18 MB corpus; Gujarati questions cost more.
- Method
- Wikipedia articles, Claude Haiku with thinking off, Grape's server reranker, every suite run three times.
Wikipedia articles, questions written before each run, Claude Haiku with thinking off, every suite run three times. Grape used its server reranker and sent 5 passages; clearly off-topic English questions were refused with no LLM call. Vector RAG used bge-small embeddings (multilingual-e5-large for Gujarati) and the top 5 chunks. Anthropic list prices. Every question and answer is in the full report, and the benchmark is open source.
More right in English, and in more languages.
English documents, 28 questions Higher is better
Large corpus, 18 MB, 29 questions Higher is better
Gujarati documents and questions Higher is better
Off-topic questions refused (English) Higher is better
Ready a thousand times sooner.
Vector RAG re-embeds a document before it can answer from it; Grape re-indexes it in a tenth of a second. Per question, Grape reads 7% fewer tokens in English and 59% fewer on Gujarati documents.
Time to take in a changed document Lower is better
Input tokens per question Lower is better
Cost per question Lower is better
LLM calls per question Lower is better
Input tokens per question, Gujarati documents Lower is better
Finds the right passage first, more often.
Right passage ranked first Higher is better
Right passage in the top 8 Higher is better
Ranked first, with a reranker Higher is better
Search time per question (CPU) Lower is better
Eighteen megabytes, indexed in a second.
300 long Wikipedia articles. Both answer almost everything and cost about the same per question; Vector RAG spends 46 minutes of CPU embedding them first, and again after every change.
Correct answers Higher is better
Time to index 18 MB Lower is better
Input tokens per question Lower is better
Cost per question Lower is better
Asked in Gujarati, answered from English documents.
The same twelve English articles, every question written in Gujarati. Grape rewrites the question in English and searches again, which costs one small extra LLM call. An English embedding model finds almost nothing; a large multilingual one needs 12 minutes of embedding first.
Against an English embedding model (bge-small) Higher is better
Against a multilingual embedding model (e5-large) Higher is better
Cost per question, against e5-large Lower is better
Time to index the documents Lower is better
Two-part questions, and one extra call.
On Gujarati documents Grape answered 24.7 of 28 against 27, and on the 18 MB corpus 27.7 of 29 against 29: with 5 passages, a question with two parts in two articles sometimes misses one half. Gujarati questions on English documents cost about 18% more than a multilingual Vector RAG, because Grape spends a small call rewriting the question in English. On the 18 MB corpus the cost is about even (2% more tokens). Each Grape answer also takes under a second longer.