Blog / / 2 min read

Grape vs Vector RAG, measured

Same documents, same questions, same LLM, three runs each. In English Grape is more accurate and cheaper than Vector RAG, it reaches across languages, and it is ready after an edit in a tenth of a second.

Vector RAG is the usual way to let an LLM answer from your documents. It works, but it has two costs you feel later: every edited file must be embedded again, and a question in another language often finds nothing.

Grape skips the embeddings. It keeps a keyword index and ranks the lines that best answer the question. We wanted to know what that costs in answer quality, so we ran an open benchmark.

How we tested

  • Documents: Wikipedia articles. 12 English ones, 12 Gujarati ones, and 300 long English ones (18 MB).
  • Questions: 28 per set, written before each run. Some are asked in Hindi, Spanish or French, and a few are off-topic on purpose.
  • LLM: Claude Haiku, thinking off, the same instructions for both.
  • Grape: its server reranker picks the best 5 passages, and clearly off-topic English questions are refused with no LLM call.
  • Runs: every set three times. The numbers are means.

Accuracy: ahead in English, far ahead across languages

Set Grape Vector RAG
English documents 28 / 28 24 / 28
Large corpus, 18 MB 27.7 / 29 29 / 29
Gujarati documents 24.7 / 28 27 / 28
Gujarati questions, English documents 28 / 28 5 / 28

Asked in Gujarati, answered from English documents: 28 of 28 for Grape, 5 of 28 for Vector RAG with an English embedder.

Vector RAG's misses on the English set were the questions asked in other languages. An English embedding model puts them far from the English text. Grape notices when a search comes back weak, rewrites the question in the documents' language and searches again.

What it means for you: if your users don't all write in the language of your docs, Grape keeps working without a bigger, slower embedding model.

Speed: ready after an edit

Set Grape Vector RAG
12 articles 0.1 s 98 s
300 articles, 18 MB 1.1 s 46 min

An edited file is searchable again in 0.1 seconds. There is nothing to re-embed.

This is CPU time on the same machine. A GPU would shrink the Vector RAG column, but you would pay it again on every upload and every edit.

Cost: cheaper in English

English, per question Grape Vector RAG
Input tokens 1,217 1,307
LLM calls 1.12 1.00
Cost $0.00148 $0.00155

On Gujarati documents Grape reads 1,684 tokens per question against 4,137. On the 18 MB corpus the two are about even.

Where Grape loses

  • Two-part questions. With 5 passages, a question whose two halves sit in two different articles sometimes misses one half. That cost Grape 2.3 questions on Gujarati documents and 1.3 on the 18 MB corpus.
  • One extra call across languages. A Gujarati question on English documents costs about 18% more than a multilingual Vector RAG, because Grape spends a small call rewriting it in English. The multilingual embedder needed 12.5 minutes to index the documents first.

Check it yourself

Every question and both answers are in the full report. The benchmark is open source. Want to try it on your own documents? Sign up free: 500 MB, and it works in Claude and ChatGPT.