Blog / / 1 min read

Add Grape to your vector RAG in five minutes

Keep your vector store and add Grape next to it. You get questions in any language and instantly fresh documents, with one extra HTTP call and no new embedding model.

You already have vector RAG, and it mostly works. Two things still hurt: users who ask in another language get nothing back, and an edited document isn't searchable until it is embedded again.

You don't have to replace anything. Put Grape next to your vector store, search both, and hand the LLM the best passages from each.

Where Grape fits

question --+-- your vector store --+
           |                       +-- merge -- LLM -- answer
           +-- Grape search -------+

One extra HTTP call, about 10 ms. No new embedding model, and nothing to re-index when a file changes.

1. Set up (one minute)

  1. Sign up and create a project from your documents. Grape reads PDF, Office files, Markdown, code and audio.
  2. Copy your key from API keys and the project ID from the project page.
  3. Put both in your environment:
export GRAPE_KEY=grape_...
export GRAPE_PROJECT=your-project-id

2. Search both (fifteen lines)

import os, requests

def grape_search(question, limit=5):
    r = requests.post(
        f"https://grape-inc.in/grape/{os.environ['GRAPE_PROJECT']}/search",
        headers={"Authorization": f"Bearer {os.environ['GRAPE_KEY']}"},
        json={"query": question, "limit": limit},
    )
    return [f"{h['path']}:{h.get('line')} {h['text']}" for h in r.json()["hits"]]

def retrieve(question):
    vector_hits = my_vector_store.search(question, k=5)   # what you have today
    passages = [h.text for h in vector_hits] + grape_search(question)
    return list(dict.fromkeys(passages))                  # drop exact duplicates, keep order

3. Use this prompt

Answer the question using only the passages below.
If they do not contain the answer, reply exactly: "Not in the documents."
Answer in the language of the question. Cite the file and line.

Passages:
{passages}

Question: {question}

The last two lines of the instructions do the work. The LLM answers in the user's language, and it refuses instead of guessing.

That's it

You now have multilingual retrieval and documents that are fresh the moment they change, without touching your embedder. When a question in Gujarati or Spanish comes in, Grape finds the English passage your vector store missed.

Next: the docs cover the full search API, and the benchmark shows the numbers behind this.