Home All services
Start a project → Call Now

RAG development company in India

Know when your AI search misses a document — before your users do.

We build retrieval-augmented generation (RAG) systems: AI assistants that search your documents first, then answer from what they found. We score that search on its own, so a missed document shows up in testing.

  • Search scored apart from answers
  • Permissions filtered inside the search
  • Access checked by our security team, not the builders
Illustration: a question sends a beam of light into a library of blank document cards; three relevant cards lift out and flow into an AI core, which writes an answer linked back to them

In brief

What it is
An AI assistant that looks things up in your documents before it answers. The look-up step is called retrieval.
Why it matters
If the search misses the right document, the answer suffers, and grading the answer won't tell you why.
What you get
A RAG system on your documents, plus a test set of real questions and the documents each one should find.

What we build

A search you can measure, and answers built on it.

A labeled test set to score the search

Real questions, each paired with the documents it should return. We score the search against that list, apart from the answers.

Keyword and meaning search

Embeddings (numbers that capture what a passage means) match meaning. Keyword search matches exact strings such as an Indian tax ID (GSTIN or PAN) or an invoice number. Scores on your questions decide the mix.

Permissions inside the index

The index, the searchable copy of your documents, filters results by each person's permissions. It keeps its own copy of them, so a change in your source systems can take time to reach it. We plan how that change gets there.

Answering customers? See AI chatbot development. Drafting documents too? See generative AI development. For all our AI work, see AI development.

Retrieval versus the answer

A wrong answer doesn't say which half failed.

A RAG system has two halves. The retriever finds passages in your documents, and the model writes an answer from them. A grade on the answer grades both at once.

Grading the search only on what it returned has a blind spot too: a missing document leaves nothing to score. You see the miss only if someone listed, in advance, the documents each question should find. Those lists are relevance labels.

So we usually measure retrieval separately from answer quality. Anthropic, Microsoft and the TREC 2025 RAG Track, a research evaluation from the US standards agency NIST, do the same.

What two numbers tell you

  • Low retrieval score. Look at chunking (how documents are split), ranking or the index.
  • Good retrieval, bad answer. Look at the generation step.

Need retrieval at all? LLM development weighs prompts, retrieval and fine-tuning side by side.

Metrics by publisher

Only scores with ground truth can see a missing document.

Ground truth is the known right result for each question: the documents to find (relevance labels) or the answer to give.

Sources: Microsoft Foundry's RAG evaluators documentation, the Amazon Bedrock User Guide and NIST's TREC 2024 RAG Search Track data. The last column is our reading.
Metric, and whose it isWhat it scoresNeeds ground truth?What it can't see
Retrieval, MicrosoftHow relevant the retrieved chunks are to the query, judged by an AI modelNo: only the query and the contextA relevant document never retrieved
Document Retrieval, Microsoft: Fidelity, NDCG, XDCG, Max Relevance N, HolesFidelity: good documents returned out of all known good documentsYes, relevance labels for each queryThe answer. Holes counts unlabeled documents, a gap in the labels, not the system
Groundedness, and Groundedness Pro (preview), MicrosoftAnswer-side. Groundedness: whether what the answer says is backed by the retrieved textNoWhat the answer left out, and whether retrieval supplied the right context
Relevance, MicrosoftThe quality of the final responseNoA better document that retrieval never returned
Response Completeness (preview), MicrosoftHow much of the expected answer the response coversYes, a ground-truth answerWhether missing information was never retrieved or was left out
Context relevance, Amazon Bedrock (retrieve-only)How relevant the retrieved texts are to the questionsNoGround-truth information that retrieval did not return
Context coverage, Amazon Bedrock (retrieve-only)How much of the ground truth the retrieved texts coverYes, ground-truth textsWhether the answer used what was retrieved
TREC 2024 RAG Search Track data, NISTDocument and passage relevance judgments, plus nugget and citation assessmentsThey are the ground truth, for the track's own topicsYour own documents and queries

How we work

Label first, then tune the search against the labels.

Steps 01 to 03 follow Microsoft's published guidance on tuning retrieval. Grading the answer and the access check are ours.

  1. Collect and label real questions

    We start from what people have already asked. People who know your documents list the ones each question should return.

  2. Try candidate settings

    The same questions run through different search methods (meaning-based vector search, or Microsoft's semantic re-ranking), result counts per query, and chunk sizes with enough overlap to keep context.

  3. Keep the best-scoring setting

    We keep the setting with the highest retrieval score, then build and grade the answer step on top.

  4. Have access checked

    Our security team, not the builders, checks that each user gets back only what they may see, as part of LLM security testing, within our AI security work.

For your engineers

Each setting changes what the search returns before any answer is written.

Publishers document each choice; your labeled questions pick the winner.

What the published research measures

Anthropic says traditional RAG solutions "remove context when encoding information", so relevant information often goes unretrieved. Its measure is 1 minus recall@20: relevant documents missing from the top 20 chunks. Contextual Embeddings cut that by 35% (5.7% to 3.7%), and adding Contextual BM25 by 49% (5.7% to 2.9%), averaged across its tested domains with its best embedding configuration.

The TREC 2025 RAG Track overview (University of Waterloo, Microsoft and Zipf AI; NIST's TREC 34 proceedings) sets four tasks: Retrieval, Augmented Generation, full RAG and a new Relevance Judgment task. Its generation task fixes retrieval, so every team works from the same input.

How documents are split into chunks

The Amazon Bedrock User Guide documents default chunking (about 300 tokens a chunk), fixed-size chunking (you set a token count and an overlap percentage), and hierarchical, semantic and no chunking. Hierarchical chunking can return fewer results than requested, as parent chunks replace child chunks. With no chunking, citations can't show a page number and results can't be filtered by it. Anthropic notes that chunk size, boundary and overlap all affect retrieval.

Merging keyword and vector results

Microsoft fuses ranked lists with Reciprocal Rank Fusion (RRF): each list gives a document 1/(rank + k), and its scores from all lists are added. Here k is a small constant such as 60, unrelated to top_k (results per query) or the nearest-neighbor k. Ranker scales differ: keyword scores have no upper limit, cosine vector scores run 0.333 to 1.00 (0 to 1 for Euclidean and dot product), and the semantic ranker, applied after the merge, runs 0.00 to 4.00. The TREC 2025 RAG Track baseline also fused sparse (keyword) and dense (vector) retrieval.

Reranking and approximate search

Anthropic's tests found reranking better than none. Google's ranking API documentation says a ranker scores how well a document answers the query; embeddings measure only semantic similarity. A reranker reorders what came back, so it can't recover a missed document.

Approximate nearest neighbor (ANN) search trades some accuracy for speed; Microsoft says its parameters can be tuned for recall, latency, memory and disk. An exhaustive scan gives ground truth for ANN recall, as nearest vectors: it shows what the approximate index dropped, not whether the right documents were found.

Query and document embeddings

Google's Gemini API documentation embeds documents and queries differently: separate task types (RETRIEVAL_DOCUMENT, RETRIEVAL_QUERY) for its earlier model, the task written into the input text for its newer one. See AI development for where the model sits in the system.

The delay before permission changes reach the index

Microsoft's document-level access documentation, for its 2026-08-01-preview REST API, notes a timing lag before permission changes are recognized. What a failed check could disclose is covered by generative AI security testing.

Indian languages and scripts

Microsoft suggests language analyzers for non-Western text, since splitting words at spaces, hyphens and slashes suits Western languages better. The ten Eighth Schedule languages in its analyzer table are Bangla, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu and Urdu.

MIRACL, the WSDM 2023 Cup benchmark from the University of Waterloo and Huawei Noah's Ark Lab, has relevance judgments in 18 languages, three of them Eighth Schedule languages (Bengali, Hindi, Telugu), and focuses on monolingual retrieval.

The Unicode normalization FAQ says canonical-equivalent strings should compare as equal. Normalize documents and queries the same way, or a query can miss a word that looks identical on screen.

FAQ

Questions about measuring retrieval.

Wrong answers, labels, languages and erasure. For anything else, ask us directly.

A score on the answer alone can't tell you: the retriever may have handed over the wrong passages, or the model may have misused the right ones. Microsoft's RAG evaluator documentation notes that poor retrieval leaves the model less chance of a satisfactory answer. So we usually score retrieval separately.

Labels: a list of the documents each question should find. Without them, a missed document can't be measured. Microsoft's Fidelity uses them; Amazon Bedrock's Context coverage needs ground-truth texts instead. Microsoft's Retrieval evaluator and Bedrock's Context relevance need neither, and grade only what came back.

When people search by exact strings. Microsoft's hybrid search documentation names product codes, specialized jargon, dates and people's names. Anthropic's example is the error code TS-999: an embedding model may find content about error codes in general and miss the exact match. A GSTIN or a PAN has the same shape. Only scoring your real queries shows whether hybrid beats vector-only search.

It has to be rebuilt. Microsoft's vector search documentation says vector queries run against vectors made by one embedding model, and the query goes through that same model.

They change the questions asked first. Counted off Microsoft's supported-analyzer table, ten of the 22 Eighth Schedule languages appear, only Hindi has both a Microsoft and a Lucene analyzer, and twelve appear nowhere. With multilingual embeddings, Microsoft says, vector queries can find a match with no analyzer or translation. Which languages your documents hold, and whether users type them in Latin script, only you can answer.

Section 8(7)(b) of the DPDP Act, 2023, says a Data Fiduciary shall, unless retention is necessary for compliance with any law, "cause its Data Processor to erase any personal data that was made available by the Data Fiduciary for processing to such Data Processor." It commences in tranche (c), eighteen months after the Rules were published, which our DPDP Act compliance page dates to May 2027, and is subject to the Act's exemptions. For a retrieval system, the design question is whether each chunk and its embedding trace back to their source record, so an erasure reaches the index too.

Let's talk

Tell us what your system is meant to find.

Share what you're building, which documents it answers from and what your users ask. Project work is delivered remotely from India during business hours. We reply within one working day.