A labeled test set to score the search
Real questions, each paired with the documents it should return. We score the search against that list, apart from the answers.
RAG development company in India
We build retrieval-augmented generation (RAG) systems: AI assistants that search your documents first, then answer from what they found. We score that search on its own, so a missed document shows up in testing.
What we build
Real questions, each paired with the documents it should return. We score the search against that list, apart from the answers.
Embeddings (numbers that capture what a passage means) match meaning. Keyword search matches exact strings such as an Indian tax ID (GSTIN or PAN) or an invoice number. Scores on your questions decide the mix.
The index, the searchable copy of your documents, filters results by each person's permissions. It keeps its own copy of them, so a change in your source systems can take time to reach it. We plan how that change gets there.
Answering customers? See AI chatbot development. Drafting documents too? See generative AI development. For all our AI work, see AI development.
Retrieval versus the answer
A RAG system has two halves. The retriever finds passages in your documents, and the model writes an answer from them. A grade on the answer grades both at once.
Grading the search only on what it returned has a blind spot too: a missing document leaves nothing to score. You see the miss only if someone listed, in advance, the documents each question should find. Those lists are relevance labels.
So we usually measure retrieval separately from answer quality. Anthropic, Microsoft and the TREC 2025 RAG Track, a research evaluation from the US standards agency NIST, do the same.
Need retrieval at all? LLM development weighs prompts, retrieval and fine-tuning side by side.
Metrics by publisher
Ground truth is the known right result for each question: the documents to find (relevance labels) or the answer to give.
| Metric, and whose it is | What it scores | Needs ground truth? | What it can't see |
|---|---|---|---|
| Retrieval, Microsoft | How relevant the retrieved chunks are to the query, judged by an AI model | No: only the query and the context | A relevant document never retrieved |
| Document Retrieval, Microsoft: Fidelity, NDCG, XDCG, Max Relevance N, Holes | Fidelity: good documents returned out of all known good documents | Yes, relevance labels for each query | The answer. Holes counts unlabeled documents, a gap in the labels, not the system |
| Groundedness, and Groundedness Pro (preview), Microsoft | Answer-side. Groundedness: whether what the answer says is backed by the retrieved text | No | What the answer left out, and whether retrieval supplied the right context |
| Relevance, Microsoft | The quality of the final response | No | A better document that retrieval never returned |
| Response Completeness (preview), Microsoft | How much of the expected answer the response covers | Yes, a ground-truth answer | Whether missing information was never retrieved or was left out |
| Context relevance, Amazon Bedrock (retrieve-only) | How relevant the retrieved texts are to the questions | No | Ground-truth information that retrieval did not return |
| Context coverage, Amazon Bedrock (retrieve-only) | How much of the ground truth the retrieved texts cover | Yes, ground-truth texts | Whether the answer used what was retrieved |
| TREC 2024 RAG Search Track data, NIST | Document and passage relevance judgments, plus nugget and citation assessments | They are the ground truth, for the track's own topics | Your own documents and queries |
How we work
Steps 01 to 03 follow Microsoft's published guidance on tuning retrieval. Grading the answer and the access check are ours.
We start from what people have already asked. People who know your documents list the ones each question should return.
The same questions run through different search methods (meaning-based vector search, or Microsoft's semantic re-ranking), result counts per query, and chunk sizes with enough overlap to keep context.
We keep the setting with the highest retrieval score, then build and grade the answer step on top.
Our security team, not the builders, checks that each user gets back only what they may see, as part of LLM security testing, within our AI security work.
For your engineers
Publishers document each choice; your labeled questions pick the winner.
Anthropic says traditional RAG solutions "remove context when encoding information", so relevant information often goes unretrieved. Its measure is 1 minus recall@20: relevant documents missing from the top 20 chunks. Contextual Embeddings cut that by 35% (5.7% to 3.7%), and adding Contextual BM25 by 49% (5.7% to 2.9%), averaged across its tested domains with its best embedding configuration.
The TREC 2025 RAG Track overview (University of Waterloo, Microsoft and Zipf AI; NIST's TREC 34 proceedings) sets four tasks: Retrieval, Augmented Generation, full RAG and a new Relevance Judgment task. Its generation task fixes retrieval, so every team works from the same input.
The Amazon Bedrock User Guide documents default chunking (about 300 tokens a chunk), fixed-size chunking (you set a token count and an overlap percentage), and hierarchical, semantic and no chunking. Hierarchical chunking can return fewer results than requested, as parent chunks replace child chunks. With no chunking, citations can't show a page number and results can't be filtered by it. Anthropic notes that chunk size, boundary and overlap all affect retrieval.
Microsoft fuses ranked lists with Reciprocal Rank Fusion (RRF): each list gives a document 1/(rank + k), and its scores from all lists are added. Here k is a small constant such as 60, unrelated to top_k (results per query) or the nearest-neighbor k. Ranker scales differ: keyword scores have no upper limit, cosine vector scores run 0.333 to 1.00 (0 to 1 for Euclidean and dot product), and the semantic ranker, applied after the merge, runs 0.00 to 4.00. The TREC 2025 RAG Track baseline also fused sparse (keyword) and dense (vector) retrieval.
Anthropic's tests found reranking better than none. Google's ranking API documentation says a ranker scores how well a document answers the query; embeddings measure only semantic similarity. A reranker reorders what came back, so it can't recover a missed document.
Approximate nearest neighbor (ANN) search trades some accuracy for speed; Microsoft says its parameters can be tuned for recall, latency, memory and disk. An exhaustive scan gives ground truth for ANN recall, as nearest vectors: it shows what the approximate index dropped, not whether the right documents were found.
Google's Gemini API documentation embeds documents and queries differently: separate task types (RETRIEVAL_DOCUMENT, RETRIEVAL_QUERY) for its earlier model, the task written into the input text for its newer one. See AI development for where the model sits in the system.
Microsoft's document-level access documentation, for its 2026-08-01-preview REST API, notes a timing lag before permission changes are recognized. What a failed check could disclose is covered by generative AI security testing.
Microsoft suggests language analyzers for non-Western text, since splitting words at spaces, hyphens and slashes suits Western languages better. The ten Eighth Schedule languages in its analyzer table are Bangla, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu and Urdu.
MIRACL, the WSDM 2023 Cup benchmark from the University of Waterloo and Huawei Noah's Ark Lab, has relevance judgments in 18 languages, three of them Eighth Schedule languages (Bengali, Hindi, Telugu), and focuses on monolingual retrieval.
The Unicode normalization FAQ says canonical-equivalent strings should compare as equal. Normalize documents and queries the same way, or a query can miss a word that looks identical on screen.
FAQ
Wrong answers, labels, languages and erasure. For anything else, ask us directly.
A score on the answer alone can't tell you: the retriever may have handed over the wrong passages, or the model may have misused the right ones. Microsoft's RAG evaluator documentation notes that poor retrieval leaves the model less chance of a satisfactory answer. So we usually score retrieval separately.
Labels: a list of the documents each question should find. Without them, a missed document can't be measured. Microsoft's Fidelity uses them; Amazon Bedrock's Context coverage needs ground-truth texts instead. Microsoft's Retrieval evaluator and Bedrock's Context relevance need neither, and grade only what came back.
When people search by exact strings. Microsoft's hybrid search documentation names product codes, specialized jargon, dates and people's names. Anthropic's example is the error code TS-999: an embedding model may find content about error codes in general and miss the exact match. A GSTIN or a PAN has the same shape. Only scoring your real queries shows whether hybrid beats vector-only search.
It has to be rebuilt. Microsoft's vector search documentation says vector queries run against vectors made by one embedding model, and the query goes through that same model.
They change the questions asked first. Counted off Microsoft's supported-analyzer table, ten of the 22 Eighth Schedule languages appear, only Hindi has both a Microsoft and a Lucene analyzer, and twelve appear nowhere. With multilingual embeddings, Microsoft says, vector queries can find a match with no analyzer or translation. Which languages your documents hold, and whether users type them in Latin script, only you can answer.
Section 8(7)(b) of the DPDP Act, 2023, says a Data Fiduciary shall, unless retention is necessary for compliance with any law, "cause its Data Processor to erase any personal data that was made available by the Data Fiduciary for processing to such Data Processor." It commences in tranche (c), eighteen months after the Rules were published, which our DPDP Act compliance page dates to May 2027, and is subject to the Act's exemptions. For a retrieval system, the design question is whether each chunk and its embedding trace back to their source record, so an erasure reaches the index too.
Let's talk
Share what you're building, which documents it answers from and what your users ask. Project work is delivered remotely from India during business hours. We reply within one working day.