Answers from your own content
We agree which documents are the source of truth, and it answers from those, given to it directly or, where needed, looked up per question. RAG development covers how that look-up is built and tested.
AI chatbot development company in India
We build customer support chatbots that answer from content you already hold, decline to invent things like a refund policy or a delivery date, and hand over to a person when needed. We test that on your real questions before launch.
What we build
We agree which documents are the source of truth, and it answers from those, given to it directly or, where needed, looked up per question. RAG development covers how that look-up is built and tested.
It has explicit permission to say it doesn't know. Output checks compare its replies with your policies and look for promises it isn't authorized to make.
When a question needs a person, the conversation moves to a place you choose, so a refusal isn't a dead end. Missed handovers are counted in testing, not found by customers.
Should it also act in your systems? That is AI agent development, where tools start read-only. For all our AI work, see AI development.
The answer that looks right
So a convincing demo can't prove it is ready.
A customer asks how many days they have to return an order, and the assistant gives a number. If it isn't in your returns policy, it came from nowhere, in the same calm tone as every other reply. Nobody notices without the policy open beside the chat.
A demo only answers the questions somebody picked for it. Anthropic, the company behind the Claude models, says its techniques reduce invented answers without eliminating them, and that readiness should be measured by systematic evaluation. So we measure it, against criteria agreed with you before we build, on test questions that include ones it should refuse or hand over.
If one kind of answer keeps coming out wrong, the fix may be clearer instructions for the model (the prompt), better document search, or further training on your examples (fine-tuning). LLM development covers that choice, and generative AI development the wider build, including search by meaning (semantic search).
Decisions about the product
These are business decisions that no prompt tweaking can answer.
The DPDP duties above commence in May 2027 (dates and sections: DPDP Act compliance). EU AI Act explains when the EU rule applies.
How we work
Anthropic's prompt-engineering guidance puts it in this order: test cases, a first prompt, testing and refinement, final validation, then launch. We work the same way.
The questions customers already ask, plus the ones the assistant should decline or hand over.
You confirm the correct answers, the subjects it declines and where handovers go. These become the success criteria.
A support chat is several jobs in one (see the panel). We decide, prompt and test each separately.
We write the prompts and connect your documents, then grade against the test set, handovers included, until it meets the criteria.
Our security team, not the builders, verifies the guardrails (limits on what it may say), tests resistance to prompt injection (instructions, typed into the chat or hidden in what it reads, that try to override its rules) through LLM security testing, and checks replies for data leakage through generative AI security testing. See AI security.
For your technical team
We grade against Anthropic's Customer support agent guide and its page on defining success criteria and building evaluations:
Anthropic's rule is to "choose the fastest, most reliable, most scalable method". Code-based grading, such as an exact match ("output == golden_answer"), is fastest but lacks nuance. Human grading is "most flexible and high quality, but slow and expensive", and Anthropic's advice is "Avoid if possible." LLM-based grading is fast and flexible, once tested for reliability.
Anthropic's hallucination guide recommends explicit permission to admit uncertainty: "This simple technique can drastically reduce false information." Anthropic's support guide adds output checks that replies "align with your company's policies and known facts", and stripping personal data from replies unless explicitly required and authorized. The hallucination guide ends: "Always validate critical information, especially for high-stakes decisions."
Under Article 50(1) of the EU AI Act, providers must design and develop AI systems that interact directly with people so those people are told they are dealing with an AI, unless that is obvious to a reasonably well-informed person in context. India's rules on synthetically generated content are covered on generative AI security testing.
What no build can promise. No technique stops a language model inventing answers outright, so a refusal is checked by evaluation, not guaranteed.
FAQ
Refusals, handovers, readiness and personal data. For anything else, ask us directly.
Yes. The engineering that matters in a customer-facing assistant includes refusing to invent things such as refund policies or delivery dates. A language model's replies vary from one run to the next, so an assistant needs evaluation, and human review where the stakes justify it, before it is called ready. Grounding its answers in content the business already holds is covered on the RAG development page. Is it security-reviewed as well? Yes, by our security team rather than the people who built it, including guardrail verification and prompt-injection resilience testing. That work is described on the AI security page.
It can be built to decline, and the decline can be checked, but Anthropic does not claim any technique ends invention outright. Its Customer support agent guide lists a guardrail that reads "Avoid contractual commitments: Ensure the agent doesn't make promises or enter into agreements it's not authorized to make." Its hallucination guide recommends giving the model explicit permission to say it does not know. The same hallucination guide adds that its techniques cut hallucinations but "they don't eliminate them entirely". So a refusal is something an evaluation checks, not something a build can promise.
When the question needs a person, and deciding which questions those are is part of designing the product. Anthropic's Customer support agent guide names Escalation efficiency among its success criteria for a support assistant, which makes a missed handover something that is measured rather than noticed by a customer. Where the conversation arrives once it leaves the assistant is the business's decision, and it belongs in the design from the start alongside the subjects the assistant declines.
An evaluation against success criteria, not a convincing demo. Anthropic's Customer support agent guide puts it this way: "To determine the readiness of your solution, evaluate the chatbot performance using a systematic process combining quantitative and qualitative methods." Anthropic's prompt engineering overview opens with what it expects to exist first. "This guide assumes that you have:" 1. "A clear definition of the success criteria for your use case" 2. "Some ways to empirically test against those criteria" 3. "A first draft prompt you want to improve". It then adds: "If not, spend time establishing that first." The prompt is last on that list. 'It seemed fine in testing' is not a release criterion.
They can, and they give realistic questions, but a set built from real chats is a copy of those chats and can carry the same personal data. It then falls under the same retention and erasure decision as the transcripts themselves, and that decision is best made before the set is built. Under India's DPDP Act, 2023, the erasure and retention duties commence in tranche (c), eighteen months after publication of the Rules, which the DPDP Act compliance page dates to May 2027, and that page covers how they will work.
Decide before launch what it does with a user who may be under eighteen. Under section 2(f) of the DPDP Act, 2023, which is already in force, a child "means an individual who has not completed the age of eighteen years". Section 9(1) will call for verifiable consent of a parent before a child's personal data is processed, section 9(4) allows prescribed classes of Data Fiduciary and prescribed purposes to be exempted, and Rule 12 of the DPDP Rules sets those out with conditions, all of which commence in tranche (c), eighteen months after publication of the Rules, which the DPDP Act compliance page dates to May 2027.
Let's talk
Send the questions it has to answer and the ones it must decline. We deliver project work remotely from India, during business hours. We reply within one working day.