Home All services
Start a project → Call Now

AI chatbot development company in India

A support assistant that answers from your real policies — and says so when it doesn't know.

We build customer support chatbots that answer from content you already hold, decline to invent things like a refund policy or a delivery date, and hand over to a person when needed. We test that on your real questions before launch.

  • Tested on your real questions
  • Hands over to a person
  • Security review by a separate team
  • Anthropic customer support agent guide
  • Anthropic prompt engineering overview
  • DPDP Act, 2023
  • EU AI Act, Article 50(1)
Illustration: a large chat window with blank message bubbles, a small glowing AI core inside the reply bubble, and a stack of source documents behind it joined by a thin line of light, showing answers grounded in real documents

In brief

What it is
A chat assistant for your website or product that answers customers from your own policies and product information.
Why it matters
If the assistant invents a return window, nothing in the chat marks it as made up, and your customer quotes it back to you.
What you get
An assistant built to decline and hand over, graded on your real questions before launch, and reviewed by a security team that didn't build it.

What we build

Three things a customer-facing assistant must get right.

Answers from your own content

We agree which documents are the source of truth, and it answers from those, given to it directly or, where needed, looked up per question. RAG development covers how that look-up is built and tested.

A refusal instead of a guess

It has explicit permission to say it doesn't know. Output checks compare its replies with your policies and look for promises it isn't authorized to make.

A handover to a person

When a question needs a person, the conversation moves to a place you choose, so a refusal isn't a dead end. Missed handovers are counted in testing, not found by customers.

Should it also act in your systems? That is AI agent development, where tools start read-only. For all our AI work, see AI development.

The answer that looks right

A made-up answer sounds exactly like a real one.

So a convincing demo can't prove it is ready.

A customer asks how many days they have to return an order, and the assistant gives a number. If it isn't in your returns policy, it came from nowhere, in the same calm tone as every other reply. Nobody notices without the policy open beside the chat.

A demo only answers the questions somebody picked for it. Anthropic, the company behind the Claude models, says its techniques reduce invented answers without eliminating them, and that readiness should be measured by systematic evaluation. So we measure it, against criteria agreed with you before we build, on test questions that include ones it should refuse or hand over.

If one kind of answer keeps coming out wrong, the fix may be clearer instructions for the model (the prompt), better document search, or further training on your examples (fine-tuning). LLM development covers that choice, and generative AI development the wider build, including search by meaning (semantic search).

Decisions about the product

You set the assistant's limits before we build it.

These are business decisions that no prompt tweaking can answer.

  • The questions people really ask. From your support inbox, contact form and call notes, not invented at a whiteboard.
  • The subjects it declines. Exactly where its topic ends, in writing. Anthropic's task "Stay on topic" needs that line.
  • The correct answers, in writing. Return windows, delivery promises, warranty terms. Grading needs something to compare against.
  • Where a handover goes. A fixed place in the product where the assistant can send a conversation.
  • Whether it takes personal data at all. India's Digital Personal Data Protection (DPDP) Act, 2023 will limit consent to data "necessary for such specified purpose" (section 6(1)), subject to exemptions.
  • Users who may be under eighteen. The DPDP Act counts them as children. The FAQ covers parental consent.
  • Whether real chats become test questions. A set built from them copies any personal data they hold, so decide first how long to keep it.
  • Whether any users are in the EU. If so, the EU AI Act duty to tell people they're talking to an AI becomes part of the design.

The DPDP duties above commence in May 2027 (dates and sections: DPDP Act compliance). EU AI Act explains when the EU rule applies.

How we work

We write the tests before the model's instructions.

Anthropic's prompt-engineering guidance puts it in this order: test cases, a first prompt, testing and refinement, final validation, then launch. We work the same way.

  1. Collect real questions

    The questions customers already ask, plus the ones the assistant should decline or hand over.

  2. Write down answers and limits

    You confirm the correct answers, the subjects it declines and where handovers go. These become the success criteria.

  3. Split the conversation into tasks

    A support chat is several jobs in one (see the panel). We decide, prompt and test each separately.

  4. Build and grade

    We write the prompts and connect your documents, then grade against the test set, handovers included, until it meets the criteria.

  5. Security review, then launch

    Our security team, not the builders, verifies the guardrails (limits on what it may say), tests resistance to prompt injection (instructions, typed into the chat or hidden in what it reads, that try to override its rules) through LLM security testing, and checks replies for data leakage through generative AI security testing. See AI security.

For your technical team

How we grade the assistant and make refusals more likely.

What gets graded, and how?

We grade against Anthropic's Customer support agent guide and its page on defining success criteria and building evaluations:

  • Response accuracy: correct company and product information, "based on the information provided to Claude in context".
  • Difficult input: poor, harmful or irrelevant messages, and ambiguous cases even people would disagree on.
  • Tone, consistency and context: angry customers, complex issues, compliments that are really complaints, typos, rambling questions, irrelevant detail, abrupt topic shifts and facts from much earlier.
  • Escalation efficiency: correctly escalated conversations, against those that needed a person and didn't get one.

Anthropic's rule is to "choose the fastest, most reliable, most scalable method". Code-based grading, such as an exact match ("output == golden_answer"), is fastest but lacks nuance. Human grading is "most flexible and high quality, but slow and expensive", and Anthropic's advice is "Avoid if possible." LLM-based grading is fast and flexible, once tested for reliability.

What makes a refusal more likely?

Anthropic's hallucination guide recommends explicit permission to admit uncertainty: "This simple technique can drastically reduce false information." Anthropic's support guide adds output checks that replies "align with your company's policies and known facts", and stripping personal data from replies unless explicitly required and authorized. The hallucination guide ends: "Always validate critical information, especially for high-stakes decisions."

EU users and AI-generated content

Under Article 50(1) of the EU AI Act, providers must design and develop AI systems that interact directly with people so those people are told they are dealing with an AI, unless that is obvious to a reasonably well-informed person in context. India's rules on synthetically generated content are covered on generative AI security testing.

What no build can promise. No technique stops a language model inventing answers outright, so a refusal is checked by evaluation, not guaranteed.

FAQ

Questions before an assistant answers a customer.

Refusals, handovers, readiness and personal data. For anything else, ask us directly.

Yes. The engineering that matters in a customer-facing assistant includes refusing to invent things such as refund policies or delivery dates. A language model's replies vary from one run to the next, so an assistant needs evaluation, and human review where the stakes justify it, before it is called ready. Grounding its answers in content the business already holds is covered on the RAG development page. Is it security-reviewed as well? Yes, by our security team rather than the people who built it, including guardrail verification and prompt-injection resilience testing. That work is described on the AI security page.

It can be built to decline, and the decline can be checked, but Anthropic does not claim any technique ends invention outright. Its Customer support agent guide lists a guardrail that reads "Avoid contractual commitments: Ensure the agent doesn't make promises or enter into agreements it's not authorized to make." Its hallucination guide recommends giving the model explicit permission to say it does not know. The same hallucination guide adds that its techniques cut hallucinations but "they don't eliminate them entirely". So a refusal is something an evaluation checks, not something a build can promise.

When the question needs a person, and deciding which questions those are is part of designing the product. Anthropic's Customer support agent guide names Escalation efficiency among its success criteria for a support assistant, which makes a missed handover something that is measured rather than noticed by a customer. Where the conversation arrives once it leaves the assistant is the business's decision, and it belongs in the design from the start alongside the subjects the assistant declines.

An evaluation against success criteria, not a convincing demo. Anthropic's Customer support agent guide puts it this way: "To determine the readiness of your solution, evaluate the chatbot performance using a systematic process combining quantitative and qualitative methods." Anthropic's prompt engineering overview opens with what it expects to exist first. "This guide assumes that you have:" 1. "A clear definition of the success criteria for your use case" 2. "Some ways to empirically test against those criteria" 3. "A first draft prompt you want to improve". It then adds: "If not, spend time establishing that first." The prompt is last on that list. 'It seemed fine in testing' is not a release criterion.

They can, and they give realistic questions, but a set built from real chats is a copy of those chats and can carry the same personal data. It then falls under the same retention and erasure decision as the transcripts themselves, and that decision is best made before the set is built. Under India's DPDP Act, 2023, the erasure and retention duties commence in tranche (c), eighteen months after publication of the Rules, which the DPDP Act compliance page dates to May 2027, and that page covers how they will work.

Decide before launch what it does with a user who may be under eighteen. Under section 2(f) of the DPDP Act, 2023, which is already in force, a child "means an individual who has not completed the age of eighteen years". Section 9(1) will call for verifiable consent of a parent before a child's personal data is processed, section 9(4) allows prescribed classes of Data Fiduciary and prescribed purposes to be exempted, and Rule 12 of the DPDP Rules sets those out with conditions, all of which commence in tranche (c), eighteen months after publication of the Rules, which the DPDP Act compliance page dates to May 2027.

Let's talk

Tell us what the assistant must never promise.

Send the questions it has to answer and the ones it must decline. We deliver project work remotely from India, during business hours. We reply within one working day.