Home All services
Start a project → Call Now

LLM development company in India

Pick the model that answers your users best — and fix what it gets wrong.

We build software on large language models (LLMs), test it on your users' real questions, and fix each failure with the lever it needs: a better prompt, retrieval from your documents, or a fine-tune.

  • Models compared on your own questions
  • Each fix chosen for a failure we measured
  • Fine-tune trade-offs shown up front
  • OpenAI, Optimizing LLM Accuracy
  • Google, Introduction to tuning
  • OpenAI deprecations
  • Google Data residency
  • OpenAI Data controls
Illustration: a glowing glass block of stacked layers, standing for a language model, on a test bench beside a gauge dial and a clipboard of check marks

In brief

What it is
Software built on a large language model, improved by a better prompt, retrieval from your documents, or a fine-tune (training it on examples).
Why it matters
Each lever fixes a different kind of failure, so the wrong one leaves the real problem in place. A fine-tune also ties you to the model it was trained from (its base model), and to that provider's retirement dates and data rules.
What you get
A test set built from your users' questions, a model chosen by its score on it, and each fix aimed at a failure the tests found. You own the code, prompts and test set.

The fix follows the failure

Start with a prompt, then let the tests pick the next step.

Many guides describe a ladder: prompt first, then retrieval, then a fine-tune. OpenAI's Optimizing LLM Accuracy guide rejects that order: the three levers solve different problems, so you pull the one each failure needs.

OpenAI and Google both start with a prompt. Google's Introduction to tuning says to fine-tune only if needed, and to find where the model makes mistakes before adding more data.

So nobody can say up front whether you'll need a fine-tune. The answer comes one failed test question at a time.

OpenAI says many of its own largest customer deployments used only prompts and retrieval, though it gives no count. If you've been told you need a fine-tune, see why we usually push back.

Which fix for which failure

Match each failure to the lever that fixes it.

OpenAI sorts failures into missing knowledge and inconsistent behavior, and one system can have both.

Our plain-English summary of OpenAI's Optimizing LLM Accuracy guide.
FailureWhat you seeThe fixWhat it leaves behind
Missing knowledgeFacts it never learned, out-of-date facts, or your private informationGive the model the right information: in the prompt, or at scale with retrieval (RAG)A search step you adjust as well as the model. It fixes missing knowledge only.
Inconsistent behaviorWrong format or tone, or reasoning it doesn't followFine-tune on examples of the prompt and the answer you expectTraining and validation sets, and a new run for every further round
Right context, wrong useIt has the facts and still gets it wrongBetter instructions, plus a fine-tune if examples helpPrompts often don't scale to wider problems or stricter consistency
Both at onceEarly tests show both kindsRetrieval plus a fine-tune: the two combine, and some systems need bothMore complexity and a risk of regressions. In one undated OpenAI example, with no validation-set size reported, retrieval added noise and lowered a fine-tuned model's score.

Retrieval is explained on generative AI development. RAG development covers how we check it finds the right passages, and how its search index is tied to the model that built it. The decision guide on resources compares all three.

No lever is set-and-forget. OpenAI's Model optimization guide says the same question can get different answers, and a model's behavior changes between versions. That applies to a prompt too, so rerun the test set whenever the model or prompt changes.

How we work

Test first, then change one thing at a time.

Every candidate model and every change is scored on your own questions. How many rounds it takes depends on what the tests find.

  1. Build a test set with answers

    Real questions from your users, each with an agreed right answer (also called an evaluation set). OpenAI's bar before heavier methods: 20+ questions, with each failure examined and a likely cause found.

  2. Choose the model on that set

    We start with a prompt and compare two or three models, such as Claude, GPT, Gemini or Llama, on your set, not a public leaderboard.

  3. Sort each failure

    Missing knowledge, inconsistent behavior, or both. The table above shows the lever for each.

  4. Change one thing, measure again

    OpenAI's loop: evaluate, form a hypothesis, apply it, evaluate again. What a test set for a customer-facing assistant must check is on AI chatbot development.

  5. Stop at the accuracy target

    OpenAI's advice: aim for your accuracy target, not for retrieval plus a fine-tune because they look sophisticated.

Before you choose a fine-tune

Know what a fine-tune commits you to.

Three things to weigh first, though a fine-tune can still be the right call. Ask what it's for: OpenAI gives two reasons, accuracy on a specific task or efficiency, each needing different evidence.

Tied to its base model

On OpenAI, a fine-tune serves only until its base model retires. Microsoft's platform retires fine-tunes in two phases: training closes first, by default no earlier than the base model's retirement, and existing fine-tunes stay deployable until the deployment phase ends. How retirements reach your code is on AI integration; who owns the weights is on AI development.

A dated wind-down on OpenAI

OpenAI's self-serve fine-tuning is closing in stages. Since May 7, 2026, teams that have never fine-tuned there can't start; from Jan 6, 2027, even active customers can't create new jobs. The FAQ quotes each stage.

Where it can be processed

Neither Google's nor OpenAI's published tables offer tuning in India, so we settle early whether training examples may leave the country. Indian law is covered on DPDP Act compliance.

Output formats on fine-tuned OpenAI models

OpenAI's Structured Outputs guide accepts fewer JSON Schema keywords on fine-tuned models. It drops minLength, maxLength, pattern and format for strings; minimum, maximum and multipleOf for numbers; patternProperties for objects; and minItems and maxItems for arrays. With strict: true, an unsupported schema returns an error.

Its function calling guide adds that strict mode is disabled when a fine-tuned model calls several functions in one turn, and that fine-tuned models' schemas are cached and not eligible for zero data retention. So test any output format you depend on against the fine-tuned model itself, not its base.

Where fine-tuning data is processed and kept

OpenAI processes fine-tuning jobs in the United States and Europe (EEA + Switzerland), keeps them until deleted and excludes them from zero data retention, as it does its hosted vector stores. So on OpenAI, retention follows the endpoint, not the lever.

We also check whether prompts and retrieved text may leave India, and whether you need a model you can host, a route in the stack on AI development.

How much training data, and what kind

Both publishers put quality first. OpenAI names training examples that differ subtly from production as one of the most common pitfalls. Google says the data should reflect the prompts, format and context the model meets in production.

On size, OpenAI suggests starting with 50+ examples and adding more only for behavior errors, not missing knowledge. Google suggests a sizable labeled dataset, about 100 examples or more.

FAQ

Questions about LLM development.

Models, test sets, fine-tuning and data residency. For anything else, ask us directly.

It depends on the task, the data residency rules you live under, and your latency budget. We benchmark two or three against your own evaluation set rather than a public leaderboard, because leaderboards do not know what your users ask. Whether a given failure then calls for a better prompt, retrieval or a fine-tune is a separate decision, taken one failed evaluation question at a time.

It depends on what an evaluation shows and on what the fine-tune is for. OpenAI's Optimizing LLM Accuracy guide separates missing knowledge from inconsistent behavior, and for the second, which it calls learned memory, it describes the fix as "showing examples of the prompt and the response you expect, and the model learning from those". The same guide says "Fine-tuning is typically performed for one of two reasons:" and names them as accuracy on a specific task and efficiency. Our own position on fine-tuning is among the arguments listed on the AI development page.

OpenAI's guide sets its own bar: "If we have a set of 20+ questions and answers, and we have looked into the details of the failures and have a hypothesis of why they’re occurring, then we’ve got the right baseline to take on more advanced optimization methods." The answers matter as much as the count. OpenAI's reason for starting with prompt engineering makes the same point: "This is because it forces you to define what accuracy means for your use case - you start at the most basic level by providing an input, so you need to be able to judge whether or not the output matches your expectations." A failure with no agreed right answer cannot be sorted at all.

On OpenAI's platform it stops serving. OpenAI's deprecations page says "Inference on fine-tuned models will be disabled only when the underlying base model is deprecated." The same page dates a wind-down of OpenAI's self-serve fine-tuning in three rows. May 7, 2026: "Creating fine-tuning jobs or training is not available to organizations that have not previously run fine-tuning." July 2, 2026: "Creating fine-tuning jobs is no longer available to organizations that have not run inference on a fine-tuned model in the past 60 days." Jan 6, 2027: "Active existing customers will no longer be able to create new fine-tuning jobs on this date." How platforms retire base models more generally is covered on the AI integration page.

Not on the two published tables read for this question. Google's Data residency page counts tuning as ML processing, and its table marks the India column Supported on some inference rows but on none of its three tuning rows, which are marked for US multi-region and EU multi-region only. OpenAI's Data controls page lists India with regional storage "Yes" and regional processing "No", and says "Support for regional storage does not imply support for regional processing." So on those tables, a request and a fine-tune can get different answers for India; settle whether training examples may leave India before choosing the lever. What Indian law makes of it is covered on the DPDP Act compliance page.

No, and OpenAI published a case where it did not. In OpenAI's own Icelandic error-correction case study, adding retrieval over 1,000 embedded examples to a fine-tuned model lowered its BLEU score (a text-match score) from 87 to 83. OpenAI had called that step a lower confidence optimization. It is one undated worked example on one task, so it shows that retrieval can lower a fine-tuned model's score, not that it will. The same OpenAI guide also says the two methods stack, and that some use cases need both.

Let's talk

Tell us what the model is getting wrong.

Describe what you're building and where its answers fall short. We scope after seeing your data, not before, and deliver remotely from India during business hours. We reply within one working day.