Home All services
Start a project → Call Now

AI integration services in India

Connect an AI model to software you already run — ready for its retirement date and its bill.

We build it into your website, app or internal software, then plan for what its provider controls: retirement dates, usage limits and cost.

  • Uses your existing sign-in and audit trail
  • Every platform rule traced to its publisher
  • Reviewed by a separate security team
  • Anthropic model deprecations
  • OpenAI deprecations
  • Microsoft model migration guide
  • RBI IT Outsourcing Directions, 2023
  • DPDP Act, 2023, s.16(2)
Illustration: an AI core in the centre with plug-like connectors snapping into a database, a browser window, a phone and a stack of cards, each connection passing through a small glass checkpoint

In brief

What it is
Connecting a hosted AI model (one another company runs, such as OpenAI's or Anthropic's) to software with existing users, records and permissions.
Why it matters
That company decides when the model is switched off and how much you can use. Nothing in your code warns you when either changes.
What you get
An integration ready for retirement day and the errors a provider can return, reviewed by a security team that didn't build it.

What the integration covers

Six things we settle before the model goes live.

Your team holds some of the answers, so we ask first.

The route, sign-in and audit trail

We map where your software calls the model: directly, via a cloud platform or your own gateway. Requests use your existing sign-in and audit log, and no timeout cuts off a reply midway.

A plan for retirement day

We agree who gets the provider's retirement notices. Where allowed, we lock the model version and keep tests comparing it with its replacement. A locked version fails visibly on retirement day instead of quietly changing.

Handling for each failure

Your software tells a rejected request, a usage or spend limit and a retired model apart, and handles each differently. A streamed reply can fail after reporting success, so errors are read from the stream.

Retries that don't multiply

Providers' code libraries (SDKs) already retry, so we keep one retry layer, not two. At a spend cap, retrying can't help: on Anthropic, requests fail (a 429 with no retry-after header), SDK retries included, until access resumes.

Usage and cost you can see

Providers usually bill by the token, a small chunk of text. We record tokens and response time for every request, so you see what each costs and can tell a slow reply from a long one.

A fallback the code can reach

If your plan names a second provider or in-house work, your software must be able to switch to it. Banks, NBFCs and others bound by the Reserve Bank of India's 2023 IT outsourcing Directions must identify such an alternative. Where your data may be processed is settled first.

A dependency that moves

A model someone else runs can change or disappear.

Software you install stays at the version you tested. A hosted model is whatever its provider serves that day, until it is retired.

A passing build makes a model look settled. Microsoft's model migration guide disagrees: "Every model you run in production has a retirement date." A replacement can quietly shift tone, formatting, JSON shape, latency, cost or tool calls, and break code downstream. Anthropic adds that deprecated models are likely to be less reliable.

Where this page fits

We scope integrations separately. The wider system sits under AI development, model choice under LLM development, how much to automate under AI automation, and AI inside an ERP or CRM under AI automation and integration. Our security team, not the builders, reviews the result: see AI security.

Retirement day, by platform

What happens to your calls when a model retires.

The platform and deployment type decide whether calls fail or quietly move to a newer model your team may not have tested.

Our plain-English reading of each publisher's pages, without model IDs or retirement dates.
Platform and deployment typeOn retirement dayNotice statedAvailability and extension
Anthropic, on platforms it operatesRequests to retired models fail.At least 60 days for publicly released models, to customers with active deployments.No extension route named.
Amazon Bedrock, launched on or after September 7, 2026Removed from all AWS Regions: requests fail without a private arrangement with the provider. No automatic migration.A Legacy period first: 6 months for most models, or 45 days.Only through that arrangement.
Amazon Bedrock, launched before September 7, 2026Requests fail on or soon after the end-of-life (EOL) date, barring a private arrangement. No automatic migration.At least 6 months as Legacy before EOL.At least 12 months on Bedrock before EOL.
Microsoft Foundry, Standard, Global Standard and Data Zone StandardAuto-upgraded, region by region. Set to NoAutoUpgrade, the deployment stops working.Generally available (GA) models: at least 60 days. Preview: at least 30.A GA model's date is set at launch, 18 months out, with no separate announcement. Not extendable.
Microsoft Foundry, ProvisionedNo auto-upgrade: you migrate by hand. After retirement, every request returns 410 Gone.No separate figure.Not extendable.
OpenAIDeprecated once announced, then inaccessible from its shutdown date.GA models: at least 6 months. Specialized variants: 3. Preview models: much shorter, such as 2 weeks. Less if safety or compliance requires.Sometimes, dedicated capacity after shutdown.
Google CloudThe model versions page we read gives availability periods only.Short-term models retire 45 days after a replacement is released.Others: at least 12 months after release. Dates may be extended, never brought forward.
Anthropic models on Microsoft Foundry, unresolvedAnthropic's dates cover the Claude API, Claude Platform on AWS and Microsoft Foundry. Bedrock and Google Cloud set their own.Depends on whose schedule governs.Microsoft gives GA Anthropic, DeepSeek, Fireworks and Mistral AI models a 12-month lifecycle, not 18. The two pages disagree; we don't pick.

When the model changes

Test the replacement before a deadline forces it.

These steps follow Microsoft's model migration guide.

  1. Start logging now

    Log prompts, responses, latency and token counts now, before any replacement is named. Capture is opt-in and never retroactive: you can't evaluate traffic you didn't record.

  2. Freeze the test

    Fix the inputs, correct answers and success criteria for the whole migration, or the results stop being comparable.

  3. Replay it unchanged

    Run your current prompts and settings on the new model, to separate what the model changed from what you changed.

  4. Measure quality, latency and cost

    Score answer quality, how fast replies start and finish, and the cost per request, with tests your app team and reviewers both trust.

  5. Compare offline, keep the old route

    Regulated flows, such as health records or financial transactions, often can't expose customers to a new model, so compare outputs offline. Keep the old deployment reachable until you're confident.

For your engineers

Usage limits, backup routes and data rules, in detail.

Retries and cross-Region routing are in the FAQ.

Why your usage limit and your bill differ (max_tokens)

Microsoft reserves compute for the full max_tokens until a response finishes. Anthropic doesn't count it toward its output-token rate limit. Amazon Bedrock deducts it from quota up front: in one AWS example, a request takes 36,000 quota tokens and settles at 9,000, so fewer can run at once. In another, 1,500 leave the quota and 1,100 are billed.

On the bedrock-runtime endpoint, output counts 5x, 10x or 15x against quota for versions of one named model family (1:1 for others), so a move from a 5x to a 15x version triples how fast output uses quota. Models only on the bedrock-mantle endpoint have separate input and output quotas, so burndown doesn't apply.

Does a service level agreement (SLA) count throttling?

Amazon Bedrock's SLA (last updated October 4, 2023) counts as an error only a request that returns a 500. Throttling returns a 429, so it never counts, though none of the six listed exclusions names it. AWS promises "commercially reasonable efforts", and its credit schedule begins below 99.9%.

Why to switch traffic on gradually

A sharp rise in usage can hit Anthropic's acceleration limits (429 errors), so ramp up gradually. OpenAI's slow_down error can occur within your per-minute limits: it reflects how fast traffic grew, and endpoints that once returned a 503 for it now return a 429.

Why a backup route can't sit idle

On Amazon Bedrock, once a model's Legacy period begins, new customers can't adopt it and existing ones may lose access after 15 days of inactivity. On Microsoft Foundry, a new subscription under the same tenant doesn't inherit access, and there's no automatic failover or disaster recovery: routing between deployments is yours to build.

Where Microsoft Foundry processes your data

Microsoft's deployment-types page says Data Zone deployments process data only within a Microsoft-specified zone (US, EU or Asia Pacific), and Microsoft can add regions to a zone without prior notice. Geography-based types arrive last, with no guaranteed date, depending on capacity freed as older models retire. Microsoft's high-availability guide says to default to US or EU Data Zones; the two pages don't list the same zones.

India: the DPDP Act and RBI outsourcing Directions

Section 16(2) of the DPDP Act, 2023 says nothing in section 16 limits any Indian law that gives a higher degree of protection for, or restricts, the transfer of personal data outside India. It commences in tranche (c), which our DPDP page dates to May 2027, subject to the Act's exemptions.

The Reserve Bank of India (Outsourcing of Information Technology Services) Directions, 2023 (RBI/2023-24/102, April 10, 2023, effective October 1, 2023) bind the entities in clause 2(a), including commercial banks as defined there, Tier 3 and 4 urban co-operative banks, Top, Upper and Middle Layer NBFCs, credit information companies, EXIM Bank, NABARD, NaBFID, NHB and SIDBI. Clause 2(c) limits them to Material Outsourcing of IT Services, defined in clause 3(a)(ii). Paragraph 18's contingency plan must consider alternative providers or bringing the work back in-house; paragraph 22's exit strategy must identify alternative arrangements.

What this work doesn't decide. Notice terms in a supplier contract belong to third-party AI risk management. Whether the RBI Directions apply is for the regulated entity and its counsel to decide, using the Directions' Appendix III.

FAQ

Questions about a model you don't control.

For anything else, ask us directly.

Yes. We build AI into websites, apps and internal software already running, and the model's requests go through the sign-in and audit trail that software keeps. The model's retirement date and usage limits are set by its provider, not your team.

The platform and deployment type decide, not the model. On Anthropic's platforms, and on Amazon Bedrock without a private arrangement, requests to a retired model fail. Microsoft Foundry's Standard deployments can instead be upgraded, so calls keep succeeding against a model your team may never have tested.

Because the SDK and the API check different things. Anthropic says most SDKs keep deprecated parameters in their request types, so old code still type-checks. But on the models its table names, a non-default temperature, top_p or top_k returns a 400 error. Its Python SDK, v1.0 and later, removes them, so passing one raises a TypeError.

Check what the SDK already does. Anthropic's official SDKs retry transient failures twice by default (each client can change or disable this), and OpenAI's retry eligible 429 and 503 responses. A second loop multiplies requests to a struggling endpoint. OpenAI advises disabling SDK retries or counting them in your limits, following Retry-After or backing off with jitter, capping attempts and total time, and never retrying quota or billing errors.

Inside a geography, not necessarily your Region. On Amazon Bedrock, requests to a geography's inference profile (US, EU or APAC) stay there, and data is stored only in the source Region by default. But prompts and results may be processed, and stored for abuse detection, in other Regions, even ones you never enabled. CloudTrail's additionalEventData.inferenceRegion field shows where each request went. Check a one-country residency rule against that.

We can build the first version alongside your team, then hand it over, as our AI development page describes. Retirement dates, notice periods and usage limits stay with the model's provider, and this page sets them out.

Sources where publishers disagree
  1. Anthropic, model deprecations
  2. Microsoft, model retirements
  3. Microsoft, deployment types
  4. Microsoft, high availability

Let's talk

Tell us what the model will sit inside.

Name the system it must fit into and where your data may live. Project work is delivered remotely from India during business hours. We reply within one working day.