Home All services
Start a project → Call Now

Generative AI security testing in India

Check that your AI's answers are safe — wherever they land.

One answer from a model is harmless if only stored, runs as code in a web page, and can leak a conversation from a chat window. We test where your AI's answers land, whether they are right, and whether your guardrails hold there.

  • Scoped by where answers land
  • Every finding shows what the output did
  • Retest included
  • OWASP LLM10:2026
  • OWASP LLM07:2026
  • MITRE ATLAS AML.T0077
  • AITG-APP-05 and AITG-APP-11
  • MeitY rule 3(3)(b)
Illustration: a chat bubble, a picture and a text tile travel along a test track through three glass checkpoint gates; the picture is held under an amber light and the text tile has passed with a blue check mark

In brief

What it is
Security testing of what your generative AI sends out, and of what reads it.
Why it matters
A model can be talked into writing almost anything. The harm lands on whoever receives it: a system that runs it, or a person who trusts it.
What you get
A signed list of where your AI's output lands, findings with evidence, a CVSS v4.0 score and fix for each, and a retest.

What we test

Three questions about every answer your AI gives.

What your systems do with it

Every place that receives an answer, from a web page to a build step, is tested as untrusted input. OWASP's Top 10 for LLM Applications calls this LLM10:2026 Improper Output Handling.

What retrieval hands back

When the model searches your documents (retrieval, or RAG), can it return files the signed-in user couldn't open? We check across privilege levels and tenants. That is LLM02:2026 Sensitive Information Disclosure.

Whether the answer is right

Scored against your written definition of correct. Does it refuse a question built on a false premise? Are its citations, links and package names real? That is LLM07:2026 Misinformation.

Not sure which AI systems to test first? Start with an AI security assessment.

One answer, many readers

The same answer is a different risk in each place that reads it.

OWASP's fix for LLM10:2026 is encoding chosen "based on where the LLM output will be used", so we list every place first, as a deliverable, and treat the model "as any other user". Each finding shows the prompt, the raw answer, and what that place did with it.

Typical rows. ATLAS: MITRE's catalogue of attacks on AI systems. CWE: MITRE's list of software weakness types.
Where it landsWhat reads itWhat it needsWhat can go wrong
A JSON field a service only storesNothingNothingNo finding: the safe baseline
A web page in the browserThe HTML parserHTML or JavaScript encoding for the contextCross-site scripting
A chat window that renders MarkdownImages, link previews or iframes fetched with no clickNo fetching of web addresses the model wroteData leaving in the image's web address (LLM10:2026 risk 7, ATLAS AML.T0077 LLM Response Rendering)
A terminal, log viewer or IDE paneANSI and OSC control codesControl characters strippedFaked on-screen text, a hijacked clipboard (OSC 52), or a known terminal emulator bug
A query, shell command or eval callThe interpreterParameters, never a composed stringYour systems run injected commands or code (CWE-77, CWE-94)
An email or notification templateThe mail client, and any automation reading itEncoding, and a rule for links the model wroteHTML injection, or a trusted-looking link from your domain
A build step that installs what the model namedThe package managerEach name checked against the real registrySomeone registers the made-up name and your build installs their code (LLM04:2026 Supply Chain)
Another agentThe next model, as instructionsChecks where it is receivedInstructions followed unchecked; its permissions belong to the agent tool-permission review

Findings usually end in a familiar flaw, fixed through a web application vulnerability assessment, an API vulnerability assessment, or secure code review where the text was built.

This is a bounded engagement. Monitoring set-up, incident response and ongoing retainers are separate engagements (see managed security services). We publish no public benchmark scores, since they say little about your deployment, and never describe a named commercial model as evaluated. A test from us is not an audit, and we don't give legal advice.

How an engagement runs

Six steps, starting with what "correct" means.

  1. Write down what "correct" means

    You name owners for the model and application sides and write down what the system should and must never say, which OWASP's AI Testing Guide (AITG-APP-05) requires. If none exists, we produce it.

  2. List every place the output lands

    We walk your code and deployment, not just the API. If no list exists, we build it first.

  3. Reach the model through each input

    We place a test instruction in each input you name (prompt box, uploads, tool results, media), only far enough to reach the model. We don't assume access to the system prompt or retrieval corpus. LLM security testing covers inputs in depth.

  4. Test each place as untrusted input

    We get the model to emit Markdown images, link previews, iframes and control codes, then record what each client does without a click and whether your guardrails catch it.

  5. Check the answer and what retrieval gave back

    We score answers against your definition, test false premises, check that every citation, package, endpoint or command name is real, and see what retrieval returned to each user.

  6. Test media labels, where they apply

    If you publish generated media in India as an intermediary, we try to strip its label and embedded provenance through every export path.

Technical detail

The published sources behind each test.

Where OWASP draws the line between input and output

LLM10:2026 covers unsafe use of model output before it is passed downstream, LLM07:2026 output that is incorrect or misleading, and LLM01:2026 Prompt Injection the validation of model inputs. LLM03:2026 adds that sanitizing inputs and outputs is not a root control for excessive agency. LLM01:2026 also notes that a model makes no architectural distinction between instructions and data, so there is no clean equivalent to parameterized queries. Splitting the testing into input and output is our choice; OWASP's AI Testing Guide groups tests by capability. Jailbreaks (tricks that talk a model out of its rules) are input-side work.

NIST AI 100-2 E2025, of March 2025, adds that an indirect prompt injection is mounted by a third party and in many cases harms the primary user, so user vigilance is not a control. NIST files this under INTEGRITY VIOLATION [NISTAML.02].

Which CWE and AI Testing Guide identifiers a finding carries

CWE-1427 (improper neutralization of input used for LLM prompting) is marked Vulnerability Mapping ALLOWED. CWE-1426 (improper validation of generative AI output) is marked DISCOURAGED, citing "Potential Major Changes, Frequent Misinterpretation", with CWE-77, CWE-94 and CWE-116 suggested instead, yet OWASP's 2026 appendix makes it the primary mapping for LLM07:2026 and LLM10:2026. Both land on CWE-116, so each finding states its convention. MITRE's research-gap note on CWE-1426 says many prompt-injection flaws differ from ordinary injection only in that the input came from model output.

OWASP's AI Testing Guide v1, published 26 November 2025, has fourteen application tests; its output group is AITG-APP-05, AITG-APP-10, AITG-APP-11 and AITG-APP-12. We use AITG-APP-05 Testing for Unsafe Outputs (harm to the reader, or to the consuming system) and AITG-APP-11 Testing for Hallucinations (factuality under prompt framing, debunking under user pressure). AITG-APP-10 and AITG-APP-12, named by identifier only because two OWASP files title the first differently, belong with AI governance. Downstream, the guide points to the OWASP Top Ten, the API Top Ten and the Web Security Testing Guide.

Retrieval, citations and made-up names

LLM09:2026 Vector and Embedding Weaknesses owns the retrieval mechanism, and LLM02:2026 Sensitive Information Disclosure the regulatory consequence. LLM02:2026 states that cosine similarity does not respect access control lists. LLM09:2026 adds that similarity search frequently runs across the full index before access control applies, and that a backup holding only embeddings is not a safe-harbor classification. That mechanism and the index's write path belong to the RAG access-control checks in LLM security testing.

ATLAS files a wrong answer dressed to look sourced under AML.T0067, with Citations as a sub-technique, and separates discovering a hallucinated name (AML.T0062) from registering it (AML.T0060).

Labels on generated media under India's IT Rules

Under India's IT Rules, as served by MeitY updated as on 10.02.2026, rule 3(3)(a)(ii) asks for a prominently visible label, or a prominently prefixed audio disclosure, plus embedded "permanent metadata or other appropriate technical provenance mechanisms", including a unique identifier, to the extent technically feasible. Rule 3(3)(b) forbids the intermediary from enabling the modification, suppression or removal of either, and MeitY's FAQ rules out "remove watermark" and "export without metadata" features. We try export, re-encoding, cropping, screenshots, integrations and the API against both.

The duty binds an intermediary offering a computer resource that enables or facilitates synthetic generation; if you are not one, the report names the rule and the step leaves scope. Text-only output is not synthetically generated information under rule 2(1)(wa), per MeitY's FAQ. Where rule 4(1A) applies, we test the automated verifier, not the declaration. The amendment is G.S.R. 120(E) of 10 February 2026, in force 20 February 2026 per that FAQ; the ten percent surface-area test was in the October 2025 draft only. Whether a label was applied at all is a conformity question, like the EU AI Act's marking duty: see EU AI Act compliance.

FAQ

Questions about generative AI security testing.

For anything else, ask us directly.

Testing what a generative AI system sends out, and what your own systems do with it. We test every place a model's answer lands as untrusted input, and score answers against your own definition of correct: OWASP's LLM10:2026 Improper Output Handling and LLM07:2026 Misinformation.

By where each one stops. LLM security testing works through the OWASP Top 10 for LLM Applications entry by entry, and its agent tool-permission review covers LLM03:2026 Excessive Agency: what an agent may call, and under which credential. This engagement goes deep on the systems that receive the answer, which decide whether a harmful answer that gets through does real damage or just sits in a chat log.

Only partly. A key returns the response text, not the browser, chat window or build step that acts on it. OWASP lists client renderers among the aggravating conditions for LLM10:2026, and MITRE ATLAS (release v2026.08) records the same event as AML.T0077.

Six things: your statement of what the system should and must never say; a list of systems that consume model output, or agreement that we build it; access to the rendering clients; an evaluation set or other source of truth, including false premises; a non-production environment or written authorization for destinations a test can change (shell, SQL, outbound email, deployment), test accounts at each privilege level and two tenants if multi-tenant; and owners for the model and application sides. Also tell us if you publish generated media as an intermediary.

Under OWASP's 2026 edition, yes, judged by consequence rather than truth. LLM07:2026 Misinformation covers output that is incorrect, incomplete, unsupported or misleading, and credible enough to influence a human decision, an automated workflow or an agent action. Against your own definition of correct, a system that won't refuse a false premise gets a finding.

Not by itself. On Amazon Bedrock, the prompt-attack filter covers input fields only; prompt-leakage detection is on the Standard tier only; for some API operations, prompt attacks are not filtered unless the input is tagged; and the filter does not evaluate tool results or tool definitions. Contextual grounding and Automated Reasoning checks evaluate responses, and content and sensitive-information filters run on both sides. We test whether yours hold where your code renders output.

It is a 2025 citation. In the 2026 edition, seven of the ten entries move and one is renamed: Improper Output Handling is LLM05:2025 and LLM10:2026, Misinformation is LLM09:2025 and LLM07:2026, and System Prompt Leakage became LLM08:2026 Hidden Context Exposure. OWASP's per-risk web pages still use the 2025 numbering. Tenth place is not a demotion: the community vote carries three-quarters of the weight and incident data a quarter, so the order is not a severity ranking.

A destination list we build and sign, naming what reads the output at each place and which team owns it. Findings in two lists, wrong answers and unsafe handling, each with evidence, its ordinary weakness and a CVSS v4.0 score. You sign the destinations in scope, and your owners sign off their fixes. Media label checks where in scope. Retest included.

Let's talk

Tell us where your AI's answers go.

List the clients that show a response, the services that use one and anything that runs what the model names. Project work is delivered remotely from India during business hours. We reply within one working day.