Prompt injection & jailbreaks
Instructions override intended policy or system behaviour.
Uncover exploitable model, agent, RAG, tool-use, privacy and governance weaknesses before they become production incidents. DataConsultant designs authorised, threat-led tests around the complete AI system so engineering, security and risk teams can see what failed, why it matters, how to reproduce it and what to fix next.
Testing is performed only within an agreed, authorised scope. Timeline, environments, access, evidence handling and commercial terms are confirmed after scoping.
AI applications create attack paths that conventional application testing can miss because instructions, retrieved content, model behaviour and connected tools can influence one another. The objective is to test realistic misuse paths across the system boundary, not to prove that a model is universally secure.
Instructions override intended policy or system behaviour.
Agents invoke actions beyond the intended business boundary.
Sensitive content escapes through prompts, retrieval, logs or outputs.
Untrusted content changes context, answers or downstream actions.
Weak identity or authorisation boundaries expose higher-risk capabilities.
Attackers steer outputs toward unsafe, misleading or prohibited behaviour.
Abuse succeeds without sufficient detection, evidence or escalation.
Downstream systems treat model output as trusted instructions or data.
Internal instructions or policy logic become visible to unauthorised users.
Models, plugins, libraries or external services expand the trusted boundary.
Controls that work on simple prompts fail in multi-turn or chained attacks.
Release and risk decisions lack reproducible tests, ownership and retest proof.
Move from ad-hoc prompt experiments and unknown attack surface to a structured adversarial-testing programme with traceable scenarios, reproducible findings and remediation evidence.
Start with the system boundary, attack surface, data flows, user roles, agent permissions and the business consequences that matter most.
End-to-end testing can span the complete enterprise AI stack, from applications and agents to models, retrieval, tools, data, identity, guardrails and human approval points. Final coverage follows the authorised scope.
Prompts, sessions, outputs, memory and user journeys.
Ingestion, retrieval, source trust, permissions and leakage.
Goals, memory, tool chains, approvals and action boundaries.
Parameters, downstream trust, secrets, permissions and abuse paths.
Routing, prompts, policies, model switching and control logic.
Sensitive content, classification, access, provenance and tenancy.
Authentication, authorisation, policy enforcement and boundaries.
Validation, logging, escalation, approval and incident evidence.
A threat-led test plan combines attack classes that fit the system architecture and business risk. The matrix below is illustrative; not every engagement requires every test family.
| Prompt & Instruction Attacks | Data / RAG Attacks | Agent & Tool Abuse | Identity & Authorization | Model Behavior & Safety | Privacy & Confidentiality | Availability & Resource Abuse | Supply Chain & Integration | Governance & Evidence |
|---|---|---|---|---|---|---|---|---|
| Prompt injection Jailbreaks Instruction conflicts System-prompt leakage | Context poisoning Retrieval manipulation Malicious documents Cross-tenant leakage | Tool misuse Unauthorised actions Excessive agency Unsafe code generation | Privilege escalation Broken access control Identity spoofing Cross-user access | Harmful generation Bias and policy bypass Model manipulation Misleading outputs | Sensitive-data disclosure PII leakage Inference attacks Memory/log exposure | Denial of wallet Resource exhaustion Token-cost abuse Service degradation | Insecure plugins Third-party tool risk Dependency compromise Service disruption | Control bypass Lack of auditability Unowned findings Weak incident evidence |
The attack surface extends across trust boundaries. Threat scenarios are mapped to the components an attacker can influence and to the controls expected to prevent, detect, contain or evidence abuse.
Define scenarios around identities, retrieval, tools, agent actions, model behaviour and the enterprise systems that receive AI output.
Before execution, confirm whether the system can be tested safely and whether evidence is sufficient to interpret findings.
Illustrative readiness view only. Actual readiness findings are evidence-based and are not inferred from these example bars.
Prioritise tests by the business action or sensitive asset an attacker could influence, not by attack names alone.
| Business Use Case | Sensitive Asset | Threat / Abuse Goal | Attack Technique | Control to Validate |
|---|---|---|---|---|
| Customer-support copilot | Customer data and PII | Extract another user’s information | Prompt injection / data exfiltration | Retrieval filtering, tenancy, output controls |
| Internal knowledge assistant | Confidential documents | Manipulate retrieved context | RAG poisoning / indirect injection | Source trust, ingestion, permissions |
| Automated service agent | Business systems and actions | Trigger unauthorised task | Tool abuse / privilege escalation | Tool scoping, approvals, identity |
| Developer assistant | Code, secrets, repositories | Expose credentials or unsafe code | Context leakage / output misuse | Secrets isolation, output validation |
Adversarial testing needs explicit authority, safe operating boundaries and reproducible evidence. The delivery method is adapted to the system, but the control points below remain important across most enterprise engagements.
Authorise systems, users, actions and stop conditions.
Map assets, actors, abuse goals and trust boundaries.
Inventory prompts, RAG, identities, tools and dependencies.
Select scenarios, techniques and evidence criteria.
Run creative multi-step and context-aware attacks.
Execute reusable variants and regression cases.
Confirm reproducibility and realistic prerequisites.
Record inputs, outputs, logs, actions and limitations.
Link impact, likelihood and control strength.
Translate failures into control and design actions.
Verify fixes and preserve reusable test coverage.
A useful finding explains exploitability, business consequence, affected control, evidence and the condition that must change. Severity should support remediation decisions rather than replace judgement.
| Finding Type | Exploitability | Business Impact | Data Sensitivity | Detectability | Overall Severity |
|---|---|---|---|---|---|
| Agent tool misuseUnauthorised action path | High | High | Medium | Medium | Critical |
| RAG leakageCross-user content exposure | Medium | High | High | Low | High |
| Prompt injectionGuardrail bypass | High | Medium | Medium | Medium | High |
| Unsafe output handlingDownstream trust issue | Medium | Medium | Low | High | Medium |
| System-prompt disclosureInternal instructions exposed | Medium | Low | Low | High | Low |
Outputs are designed for both remediation teams and accountable decision-makers. Exact documents depend on scope, access, evidence and whether remediation validation is included.
Material attack paths, business implications, limitations and decision priorities.
System boundaries, trust relationships, data flows, identities, tools and dependencies.
Threat actors, abuse goals, attack hypotheses, assets and priority scenarios.
Test cases, techniques, system components, evidence rules and execution status.
Inputs, outputs, tool calls, logs, prerequisites, replay steps and limitations.
Risk rationale, business impact, affected controls, owners and remediation priority.
Technical, product, process, monitoring and governance actions linked to findings.
Preventive, detective and response controls mapped to observed attack classes.
Verification of agreed fixes, unresolved conditions and residual-risk observations.
Reusable scenarios to test material failure modes after model, prompt or workflow changes.
Connect exploit evidence to the control owner, fix, acceptance criteria and retest scenario needed to close the risk.
DataConsultant does not publish a fixed fee for adversarial testing. Engagement shape, timeline and price are confirmed after the AI system, authorised scope, risk, access, test depth and evidence requirements are understood.
One defined AI application or narrow decision scope with selected high-priority attack classes.
Request a QuoteBroader human-led attack exploration across prompts, RAG, tools, identities, models and workflows.
Request a QuoteSpecialised testing for retrieval trust, vector-store boundaries, autonomous actions and tool permissions.
Request a QuoteRe-execute agreed material findings and preserve attack cases for future model or workflow changes.
Request a QuotePeriodic adversarial evaluation aligned to material system changes, risk triggers and governance needs.
Request a QuoteMarket-guidance note: references reviewed 8 September 2026. They are shown to help buyers understand scope sensitivity, not as an official DataConsultant price, market average, benchmark or commitment. Supplier prices and inclusions can change. DataConsultant pricing is provided only after scoping.
Adversarial testing is most useful when the AI system, decision, authorised environment and expected evidence are sufficiently defined. A different service may be more appropriate when the need is primarily conventional cybersecurity, certification or general AI advisory.
Recognised references can help organise threats, evidence and governance discussions. They are used as relevant lenses rather than treated as automatic proof of compliance or complete security coverage.
Useful for current LLM application risk categories such as prompt injection, sensitive-information disclosure, supply-chain risk, data/model poisoning, improper output handling, excessive agency and other GenAI-specific weaknesses.
Review OWASP prompt-injection guidance ↗A living knowledge base of adversary tactics and techniques involving AI-enabled systems, including predictive, generative and agentic AI attack paths and mitigations.
Review MITRE ATLAS ↗NIST AI 600-1 provides a cross-sectoral profile for incorporating trustworthiness considerations into the design, development, use and evaluation of generative AI systems.
Review NIST AI 600-1 ↗An AI management-system standard that can inform governance, risk, accountability and continuous-improvement discussions. Adversarial testing does not by itself establish conformity or certification.
Review ISO/IEC 42001 ↗Share the AI system, business use case, data sensitivity, agent or RAG architecture, current controls and the decision your testing evidence needs to support.
Answers to common enterprise questions about systems in scope, red teaming, RAG and agents, production testing, evidence, remediation, pricing, duration and assurance limitations.
Share your contact details and a concise requirement. DataConsultant can review likely scope, test prerequisites, evidence needs and the appropriate engagement model.