See How Fidelis Deception® Turns Attacker Activity Into Actionable Evidence

What is Prompt Injection?

Prompt Injection Defined

Prompt injection is a security attack in which an attacker manipulates the instructions given to an artificial intelligence (AI) system, particularly a large language model (LLM), to influence its behavior. The goal may be to make the AI ignore its original instructions, reveal sensitive information, perform unintended actions, or generate content outside its intended boundaries.

What Is Prompt Injection?

Prompt injection occurs when malicious or carefully crafted input causes an AI model to follow attacker-controlled instructions instead of, or in conflict with, the instructions provided by the application.

For example, an AI assistant may be instructed to summarize documents while protecting confidential information. An attacker could include hidden or misleading instructions inside a document telling the AI to ignore its previous rules and reveal restricted information.The risk becomes more significant when AI systems can access internal data, external applications, APIs, databases, or automated tools.

How Does a Prompt Injection Attack Work?

LLMs process natural-language instructions and other content as part of their context. Attackers exploit this behavior by inserting instructions designed to change how the model responds or acts.

A prompt injection attack can be direct or indirect.

In direct prompt injection, the attacker enters malicious instructions directly into the AI application’s prompt or chat interface. The attacker may attempt to override system instructions, bypass restrictions, or obtain information the application should not provide.

Indirect prompt injection occurs when malicious instructions are placed in external content that an AI system later processes. This could include a webpage, email, document, or other data source. When the AI reads the content, it may interpret the embedded instructions as something it should follow.

What Are the Risks of Prompt Injection?

The impact of prompt injection depends heavily on what the AI system can access and what actions it is allowed to perform.

Potential risks include unauthorized disclosure of sensitive information, manipulation of AI responses, misuse of connected tools, unintended transactions, and bypassing application-level safeguards.

Prompt injection can be especially concerning agentic AI systems because these systems may be able to perform actions rather than simply generate text. If an AI agent has excessive permissions, manipulated instructions could potentially influence actions across connected systems.

Prompt Injection vs. Jailbreaking

Prompt injection and AI jailbreaking are related but are not exactly the same. Prompt injection generally targets an AI-enabled application by introducing instructions that interfere with its intended behavior or trusted instructions. Jailbreaking typically focuses on getting an AI model to bypass its built-in restrictions or safety controls.

Both techniques attempt to manipulate model behavior, but their targets and objectives can differ.

How Can Organizations Prevent Prompt Injection?

There is no single control that completely eliminates prompt injection risk. Organizations should use multiple security measures to reduce both the likelihood and potential impact of an attack.

These measures can include separating trusted instructions from untrusted content, validating model inputs and outputs, limiting the permissions available to AI systems, requiring human approval for sensitive actions, and monitoring AI interactions for suspicious behavior.

Organizations should also apply least-privilege principles to AI agents and restrict their access to sensitive systems and data.

Ultimately, prompt injection security requires organizations to treat external prompts and content as potentially untrusted. Combining strong access controls, continuous monitoring, input and output safeguards, and AI-specific security testing can help reduce the risks associated with prompt injection attacks.

Five Ways You Can Use Deception in the Mythos-like AI Era
use deception for ai-threats Cover

Want to Dive Deeper?

Enhance your perspective with additional analysis and experts take!

One Platform for All Adversaries

See Fidelis in action. Learn how our fast and scalable platforms provide full visibility, deep insights, and rapid response to help security teams across the World protect, detect, respond, and neutralize advanced cyber adversaries.

Proactive Threat Hunting: What It Is and What It Isn’t

Debunk the myths around proactive threat hunting and discover how it helps uncover hidden threats and attacker activity.

Insights from the Latest Global Network Security Report
Read the report on emerging cyber threats, AI-powered attacks, and strategies to strengthen security and resilience.