Prompt Injection
Prompt injection is an attack in which text crafted by an adversary is fed to an AI model so that the model follows the attacker's instructions instead of those of its user or developer.
What Is Prompt Injection?
A large language model receives its instructions and its data through the same channel: text in the context window. It has no reliable, built-in way to tell "this is a command from my user" apart from "this is content I was asked to read." Prompt injection exploits that gap. If an attacker can get words into the model's context, those words can try to override the system prompt, change the task, or trigger actions. The term was coined in 2022 by analogy to SQL injection, where data is mistaken for code. Unlike SQL injection, there is no equivalent of parameterized queries that cleanly separates the two, which is why prompt injection is often described as a problem to be managed rather than one that can be fully fixed.
Direct vs. Indirect Injection
Direct prompt injection is when the user typing into the model is the attacker, for example by telling a chatbot to ignore its instructions and reveal its system prompt. It overlaps with jailbreaking, which aims to get around a model's safety training.
Indirect prompt injection is more dangerous. Here the user is the victim. The malicious instructions sit in content the model reads while doing its job: a web page, an email, a PDF, a code comment, a calendar invite or a tool result. The user asks an agent to summarize a page; the page tells the agent to forward the user's files somewhere. Because AI agents read untrusted content constantly and can act on it, indirect injection is the main security problem of the agent era.
Injection Through Tools, Repositories and Images
Coding agents and browser agents are the most exposed, because they combine untrusted input with real permissions: file access, shell commands, credentials and logged-in sessions.
Ghostcommit
In July 2026, researchers disclosed Ghostcommit, a proof-of-concept attack that hid a prompt injection inside a PNG image. The repository's AGENTS.md file pointed coding agents at the image, and the instructions in it made Cursor and Google's Antigravity leak secrets from .env files. Claude Code refused under every model the researchers tested. The attack was a research demonstration and has not been seen in the wild, but it showed that injections can arrive through images and through the configuration files agents trust most.
Browser Agents
Agents that browse the web on a user's behalf, a form of computer use, read whatever pages they visit. When Anthropic made Claude in Chrome generally available on all paid plans on August 26, 2026, it published red-team attack success rates for prompt injection: 16.7% for Claude Opus 4.5, 0% for Claude Opus 5 and Sonnet 5, and 0.3% for Fable 5. Even with those figures, Anthropic cautioned that "prompt injection remains a moving target." Low measured rates reflect a particular test set; attackers adapt.
Relation to MCP and Tool Poisoning
The Model Context Protocol lets agents connect to thousands of external tools and data sources, and every one of them is a potential injection channel. Two forms matter most. First, a tool's output, such as a web search result or a support ticket, can carry injected instructions just like any document. Second, in tool poisoning, the tool's own description or metadata, which the model reads to decide how to use the tool, contains hidden instructions. A malicious or compromised MCP server can then steer an agent without the user ever seeing the text. As marketplaces grow, such as the Claude Marketplace launched in September 2026 with more than 2,000 connectors and plugins, vetting third-party tools becomes part of the defense. The MCP specification released on July 28, 2026 aligned authorization with OAuth and OpenID Connect, which helps control what a tool can access, though it does not stop injected text by itself.
Defenses
No single defense is sufficient, so practitioners layer several:
- Model training and classifiers. Models are trained to treat content as data, and separate classifiers scan inputs and tool results for injection attempts before or as the model reads them. The drop in Claude in Chrome's measured attack rates reflects this kind of work.
- Human approval. Requiring a person to confirm consequential actions, such as sending messages, making purchases, pushing code or visiting sensitive sites, limits what a successful injection can do.
- Sandboxing and least privilege. Running agents in isolated environments, with only the credentials and network access a task needs, contains the damage. Ghostcommit worked because secrets were within the agent's reach.
- Separating trusted and untrusted input. Agent harnesses can mark which text came from the user and which from the outside world, and restrict which tools may be called after untrusted content has been read.
- Vetting tools and configuration. Treating MCP servers, plugins and files such as AGENTS.md as code that needs review, not as inert settings.
Why It Matters
Prompt injection is the reason agent autonomy and agent security are in tension. Every gain in what agents can do, from reading email to running shell commands to paying for things in the agentic economy, raises what an attacker can achieve by slipping text in front of them. Results like 0% attack success on a red-team set are real progress, but the risk does not disappear while models still read instructions and data through the same channel. For anyone deploying coding agents or browser agents, prompt injection belongs at the center of the threat model, alongside the broader concerns covered in AI safety and AI in cybersecurity.
Further Reading
- Ghostcommit hides prompt injection in images to fool AI agents — BleepingComputer — The July 2026 image-borne injection against coding agents
- Claude in Chrome is generally available — Anthropic — Published prompt-injection attack success rates by model
- Prompt injection attacks against GPT-3 — Simon Willison — The 2022 post that named the attack
- Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection — arXiv — The foundational paper on indirect injection
- LLM01: Prompt Injection — OWASP — The top entry in OWASP's LLM risk list, with mitigations
- MCP protocol updates — Model Context Protocol blog — The 2026-07-28 spec and its authorization changes
- Claude Marketplace — Anthropic — The connector and plugin marketplace built on MCP