Insights / 8 min Read
Prompt Injection: Why the Fix Is Permissions, Not a Better Prompt
Prompt injection is not a prompting bug, and it is not patched at the prompt layer. Here is what it actually is, and what genuinely reduces the damage.
Published 2026-09-10
Prompt injection is usually introduced as a trick: someone types "ignore your previous instructions" into a chatbot and it does something silly. That framing is why it gets under-rated. The silly version is a demo. The version that matters involves no typing by an attacker at all, and it is not fixed by writing a firmer system prompt.
The actual mechanism
A language model receives one stream of text. Your system prompt, the user's message, a retrieved document, the contents of an email, the body of a fetched web page - by the time the model sees them they are the same kind of thing: tokens in a context window. The model has no channel that marks some of those tokens as data and others as instructions, because there is no such channel to have.
That is the whole vulnerability, stated once: a language model cannot reliably distinguish content it was asked to process from an instruction addressed to it.
Everything else follows. If your agent summarises support tickets, whoever writes a ticket is writing into the model's instruction stream. If it browses, whoever controls the page is. If it reads a shared document, or a calendar invite, or a code comment, or a row in a table it queried - each of those is a place a sentence can be planted by someone who never touched your interface. This is indirect prompt injection, and it is the form that shows up in real incidents, because it needs no access to your product at all.
Why prompt-layer defences do not close it
The instinctive fix is to write better instructions: tell the model to ignore any instructions found inside documents, wrap untrusted text in delimiters, add a classifier that screens input for injection attempts. All three help. None of them is a boundary.
Instruction hierarchies are requests, not enforcement - you are asking a probabilistic system to reliably follow one rule about all future inputs, and the failure rate is not zero. Delimiters are guessable and can be closed by the injected text itself. Classifiers face an open-ended input space where an attacker gets unlimited attempts and only has to succeed once: attacks arrive in other languages, in encodings, split across documents, embedded in images, or phrased so innocuously that a screen tuned to catch them would also block ordinary work.
The general shape of the problem: these defences reduce the probability of a successful injection. They do not bound the consequence of one. A control that works ninety-nine times in a hundred is a rate reduction, and against an attacker who can retry it is barely that.
The question that actually matters
Stop asking whether the model can be tricked - assume it can - and ask what happens next. If an attacker got to write the agent's next instruction, what would the agent be able to do?
That reframing turns an unbounded AI problem into an ordinary security problem, which is a much better position, because ordinary security problems have known answers. An agent that can only read public documentation and write a draft is a nuisance when compromised. An agent holding a credential that can issue refunds, email customers, or read an entire file store is a breach.
Two failure modes are worth naming separately because teams tend to see only the first. Unauthorised action is the agent doing something it should not - sending, deleting, paying, granting. Exfiltration is quieter: an injected instruction tells the agent to take something sensitive from its context and put it somewhere the attacker can read, and the delivery mechanism can be as mundane as a URL the agent fetches with the data in the query string, or a Markdown image the client renders. Any capability to send an outbound request is a capability to exfiltrate.
What actually reduces the damage
Scope the credentials to the job, not to the team
Agents inherit permissions casually - a service account made for a pilot, reused, never narrowed. The rule is the one that already applies to any other process: an agent gets the narrowest credential that lets it do its specific job, its own identity rather than a shared one, and read-only wherever reading is enough. Most of the alarming findings in an agent security review are permissions nobody deliberately granted.
Put a human in front of the irreversible things
Not in front of everything - approval fatigue is real, and a gate that fires forty times a day gets clicked through without being read. In front of the actions that cannot be undone: money moving, external communication, deletion, permission changes. The gate has to show what is about to happen in terms a person can evaluate in a few seconds, and it has to be enforced somewhere the model cannot reach.
Separate the privileged from the untrusted
The most reliable structural fix is to stop letting one context hold both. An agent that reads untrusted content does not also hold the dangerous credentials; it hands a structured, constrained request to a component that does, and that component validates the request against rules of its own. This is unglamorous plumbing and it is the closest thing to an actual boundary that exists today.
Control what can leave
Treat outbound requests as the exfiltration channel they are. Allowlist the domains an agent may reach rather than blocklisting the ones it may not, be careful about rendering model-authored links and images in clients, and be deliberate about what sits in a context window at the moment an untrusted document arrives in it.
Record enough to reconstruct what happened
None of the above will hold perfectly. When something does get through, the difference between a contained incident and an indefinite one is whether you can reconstruct the run: the input, the retrieved documents, the tool calls, the identity that authorised them, the outcome. That has to be instrumented before the incident, which is the part that gets deferred, because it is invisible right up until it is the only thing anybody wants.
The honest summary
Prompt injection is not solved, and treating it as solvable at the prompt layer is the mistake that produces the worst outcomes. The realistic posture is to assume the model will occasionally be made to do something you did not intend, and to make sure that what it can do in that moment is small, reversible, visible and recorded. That is a permissions and architecture problem. It is tractable in a way the model problem is not.
General information, not legal or professional advice about your organisation. Regulatory obligations depend on specific facts; take advice from qualified counsel before acting. See the Terms of Use.
Questions This Raises
What is the difference between prompt injection and jailbreaking?
Jailbreaking is a user getting a model to bypass its own safety training - the person talking to it is the attacker. Prompt injection is a third party planting instructions in content the model processes, so the user and the attacker are different people and the user may never know. Indirect prompt injection is the form that matters for agents, because it needs no access to your product at all.
Can prompt injection be fixed with a better system prompt?
No. A system prompt is a request to a probabilistic system, not an enforced boundary, and the model has no way to mark some tokens as data and others as instructions. Instruction hierarchies, delimiters and input classifiers all lower the success rate; none of them bounds what a successful injection can do. The consequence is bounded by permissions, approval gates and architecture instead.
What is indirect prompt injection?
An attack where the malicious instruction is planted in content the model reads rather than typed by the user - a web page, an email, a shared document, a support ticket, a code comment, a retrieved database row. It matters more than the direct form because the attacker never needs access to your application, and the legitimate user is unaware anything happened.
How do I know whether my agent is exposed?
Trace the path from every source of untrusted text into the agent's context, then list every tool, API and credential reachable from that context, marking which of those actions are irreversible or can send data outbound. Where an untrusted path reaches a privileged action with nothing enforced in between, that is the exposure - and it is a review you can do on paper before you test anything.