
Prompt Injection: The Vulnerability Class Your AI Feature Inherits
If your application sends untrusted content to a model that can act, you have inherited a vulnerability class with no complete fix.
GuardsArm Team
Security Experts

If your application passes content into a language model, and that model can take actions or reach data, you have inherited a vulnerability class that does not yet have a complete solution.
The root cause is architectural. Models process instructions and data through the same channel. A conventional application can separate code from input — that is what parameterised queries do. A model receives one stream of text and has no reliable mechanism for deciding which parts of it are authoritative.
Direct and indirect
Direct injection is a user telling the model to disregard its instructions. It matters mainly where the model's own instructions are sensitive or where it holds privileges the user does not.
Indirect injection is the serious case. Instructions are planted in content the model will process later — a web page it retrieves, a document it summarises, an email it reads, a support ticket, a code comment, a calendar invitation.
The attacker never interacts with your application. They place content somewhere your system will eventually ingest, and wait.
Where it becomes dangerous
Injection alone is a curiosity. It becomes a vulnerability when the model can do something.
| Model capability | What injection enables |
|---|---|
| Read-only, output to one user | Misleading output for that user |
| Retrieval across a document store | Retrieval of documents the user should not see |
| Sending email or messages | Exfiltration, or action taken in the user's name |
| Calling internal APIs | Whatever those APIs permit |
| Executing code | Whatever that environment permits |
| Autonomous multi-step operation | All of the above, without a human checkpoint |
Each row down that table increases the impact. The dangerous combination is untrusted input plus privileged capability plus no human confirmation, and that combination is exactly what an autonomous agent is.
Why filtering does not solve it
Input filtering for injection attempts fails for the same reason spam filtering never finished: the space of phrasings is unbounded, the attacker adapts, and attempts can be encoded, translated, split across documents or hidden in markup.
Guardrail models help and are worth using. They are a probabilistic layer reducing the rate, not a boundary you can rely on. Designing as though filtering is sufficient produces a system that fails silently the first time someone phrases it differently.
What actually contains it
Least privilege for the model. It should hold only the access required for its function, scoped to the user on whose behalf it acts. A retrieval system that queries with the user's permissions, not a service account's, removes an entire class of exposure.
Human confirmation before consequential action. Sending, paying, deleting, modifying records — a person confirms. This is the single most effective control and the one most often removed for convenience.
Treat model output as untrusted input. If output flows into another system, validate it there. Injection frequently chains through a second component that trusted what the model returned.
Separate trust domains. A model processing untrusted external content should not be the same one holding privileged internal capability.
Log everything — inputs, retrieved context, outputs, actions. Investigating an incident without the retrieved context is close to impossible.
Questions for a vendor selling you an AI feature
- What can the model do on my behalf without a human confirming?
- What untrusted content does it process?
- Does retrieval respect the individual user's permissions?
- What is logged, and can I see it?
- How do you test for injection, and can I see results?
See AI vendor risk for the broader assessment, and secure development for where this fits in the lifecycle.
Where to start
List what your AI features can actually do without a human in the loop. If anything on that list is consequential and also processes content from outside your organisation, that is where to put a confirmation step — and it is usually a small change.
GuardsArm tests AI-enabled applications and reviews their architecture. See application penetration testing or book a scoping call.
Written by GuardsArm Team
Our team of cybersecurity experts brings decades of combined experience in penetration testing, compliance auditing, and incident response. We're dedicated to helping organizations strengthen their security posture.
Take the next step on this topic
Talk to the GuardsArm team about how these services apply to your environment.