Back to Blog
Application Security
4 min read

Prompt Injection: The Vulnerability Class Your AI Feature Inherits

If your application sends untrusted content to a model that can act, you have inherited a vulnerability class with no complete fix.

GuardsArm Team

Security Experts

September 25, 2026

Prompt injection in business applications

If your application passes content into a language model, and that model can take actions or reach data, you have inherited a vulnerability class that does not yet have a complete solution.

The root cause is architectural. Models process instructions and data through the same channel. A conventional application can separate code from input — that is what parameterised queries do. A model receives one stream of text and has no reliable mechanism for deciding which parts of it are authoritative.

Architectural, not a bug
Instructions and data share one channel
Indirect is the real risk
Planted in content you will ingest later
Capability is the multiplier
Injection matters when the model can act

Direct and indirect

Direct injection is a user telling the model to disregard its instructions. It matters mainly where the model's own instructions are sensitive or where it holds privileges the user does not.

Indirect injection is the serious case. Instructions are planted in content the model will process later — a web page it retrieves, a document it summarises, an email it reads, a support ticket, a code comment, a calendar invitation.

The attacker never has to touch your application
Indirect injection works by planting instructions in a web page, a document, a support ticket or an email that your system will process later. There is no login attempt, no malicious request to your endpoint, and nothing for a perimeter control to see.

The attacker never interacts with your application. They place content somewhere your system will eventually ingest, and wait.

How indirect prompt injection reaches youAn attacker plants instructions in content your system will later retrieve; the model processes it as instruction and acts with whatever privileges it holds.Attackerplants instructionContentpage, document, ticketIngestyour system retrieves itModelreads it as instructionActionwith the model’s privileges
Every control worth having is in the last box: limit what the model may do.

Where it becomes dangerous

Injection alone is a curiosity. It becomes a vulnerability when the model can do something.

Model capabilityWhat injection enables
Read-only, output to one userMisleading output for that user
Retrieval across a document storeRetrieval of documents the user should not see
Sending email or messagesExfiltration, or action taken in the user's name
Calling internal APIsWhatever those APIs permit
Executing codeWhatever that environment permits
Autonomous multi-step operationAll of the above, without a human checkpoint

Each row down that table increases the impact. The dangerous combination is untrusted input plus privileged capability plus no human confirmation, and that combination is exactly what an autonomous agent is.


Why filtering does not solve it

Input filtering for injection attempts fails for the same reason spam filtering never finished: the space of phrasings is unbounded, the attacker adapts, and attempts can be encoded, translated, split across documents or hidden in markup.

Guardrail models help and are worth using. They are a probabilistic layer reducing the rate, not a boundary you can rely on. Designing as though filtering is sufficient produces a system that fails silently the first time someone phrases it differently.


What actually contains it

Controls that actually contain prompt injectionLeast privilege scoped to the user, human confirmation for consequential actions, and treating model output as untrusted are the effective controls. Filtering reduces the rate but is not a boundary.Least privilege, scoped to the userRetrieval with the user’s permissions, not a service accountHuman confirmation for consequential actionsThe most effective control, and the first one removedTreat model output as untrustedValidate it wherever it lands nextSeparate trust domainsUntrusted content and privileged capability in different placesGuardrail models and filteringReduces the rate; not a boundaryComprehensive loggingIncluding retrieved context, or investigation is impossible
Note where filtering sits: useful, and not something to rely on.

Least privilege for the model. It should hold only the access required for its function, scoped to the user on whose behalf it acts. A retrieval system that queries with the user's permissions, not a service account's, removes an entire class of exposure.

Human confirmation before consequential action. Sending, paying, deleting, modifying records — a person confirms. This is the single most effective control and the one most often removed for convenience.

Treat model output as untrusted input. If output flows into another system, validate it there. Injection frequently chains through a second component that trusted what the model returned.

Separate trust domains. A model processing untrusted external content should not be the same one holding privileged internal capability.

Log everything — inputs, retrieved context, outputs, actions. Investigating an incident without the retrieved context is close to impossible.


Questions for a vendor selling you an AI feature

  • What can the model do on my behalf without a human confirming?
  • What untrusted content does it process?
  • Does retrieval respect the individual user's permissions?
  • What is logged, and can I see it?
  • How do you test for injection, and can I see results?

See AI vendor risk for the broader assessment, and secure development for where this fits in the lifecycle.


Where to start

List what your AI features can actually do without a human in the loop. If anything on that list is consequential and also processes content from outside your organisation, that is where to put a confirmation step — and it is usually a small change.

GuardsArm tests AI-enabled applications and reviews their architecture. See application penetration testing or book a scoping call.

Written by GuardsArm Team

Our team of cybersecurity experts brings decades of combined experience in penetration testing, compliance auditing, and incident response. We're dedicated to helping organizations strengthen their security posture.

Take the next step on this topic

Talk to the GuardsArm team about how these services apply to your environment.