First-line triage without first-line headcount
Explanations and prioritisation arrive attached to the alert, so a smaller team covers more ground. This is a capacity argument rather than a replacement one.
An on-prem AI analyst that triages, correlates and explains — using a local LLM, so your data never leaves your network.
Most AI security tooling sends your telemetry to a vendor model. The GuardsArm AI SOC Analyst runs a local language model on your own infrastructure, which is what makes it usable in environments that cannot export security data at all.

What is at stake
A small team facing an unbounded queue either misses things or burns out, and the usual remedy — a hosted AI analyst — is unavailable to anyone who cannot export telemetry.
Explanations and prioritisation arrive attached to the alert, so a smaller team covers more ground. This is a capacity argument rather than a replacement one.
An alert that explains itself in plain language shortens the ramp for new staff and makes handovers between shifts less lossy.
Because inference runs on your hardware, regulated and air-gapped operators get the capability rather than having to decline it. For a great many buyers this is the only version of AI-assisted triage they are permitted to adopt.
Who carries this
The CISO makes this case on staffing and coverage. Legal and compliance sign off on it because nothing leaves the perimeter.
The problem
The useful thing an AI analyst does is read an alert in context and say what it means. Doing that requires the alert, the hostname, the username, the command line, and ideally the history of the asset. That is, by some distance, the most sensitive data a security team holds — and in almost every AI-assisted security product, it is posted to a third-party inference API to get an answer back.
For a regulated organisation this is frequently not a trade-off to be weighed but a condition that cannot be met. A HIPAA-covered entity, a government department, or any operator running air-gapped has no version of this where the data leaves.
GuardsArm runs the model on your hardware. There is no external AI provider in the path, which is why the capability is available to organisations that would otherwise have to decline it entirely.
A local model is a smaller model. The design compensates by never letting the platform depend on it.
That is the honest trade, stated directly: on-premises inference means a local language model rather than a frontier hosted one. The architecture is built around that fact. The AI is a downstream consumer of detections, never the source of truth. Detection, severity grading, ATT&CK mapping and correlation all happen in the detection engine and reach SOC Operations independently. If the AI service is stopped, the platform keeps detecting and keeps alerting; what you lose is the explanation layer, not the security control.
Capabilities
Separate models handle triage, correlation and reporting rather than one general-purpose prompt.
The language model runs on your hardware. There is no call to an external AI provider, which means no egress of alert content, asset names or user identities.
A RAG knowledge base lets the analyst reason over your estate and your past incidents, so its explanations reference what you actually run.
Each alert can be explained in ordinary language — what fired, why it matters here, and what the analyst should check next.
Triaged incidents arrive in a queue, and actions that need sign-off wait in an approval center rather than executing unsupervised.
Analyst feedback on verdicts feeds back into the service to reduce false positives over time.
Analyst playbooks ship with the service, and programmatic access uses scoped keys so automation gets only the permissions it needs.
How it fits the platform
The architectural principle matters as much as the capability: the AI is a downstream consumer of detections, never the source of truth. Alerts reach SOC operations independently, so the platform keeps working — and keeps detecting — whether or not the AI service is running.
How it works
Note where the AI enters: after the finding already exists and has already reached the console.
The detection engine promotes the event, assigns severity from the shared scale and writes the finding. SOC Operations sees it immediately. This step does not involve the AI service at all, and that is deliberate.
As a downstream consumer, the AI picks up findings and incidents from the same pipeline the console reads. It is a reader of the detection record, never a writer of the detection decision.
A retrieval-augmented knowledge base supplies the relevant history — this asset, this user, prior incidents that looked like this one, your own documented environment. Generic security knowledge without your context produces generic conclusions.
Triage, correlation and reporting are handled by separate purpose-built models rather than one prompt attempting all three. Each has a narrower job and a correspondingly narrower failure mode.
The language model runs on your hardware. The alert, the asset names, the usernames, the command lines — none of it is sent to a third-party AI provider, because there is no third-party AI provider in the path.
Explanations attach to alerts; proposed actions enter the approval center rather than executing. Analyst feedback feeds the learning loop, which is how false-positive patterns get suppressed over time.
The agents
Triage, correlation and reporting are different tasks with different failure modes. Splitting them is what makes the output auditable.
Privacy
The comparison is with the delivery model rather than any named AI provider, because what differs is structural.
| Aspect | Hosted AI in a security product | GuardsArm AI SOC Analyst |
|---|---|---|
| Where inference runs | In the AI vendor's cloud, on their accelerators. | On your hardware, inside your perimeter, on a local model. |
| What leaves the network | Alert content, hostnames, usernames, command lines — whatever is in the prompt. | Nothing. There is no outbound call in the inference path. |
| Air-gapped deployment | Not possible; the model is reached over the internet. | Supported. Dependencies are self-hosted, so the stack runs with no egress. |
| What happens if the AI is down | Depends on the product. Often triage stops. | Detection and alerting continue unaffected. The AI is downstream by design. |
| Who the learning benefits | Typically the vendor's aggregate model. | Your deployment. The knowledge base and feedback stay yours. |
Guardrails
An AI analyst nobody can check is a liability. These four things are what make the output something a team can act on.
Structured procedures the AI follows for common investigation types, so its output is consistent between runs rather than reinvented each time. Consistency is what makes AI output reviewable.
Automation reaches the AI service through keys scoped to specific capabilities. An integration that should only read explanations cannot approve a containment action.
Anything the AI proposes that would change the state of your environment waits for a human. The gate is on by default, and what may bypass it is an explicit policy decision rather than a default.
Analysts mark explanations and classifications right or wrong. That signal is retained in your deployment and shapes subsequent triage — a loop that cannot exist where the model is shared across a vendor customer base.
Background
Plain explainers on the underlying ideas, written for someone evaluating rather than buying.
Questions
No. It runs a local language model on your own infrastructure. There is no call to an external AI provider, which means no egress of alert content, asset names or user identities.
Nothing. The AI is a downstream consumer of detections, never the source of truth. Alerts reach SOC operations independently, so the platform keeps detecting and responding whether or not the AI service is running.
A retrieval-augmented knowledge base lets it reason over your estate and your past incidents, so explanations reference what you actually run rather than generic descriptions.
Actions needing sign-off wait in an approval center rather than executing unsupervised. That is deliberate: automation with significant consequence and imperfect confidence turns a false positive into an outage, and the first time that happens to something important the automation gets switched off.
A feedback and learning loop takes analyst verdicts back into the service to reduce false positives. How much that helps depends on how consistently analysts record feedback.
Works with
Every module runs in the same self-hosted stack and shares the same telemetry, severity model and response engine.
We will walk through the console, the deployment model and what it takes to stand it up in your environment.