Skip to content
Technology

LLM Phishing Classification: BEC Detection with a Local LLM in the Mail Pipeline

|
|
6 min read

Classic spam and phishing filters check reputation, RBLs, SPF/DKIM/DMARC and content rules. You need them. But against the phishing of 2026, they are not enough. Today's phishing mail comes from a valid domain that was taken over. It is DKIM-signed and has clean SPF. The text is in perfect German without typos. This is where LLM-based classifiers come in.

What an LLM Spots in a Mail That a Regex Misses

  • Odd tone: "Accounting needs an urgent transfer of €87,500 today to the following IBAN". The meaning matches business email compromise (BEC), but no keyword list would catch it.
  • Fake reply chains: the LLM flags a made-up mail history as "appears artificially constructed".
  • Fake branding: "We have detected unusual activity on your account" in a Microsoft layout, but with a tone that is slightly off.
  • Pressure and urgency: "please immediately", "strictly confidential, do not tell anyone". These are the mental triggers that awareness training talks about.
  • Attachments out of context: a mail thread with an attachment whose topic has nothing to do with the thread.

Why a Local LLM (Ollama) Instead of a Cloud API

Should you push mails through a cloud LLM API? A GDPR DPIA gives a clear answer: no. Not even when the provider offers an EU region. So the answer is a local LLM via Ollama with an optimized model (Llama 3 or smaller BERT variants). This has four benefits:

  • Data sovereignty: no mail content leaves your infrastructure.
  • Works offline: the mail filter stays up even when the internet is down.
  • Cost: no price per token.
  • Same result every time: model versions are pinned. A verdict on a mail today is the same tomorrow (unlike "GPT-4 update kills consistency").

How It Fits Into the Mail Pipeline

Inbound mail first passes the usual pipeline: RBL, SPF/DKIM/DMARC, hash reputation, ClamAV, YARA and the CAPE sandbox. Say the mail is "clean" up to this point, but certain heuristic triggers fire. Examples are a new sender, an external sender asking an internal recipient to act, or an Office maldoc attachment type. Then the mail goes to the LLM:

  1. Mail body and subject go to the local model via the Ollama API.
  2. The prompt asks: "Classify this mail – legitimate, suspicious, phishing, BEC. Justify in one sentence."
  3. The output comes back in a fixed structure (JSON with verdict, confidence and reasoning).
  4. If the verdict is not "legitimate", the mail goes to quarantine. The operator sees the LLM reasoning as the explanation.

Tuning False Positives

Three tools keep the false positive (FP) rate under control:

  • Confidence threshold: only verdicts above 75 % confidence send a mail to quarantine. The rest is delivered with the tag "LLM-uncertain".
  • Sender reputation: a known sender (over 6 months of mail traffic without complaints) gets a bonus.
  • Manual override: the operator can whitelist a given sender/recipient pair for good.

Real-World Performance

Llama 3 8B on an RTX 4090 handles about 5–8 mails per second with 200 ms latency. A typical mid-sized company gets 5,000 mails a day (β‰ˆ 0.06 mails/sec on average). For that load, this setup is far larger than needed, but it copes with the morning peaks. CPU-only also works, but latency then rises to 2–4 s per mail.

Compliance Mapping

LLM-based classification covers requirements from:

  • NIS-2 Art. 21(2)(b): measures to detect incidents.
  • BSI IT-Grundschutz APP.5.3: extra filter stages for email security.
  • ISO 27001 A.5.7: threat intelligence, with detection depth as evidence.

Conclusion

In 2026, LLM-based phishing classification marks the gap between a spam filter that sorts by gut feeling and one that understands what a mail means. Run locally with Ollama, it is GDPR-compliant and gives the same result every time. In the SecTepe.Comm mail pipeline it is an optional add-on. Teams that turn it on usually see 30–50 % more phishing detections, with a slightly higher FP rate.