papersSEP 12 04:00 UTC
DriftNet: Dual-Head Transformer Detects and Locates Prompt Injection in LLM Agents
A new arXiv paper introduces DriftNet, a dual-head trajectory transformer designed to detect indirect prompt injection in LLM agents and pinpoint where in the agent's action sequence the compromise occurred. The approach treats a successful attack as a visible behavioral pattern: a benign run of tool calls, a poisoned observation, then attacker-serving actions. This would give operators more granular visibility into agent security incidents than a simple pass/fail detection signal.