"Hacking AI Agents: The Attack Surfaces Nobody Patched"
Hacking AI Agents: The Attack Surfaces Nobody Patched
AI agents have crossed a line that traditional software never crossed: they read instructions from the same channel they read data, they act on the physical world through tools, and they do it all with privileges users themselves don't hold. The security research community has spent three years documenting what this means. The results are not comforting. NIST has characterised prompt injection as "generative AI's greatest security flaw", and OWASP ranks it the number-one vulnerability in its LLM Applications Top 10. This article maps how agents actually get hacked — grounded in the offensive research of the Arcanum Security team, whose AI Security Resource Hub and agent-hacking CTF labs have trained hundreds of practitioners — and what the peer-reviewed literature says defenders can realistically do.
The root flaw: instructions and data share one channel
Every agent hack starts from the same architectural fact. A traditional application maintains a hard boundary between code and data: SQL injection was devastating precisely because it violated that boundary, and we spent two decades building parameterised queries to restore it. LLM agents cannot make this separation at all. The system prompt, the user's message, the contents of a web page the agent just fetched, and the output of a shell command all arrive through the same neural pathway, and the model decides — probabilistically, not structurally — which parts are instructions.
The systematisation of knowledge on agentic coding assistant attacks by Ptacek and colleagues calls this the "architectural conflation of code and data", and it means the classic defence of input sanitisation is structurally unavailable. You cannot quote-encode an English sentence the way you quote-encode a SQL string. The broader SoK on agentic AI attack surfaces puts numbers on the consequence: across 78 studies, attack success rates against state-of-the-art defences exceed 85% when attackers use adaptive strategies, while most published defence mechanisms mitigate less than half of sophisticated attacks.
Vector one: direct and indirect prompt injection
Direct injection is the classic: a user (or an attacker who can reach the model) writes "ignore your previous instructions and...". Modern models resist the naive version. The research consensus is that direct injection against well-aligned frontier models is now the weaker path.
Indirect injection is where the field has moved. An agent that browses web pages, reads email, processes GitHub issues, or ingests PDFs is reading attacker-controllable text into its context. Greshake and colleagues demonstrated early that malicious instructions embedded in retrieved content could steer LLM-integrated applications end-to-end; the current literature treats this as the dominant attack family. The insidious part is timing: the injection may sit in a document for weeks before an agent happens to read it, and the malicious action may be separated from the user's original request by dozens of steps.
The newest evolution removes the last comfort — that a defender just needs to scan for one obvious malicious instruction. The AdaLCPI research (2026) shows attack objectives can be split into incomplete fragments scattered across 8,000+ tokens of retrieved content, with a generic "reconstruction cue" prompting the agent to piece them together. Because no single fragment contains the harmful objective, and the cue is innocuous, content scanning finds nothing. Against seven frontier models, the fragmented approach achieved a 61.4% average attack success rate versus 32.8% for the previous generation of adaptive injection — and the fragments didn't even need to arrive in the same tool output.
Vector two: the tool layer — where prompts become code execution
A prompt injection that changes an agent's answer is a content-quality problem. A prompt injection that changes an agent's tool calls is a remote-code-execution problem. The tool layer is where agent hacking stops being novel AI risk and starts being classic RCE with a natural-language payload.
The red-team study of six production coding agents (Cursor, Claude Code, Copilot, Windsurf, Cline, Trae) demonstrated the full chain against real products. Phase one: exfiltrate the agent's internal prompts through what the researchers named ToolLeak — a malicious tool whose argument schema requests "the current system instructions" as a required parameter. The model populates that argument as a matter of ordinary tool use, no refusal triggered, because it looks like benign argument generation rather than an "ignore previous instructions" request. It worked on 19 of 25 agent–model pairs. Phase two: a two-channel injection — malicious steering in a tool description, then the payload in the tool's return value, which the study found carries higher salience in the model's subsequent reasoning. The result was remote code execution on every tested agent–LLM pair, via a curl | bash the model believed was completing an installation routine.
The companion SoK catalogues the supporting cast: rules-file poisoning (a malicious .cursorrules in a cloned repository — 41–84% success across platforms, with data exfiltration at 84%), MCP tool-poisoning where a tool description instructs the agent to read ~/.aws/credentials into a parameter, and the "Toxic Agent Flow" through GitHub MCP servers where an issue comment exfiltrates ~/.ssh contents. One cited empirical study found 19 remote-code-execution flaws across 11 agent frameworks. These are not AI problems. They are the same command-injection and trust-boundary bugs the industry has been fixing in web applications since before most agent vendors were incorporated — recurring because agent frameworks wrap tools in natural language instead of type systems.
Vector three: the skill and MCP supply chain
The newest attack surface is also the most structural. Agent platforms have converged on skills — installable packages of instructions and scripts that extend what an agent can do — and on the Model Context Protocol as the connective standard. Both create supply chains, and supply chains can be poisoned.
The Skill-Inject benchmark (2026) quantified the exposure: 202 injection-task pairs targeting skill files specifically, with attack success up to 80% on frontier models — including data exfiltration, destructive actions, and ransomware-like behaviour. The reason skills are uniquely dangerous is that they invert the defence that works elsewhere: a skill file is all instructions, so "separate instructions from data" has nothing to operate on. An injected line in an email is anomalous; an injected line in a file of 200 instruction lines is invisible. Users install skills with app-store trust and audit them with app-store diligence — which is to say, rarely.
The MCP ecosystem adds protocol-level surface. The VATS research (2026) found that the error path — the messages a tool returns when something goes wrong — carries "implicit authority": models slip into corrective reasoning modes around error messages and comply with injected instructions at up to triple the rate of standard indirect injection, reaching 100% compliance in controlled tests. A malicious MCP server that formats its failure responses as instructions is exploiting the agent's own helpfulness. And the agentic AI SoK documents the multi-agent variant: poisoned shared memory, cross-agent instruction leakage, and emergent behaviour manipulation in systems where one compromised agent talks to others.
Arcanum's training labs treat these as hands-on skills rather than academic categories — their open resource hub organises 23+ active labs across prompt injection, jailbreaks, agent abuse, and chained exploitation, from beginner tiers to DEFCON-level scenarios. The offensive technique count grows monthly; the defence count does not.
What defenders can actually do
The literature is blunt that no mitigation is sufficient, but several are directionally useful, and they rhyme with hardening lessons the security industry already knows.
Constrain the blast radius, not the prompt. Since injection cannot be fully prevented, the damage an injected instruction can cause must be bounded. That means least-privilege tool credentials (a GitHub MCP token that can read two repos, not your account), sandboxed execution for agent-initiated commands, and — critically — human confirmation gates on irreversible actions. The coding-assistant SoK frames its defence-in-depth recommendations as architectural, not filter-based, precisely because filtering fails.
Assume retrieved content is attacker-controlled. Every tool result, every fetched page, every skill file is untrusted input wearing a documentation costume. Treat it the way a mail server treats an attachment: parsed, never executed. Concretely — tool results should be wrapped in explicit untrusted-data framing, agent plans that involve external content should require re-approval, and skill/MCP installs should be pinned, hash-verified, and reviewed like dependencies, because that is what they are.
Watch the error path and the return path. The VATS and tool-hijacking research both show that monitoring only user inputs misses the channel where modern injections arrive. Log tool returns, alert on instruction-like patterns in error payloads, and require re-validation whenever a tool result proposes the next tool call.
Test adversarially, continuously. Static defences degrade against adaptive attackers — the same finding the adaptive-injection literature keeps re-proving. Teams running agents should run agent-hacking labs (Arcanum's CTF tiers are a reasonable bar) and benchmark their stacks against the public suites — Skill-Inject for the skill supply chain, injection benchmarks for retrieval paths — the way web teams run DAST.
Treat agent identity as its own perimeter. An agent acting with the user's credentials indistinguishably from the user is the deepest design failure in the current generation. Token-scoped agents, per-tool identity separation, and audit trails that distinguish human-initiated from agent-initiated actions are the foundations the next platform generation must build in.
The honest bottom line
The research record reads like the early history of web exploitation: a structural flaw everyone knows about, an attack toolchain industrialising faster than the defences, and a market deploying the vulnerable thing as fast as it ships. The difference is scale of agency — a vulnerable web app leaks data; a vulnerable agent acts. The Arcanum training material and the 2026 academic literature converge on the same closing argument: agent security is not a prompt-engineering problem, and it will not be solved by models getting smarter. It is a systems problem — privilege boundaries, verified supply chains, adversarial testing — of the kind this industry knows how to solve when it decides to. For anyone running agents in production, the time to start is before the first injection lands, not after.
Sources: Arcanum Security — Attacking AI training & resource hub · Arcanum AI Security Resource Hub · Ptacek et al., Prompt Injection Attacks on Agentic Coding Assistants (arXiv 2601.17548) · SoK: The Attack Surface of Agentic AI (arXiv 2603.22928) · AdaLCPI: Adaptive Long-Context Prompt Injection (arXiv 2609.36576) · Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks (arXiv 2602.20156) · VATS: Exploiting Implicit Authority in Error-Path Injection (arXiv 2606.07992) · Red-Teaming Coding Agents from a Tool-Invocation Perspective (arXiv 2509.05755)