Tool Name Exploitation and Confused Deputy Vulnerabilities in AI Agents
AI agents inherit a decades-old security flaw that tool descriptions now weaponize at scale.

A deputy confused about who it's working for has a long history in computer security, but AI agents revive this old vulnerability with a new twist. An agent reads operator instructions and untrusted outside content through the same inference pathway. There is no hard wall separating code from data, no gate that tells the model "this part is a command, this part is just text to read." Everything arrives as words, so the model has to decide, token by token, what those words mean.
Norm Hardy named this failure mode back in 1988: authority should travel with whoever is asking for an action, not sit baked into the program carrying it out. A compiler that holds broad file-system rights and blindly trusts whatever filename a user hands it is a deputy confused about who it's actually working for. Swap "compiler" for "AI agent" and the picture barely changes. Agents hold ambient credentials. They act on inputs they have no reliable way to verify. Almost four decades on, Hardy's insight still fits today's agent stack.
The Cloud Security Alliance's 2026 analysis names three things that make this worse right now. Agents are built to treat anything in their context window as potentially instructive, so the data/code boundary that older security models leaned on isn't there. The broad permissions that make an agent genuinely useful, booking flights, merging code, moving money, are the same permissions that make a successful injection severe and sometimes impossible to undo. And multi-agent setups create paths where one compromised step can travel across organizational lines with no human checking in along the way.
That leaves a simple, uncomfortable fact: any input channel an agent reads is a potential attack vector. An email inbox, a GitHub issue tracker, a retrieved PDF, a tool's own output text, all of it. Whatever the agent interprets as an instruction, it carries out using the operator's full credentials.
The four-stage anatomy of a confused deputy attack chain
A confused deputy attack on an AI agent follows the same four steps regardless of the incident's specifics. Think of it as a diagnostic lens: once you can spot each stage, you can recognize the pattern in almost any case that gets reported.
Stage one is the injection vector. An adversary plants malicious instructions wherever the agent will read them, an issue title, a code comment, a retrieved document. They don't need to write code. They just need to control content.
Stage two is instruction interpretation. The agent processes natural-language commands and natural-language content through the same mechanism, so the planted instruction looks, from the inference engine's point of view, no different from something the operator typed directly. Nothing downstream can tell the two apart.
Stage three is credential exercise. The agent acts using whatever authority it already holds, OAuth tokens, API keys, shell access, a repository connection, granted earlier by the operator for entirely legitimate reasons. ScopeGate's threat modeling sharpens this stage into a single concrete goal for the attacker: get the agent to emit a side-effecting tool call with attacker-chosen arguments. Repoint a payout account. Issue a refund that shouldn't exist. Pull a secret out of a config file. Fetch a URL that triggers a server-side request forgery. The attacker never touches the tool itself. They only control the text the model reads.
Stage four is the outcome, and it's often one-way. A package gets published. A file gets overwritten. A payment goes out. Credentials leave the building. If the injected instruction told the agent to stay quiet about what it did, the agent may do exactly that.
How MCP tool descriptions become an instruction channel attackers control
The underlying protocol that lets agents connect to external tool servers is the infrastructure that took this general flaw and turned it into something attackers can run at scale. MCP lets an agent connect to many tool servers in a single session, reading each server's tool descriptions to learn what the tool does and how to call it. Those descriptions are text. The agent treats them as trustworthy instructions about its own capabilities. Whoever writes a tool description gets a direct line into the agent's reasoning.
Tool poisoning exploits this directly: an adversary hides instructions inside a tool's description that tell the agent to exfiltrate data, take unauthorized action, or suppress any notification about what it just did, all before a human ever sees a prompt. Cross-server tool shadowing takes it further. Because one session can link an agent to several MCP servers at once, a malicious server can inject descriptions that redefine how the agent understands tools from a completely different, trusted server sitting right next to it.
Research on compositional threats in LLM agent systems places this kind of attack in its own category: a harness-level implant, something embedded in a skill, a plugin, a tool description, a hook, or a configuration file. That's distinct from a one-time prompt injection that fires and disappears, and distinct from memory poisoning that persists across sessions. Each needs its own defense, because each lives in a different part of the system.
The MCPoison pattern, tracked as CVE-2025-54136 in Cursor IDE, shows how patient this can be. An attacker commits a harmless-looking MCP configuration to a shared repository. A developer reviews it, approves it inside Cursor. Later, the attacker swaps in a malicious payload through a follow-up commit, and because the configuration was already approved once, every subsequent session runs the attacker's commands with no re-approval prompt.
Living-off-the-Agent: when the attack needs no payload of its own
Some of the most dangerous confused deputy attacks skip the malware step. Living-off-the-Agent, or LotA, describes an attack where the adversary brings no custom payload. The agent's own existing tool set, the same tools it was given to do its job, supplies everything needed for data exfiltration, lateral movement, or privilege escalation.
A legitimate tool call becomes the channel that moves the stolen data out. A confused deputy chain can do what would otherwise need privilege escalation malware. The agent does the work, using permissions the operator already handed it.
Intruders already do this in endpoint security, where they use a machine's own built-in utilities instead of dropping custom tools onto disk, and this is its direct descendant. Defenders can't block those everyday scripting and installer tools because they are legitimate and needed every day. They have to catch the misuse in context, which is a much harder problem than catching a known-bad file.
That's the core reason LotA attacks are so hard to stop with traditional tooling. Signature-based scanning and payload inspection look for a malicious artifact, and there isn't one to find. The malicious element is an instruction sitting in plain text, not a piece of software.
Research modeling the "Order 66" scenario, named for the fictional moment when a trusted population turns on itself at a single trigger phrase, shows how this compounds with dormancy. A harness-level implant can sit completely inactive through testing and evaluation, pass every check, then activate later on some trigger, and exercise the agent's full granted authority the moment it wakes up. Pre-deployment scanning catches what's active at the time of the scan. It has nothing to say about what's sleeping.
Four concrete attacks that show the chain executing in production
Four disclosed incidents in 2026 show this four-stage chain firing in real systems, each through a different door but with the same structure.
In February 2026, the Cline AI coding assistant was compromised through a crafted GitHub issue title [1][16]. That title carried a malicious instruction (the injection vector); an authenticated Claude coding session read it, interpreted it as a legitimate directive (instruction interpretation), then acted on it using its existing credentials to install an attacker-controlled package (credential exercise). That package went out as an official update to roughly 4,000 developer machines (the outcome), the Cloud Security Alliance's 2026 analysis found.
GhostApproval, disclosed July 8, 2026, worked through symlinks. A malicious repository contained a workspace path that resolved, via symlink, to a sensitive file sitting outside the repository entirely, something like ~/.ssh/authorized_keys or ~/.zshrc. The approval dialog shown to the developer displayed the harmless-looking workspace path. The agent then wrote to the real, resolved, external target instead, a textbook case of CWE-451, UI misrepresentation of critical information. In several documented instances, the agent's own internal reasoning correctly identified the dangerous target, and the confirmation prompt still hid that fact from the person approving it. Wiz's documented remediation guidance is direct: resolve the canonical target before any read, write, or approval prompt, flag anything outside the workspace explicitly, and never let a write reach disk before authorization is actually granted.
The Azure DevOps MCP case, disclosed July 22, 2026, moved the vector into code review itself. Manifold Security researchers found that Microsoft's official Azure DevOps MCP server had applied prompt injection defenses, a technique called "spotlighting", to several of its tools, but not to the pull request retrieval function. HTML comments invisible in the rendered PR interface were returned verbatim by the underlying REST API. An attacker submitting a pull request could hide instructions inside one of those comments, invisible to any human reviewer but fully visible to a more privileged reviewer's agent reading the raw API response. MSRC acknowledged and triaged the report.
The exfiltration mechanism wasn't a new tool smuggled onto the machine. It was the agent tooling already installed there, doing what it was built to do, just on the attacker's instructions instead of the operator's.
Four different entry points, one coding assistant, one IDE symlink trick, one DevOps review pipeline, one package ecosystem, and the same four stages run in every one of them.
Visual grounding failures as a separate but structurally identical attack class
Not every agent reads text. Computer-using agents, or CUAs, act directly on what they see on a screen, and that visual channel carries its own version of the exact same confused deputy problem. A formal analysis of this failure mode argues that perception, for a CUA, isn't just an input modality the way a camera feed is input to a self-driving car. Perception is the trust anchor the agent uses to decide what action it's even authorizing. When that anchor slips, the result is a security failure, not a performance blip.
Three things can cause it. Visual grounding errors happen when the agent simply misreads what's on screen and clicks or authorizes an action against the wrong object, and this already happens often enough in ordinary, non-adversarial use. Adversarial screenshot manipulation happens when a compromised runtime or a sitting-in-the-middle process swaps or alters the image the agent is actually looking at. And TOCTOU races, time-of-check to time-of-use, happen when the screen changes between the moment the agent decides on an action and the moment that action executes, so the agent ends up acting on a state that no longer exists by the time it hits "confirm."
Researchers demonstrated an attack called ScreenSwap that needs only eight lines of code to swap pixels and induce privilege escalation this way, and the result is indistinguishable from an ordinary CUA misclick. The line between a routine perception mistake and a deliberate exploit turns out to be thin enough that telling them apart after the fact may not even be possible.
Bigger models don't obviously fix this. Grounding accuracy scales weakly with model size, and even the best specialist models built for GUI interaction still show substantial failure rates on professional software interfaces. So this looks like a systems-level weakness running across architectures, not a rough edge that better training eventually smooths out.
The defense the researchers propose operates outside the agent's own perceptual loop entirely: dual-channel contrastive classification, which independently checks the visual click target and the agent's stated reasoning against a separate knowledge base, and blocks the action if either channel raises a flag. That's structurally the same move as per-call authorization for text-based agents: a check that sits outside the thing being checked, because asking the agent to police its own perception is asking the confused deputy to confirm its own confusion.
Capability Gating Without Per-Call Authorization
A lot of teams believe they've solved this problem because they've restricted which tools an agent can see. That belief is the gap attackers are currently walking through.
Capability gating is a static decision. It answers one question, once, ahead of time: which tools exist on this agent's menu? Per-call authorization is a dynamic decision, and it answers a different question every single time a tool gets called: is this specific call, with these specific argument values, in this specific session, actually allowed right now? Those are not the same control, and a system with only the first one is not protected against the attacks described above.
A runtime becomes a confused deputy the moment its only checks are whether a tool name exists and whether its arguments match a schema. At that point, the untrusted model is effectively supplying both the proposed action and the fact that the action is authorized, since nothing independent is checking the second part.
A cross-framework audit covering LangChain/LangGraph, LlamaIndex, and the Stripe Agent Toolkit, run against pinned public-source commits, found all three provide capability gating by default. None of them gives you a deterministic, fail-closed, per-call value authorization gate by default. Those frameworks behave exactly as documented; the gap is between what teams assume a default gives them and what the default actually checks.
Schema validation can confirm that an argument is well-formed. It cannot confirm that a given account identifier is a legitimate destination for a payout, because "legitimate" is a policy question that has to be answered by something external to the model making the call. The same audit found that cost-optimized, deployment-tier models try unauthorized calls materially more often than flagship models do, so this exact gap runs widest in the high-volume, lower-cost deployments that make up most of production traffic today.
Least Privilege for Agents at Runtime
Closing the gap means building a check that runs on every tool call, not just once when the tool is installed. ScopeGate's own design lays out five stages that form a template for what a real control looks like: scope, authorization, a money ceiling, idempotency, and a default-deny posture when none of the above can confirm the call is safe. In testing, this design reported zero static bypasses out of 48 attempts, zero unauthorized actions across a 40-iteration adaptive attack run, zero false denials against 10 benign requests, and full containment across 10 out of 10 payment-agent trials.
Least privilege, applied to an agent at runtime, means the agent's credentials are necessary but not sufficient. A call also needs an independent policy decision, made at the moment of the call, using the actual argument values the model produced, checked against a ceiling, a scope, and a record of what's already been done. Static tool lists and schema checks still matter.
Sources
- Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
- Confused Deputy Attacks on Autonomous AI Agents
- Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks
- Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario
- “Living Off the Agent”: AI Agents as Lateral Movement
- Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem


