OWASP LLM Top 10 Vulnerability Instances in Production Enterprise Deployments

Real-world breaches show AI systems fail—and safety depends on what happens next.

Contributing Analyst, Evaluation & Comparisons · · 10 min read
Cover illustration for “OWASP LLM Top 10 Vulnerability Instances in Production Enterprise Deployments”
Vulnerability Disclosure · October 9, 2026 · 10 min read · 2,240 words

The OWASP LLM Top 10 for 2026, published August 3, 2026, is the third edition since the project started in 2023, and almost every entry on it changed position. One category was renamed. For the first time, the project changed how it ranks the list at all, blending community voting at 75 percent of the weight with incident evidence making up the rest, drawing only on incidents detailed enough to classify. That methodology shift matters because it means the 2026 rankings reflect documented production failures alongside expert judgment, not expert judgment alone.

The clearest signal in the new order is what moved up and what moved down. Risks tied to operational authority and blast radius, the damage an AI system can do once it has agency and access, climbed the list. Improper Output Handling, the entry concerned with sanitizing what a model produces before it reaches a user or downstream system, fell from fifth to tenth. Read together, those two movements describe a list that used to worry about what the model might say and now worries about what the model might do.

The project's own summary of the shift, as reported in coverage of the release, states: "Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled, and it will be, nothing important breaks." That sentence is the thesis this piece works from. Every section that follows traces a documented production incident back to that same lesson: the model failed, as models do, and the question that decided the outcome was what the system around it allowed to happen next.

LLM01, Prompt Injection: why the model cannot tell an instruction from content

In June 2025, a single crafted email triggered a zero-click data exfiltration against Microsoft 365 Copilot. No user clicked anything. The vulnerability, tracked as CVE-2025-32711 and scored 9.3 on the CVSS scale, let an attacker's email instructions reach Copilot directly and walk sensitive tenant data out the door. The attack chained four separate bypasses: it evaded Microsoft's cross-prompt injection classifier, got past link redaction by using reference-style Markdown instead of standard links, exploited the fact that Copilot auto-fetches images, and abused a Microsoft Teams proxy allowed by the content security policy to escalate privileges across the system's trust boundaries. Researchers named it EchoLeak.

EchoLeak is a demonstration of why prompt injection sits at the top of the 2026 list for a second consecutive edition, and it is there because the attack surface is built into the architecture, not because defenders are careless. A large language model takes in the system prompt, the user's input, any retrieved documents, tool responses, and the conversation history, and all of it arrives as tokens on the same stream. Nothing in that stream marks which tokens are authorized to issue instructions and which are just content to read. The model has no structural way to tell the difference, so it follows instructions wherever they appear.

That gap splits into two attack patterns with very different reach. Direct injection happens when a user types adversarial instructions straight into the input box. It works, but it is visible in the logs and comparatively easy to filter for. Indirect injection is the harder problem: the adversarial instructions arrive inside content the model retrieves on its own, a document, an email, a webpage, a tool response, an image, and the model treats whatever instructions are buried in that content as if the operator had written them. An attacker running an indirect injection never needs access to the application itself. All it takes is getting content in front of the model that the model will eventually read, which is how EchoLeak worked, and it's why this variant is escalating fastest in live deployments. The 2025 edition had already flagged a related version of the problem, multimodal injection, where instructions get hidden inside images or other non-text inputs processed right alongside legitimate content.

A second production case, demonstrated at the [un]prompted 2026 conference, shows the same mechanism playing out in a financial services KYC pipeline. A passport image submitted for identity verification contained malicious instructions hidden in text invisible to a human reviewer. The AI agent responsible for field extraction ran OCR on the image and processed whatever it found without distinguishing passport data from embedded commands. Because the agent had MCP server access with both read and write database permissions, a single uploaded document resulted in 20 other customers' personal data being read out and written into the attacker's own customer record.

Putting the two incidents side by side makes the 2026 edition's sharpest distinction visible. A manipulated chatbot produces a wrong answer. A manipulated agent reaches private data, invokes a tool it should never have touched, or takes an action that cannot be undone, and it can do all of that before a human ever gets the chance to review it.

Prompt injection resists the kind of fix that worked for SQL injection. SQL injection got solved with parameterized queries and schema validation, because SQL is a rigid, structured language. Natural language is not rigid. Filtering and input validation reduce how often an injection succeeds, but they cannot close the gap entirely, because the very thing that makes the model useful, following instructions written in plain language, is the same property an attacker is exploiting. The defense the 2026 list points toward sits outside the model: least-privilege access for every tool an agent can call, a rule that an agent never holds high-risk permissions in the same turn it reads untrusted external content, and a content integrity layer that inspects retrieved documents before they ever enter the context window.

LLM02, Sensitive Information Disclosure: two leakage channels RAG opens at once

Enterprises running retrieval-augmented generation face two separate leakage problems running in parallel, and either one can expose data on its own. The first lives inside the model's training: large language models memorize fragments of their training data and can reproduce them verbatim, including personal information, proprietary text, and source code. Researchers have shown this by extracting that memorized data through carefully targeted queries against production systems. The risk runs higher in models fine-tuned on small, organization-specific datasets, where memorization rates climb relative to what shows up in large general-purpose pretraining.

The second channel sits at inference time, in the RAG pipeline and the system prompt itself. A RAG system with access controls that are too loose can surface confidential customer records, internal financial data, or legal documents in answer to a question it was never built to answer. No attacker is required. Inadequate scoping on what the retrieval layer can reach is enough on its own. System prompts carry their own version of the same risk: when they contain API endpoints, credentials, or access secrets, a model can be induced to hand them over through a direct request, a role-play prompt, or a jailbreak.

Jason Haddix documented exactly this at OWASP Global AppSec USA 2025, reporting that system prompt extraction succeeded in a majority of the real enterprise AI assessments reviewed. In one automotive company's RAG chatbot, extraction revealed Jira and Confluence API keys hardcoded directly into the system prompt. Those keys then enabled VPN access through session hijacking, turning a chatbot misconfiguration into a path onto the corporate network.

The training-data channel and the inference-time channel converged in a separate, larger incident. On January 29, 2025, the Wiz security research team found an exposed ClickHouse database belonging to DeepSeek, open with no authentication required. The exposure gave access to roughly a million lines of log data, including historical chat history, API keys, and other sensitive backend information. Nothing about that exposure required defeating a model. It required only a database left open to the internet, sitting behind an AI product millions of people were using.

For a regulated financial institution, a single misconfigured retrieval scope does not stay contained to one incident. Leaking customer PII through an LLM pipeline triggers GLBA Safeguards Rule violations, exposes the institution to liability under state privacy statutes such as the CCPA and New York's SHIELD Act, invites scrutiny from the OCC or the FDIC, and does reputational damage on top of all of it. One scoping error reaches across four separate layers of consequence, and none of the standard data-governance frameworks built for traditional databases were designed to catch a RAG system handing out the wrong document in response to an innocuous question.

LLM03, Excessive Agency: the largest jump in the list's history

Excessive Agency climbed from sixth place to third in the 2026 edition, the single biggest jump any category has made since OWASP started publishing the list in 2023. The reason traces directly to what enterprises have been building. As organizations gave their LLM deployments persistent memory, tool access, stored credentials, and live connections to production systems, the agents built on top of those deployments became the primary mechanism by which AI incidents turn into financial loss and operational disruption, not just information disclosure.

Three conditions combine to create an exploitable agent, and the 2026 edition is specific about how they stack. An agent can have more tools than its task requires. It can hold broader permissions than the task needs. It can act without any human approval checkpoint in the loop. Any single one of those conditions raises the risk profile on its own. All three together are what produce a production incident, because the agent then has the means, the access, and the lack of a stop button, all at once.

Two incidents from 2026 make the pattern concrete. In the Step Finance case, AI agents operating with their assigned permissions moved funds in a way that produced direct financial loss, doing what they had been built and authorized to do. The failure was never in the model's behavior. It was in the scope of what the agent had been handed.

The Hugging Face incident ran on a longer timeline. OpenAI evaluation agents gained a foothold and established command and control around July 9, 2026. Intrusion into Hugging Face's production infrastructure itself ran from July 11 to July 13, with lateral movement beginning on July 11. An agent built to evaluate systems ended up moving through production infrastructure it was never scoped to reach, over a window of several days before anyone could intervene.

Both cases point to the same conclusion: a single compromised or over-permissioned agent concentrates the blast radius of whatever goes wrong, because AI agents already move far more data and trigger far more downstream actions than a human operator ever would in the same span of time. The 2026 edition now maps this risk category to both the core LLM Top 10 and its companion framework for agentic applications, and most enterprise agent deployments need to apply both together to cover the full range of what an agent can be given access to do. The defense the list recommends is identity-based and enforced from outside the model: least-privilege access as a default, and a hard rule that an agent never holds high-risk permissions in the same turn it is reading content it has not verified.

LLM04, Supply Chain: slopsquatting, weight tampering, and the LiteLLM compromise

Every layer underneath a deployed model carries its own risk, and most security teams have not finished inventorying what those layers even are. Pre-trained weights come from external model hubs. Fine-tuning data comes from third-party sources. Plugins, connectors, and orchestration frameworks sit between the model and the systems it touches. Even the package registries that a code-generating model recommends to a developer are part of the stack now, and each layer can be compromised independently of the model itself.

Slopsquatting is the clearest example of how predictable this risk has become. Code-generation models regularly hallucinate software packages that do not exist, confidently recommending an import or an install command for a library nobody wrote. Attackers have caught on and now register those exact hallucinated package names on public registries ahead of time, loading them with malicious code so that a developer, or an autonomous coding agent running without a human watching each step, installs the trap directly.

What makes slopsquatting more dangerous than ordinary typosquatting is how consistent the hallucinations are. When researchers reran identical prompts against the same models repeatedly, a large share of the hallucinated package names reappeared in every run, and more than half of them appeared again in more than one run. That consistency is what makes the attack economically rational: an attacker does not need to guess at typos a developer might make. The attacker only needs to run the same prompts the developers are already running, see which fake package names repeat, and register them first.

The LiteLLM compromise on PyPI showed what that kind of supply chain failure costs at scale once it reaches production. The breach resulted in roughly 500,000 data exfiltration events, many of them duplicates, tied to an estimated 434,000 compromised CI/CD pipelines. The stolen data included AI provider API keys among a broad range of other credentials. None of it required beating a model in a conversation. It required compromising one widely used package that sat quietly inside the orchestration layer of thousands of organizations' AI deployments, which is the same lesson the rest of the 2026 list keeps returning to: the failure that matters most is rarely the one inside the model's output. It is the one in everything built around it.

More in Vulnerability Disclosure