Data Exfiltration Risk Scoring for LLM-Based Enterprise Products

A framework for scoring data exfiltration risks in enterprise LLM products.

Staff Writer · · 13 min read
Cover illustration for “Data Exfiltration Risk Scoring for LLM-Based Enterprise Products”
Evaluation Criteria · September 30, 2026 · 13 min read · 2,902 words

Data Exfiltration Risk Scoring for LLM-Based Enterprise Products.

Why LLM-based products create a data exfiltration risk that existing frameworks weren't built to measure

Scoring data exfiltration risk in LLM-based enterprise products means evaluating a set of factors that generic risk frameworks were never designed to catch: data access scope, prompt injection exposure, output channel controls, and vendor transparency. This piece builds a factor-by-factor model that TPRM and InfoSec teams can apply consistently across any LLM-based product sitting in a vendor ecosystem.

Start with where these systems actually live. LLMs sit inside email, CRM platforms, IDEs, document management systems, and ticketing queues now, not in some sandboxed pilot program off to the side. Gartner's projection puts real numbers on how fast that happened: 40% of enterprise applications will integrate AI agents by the end of 2026, up from under 5% in 2025 Kellton / AI Governance and Security. That's a sharp jump. That's a cliff.

And most organizations are climbing that cliff without a rope. McKinsey's State of AI survey found 71% of organizations use generative AI in at least one business function, but only 21% report a mature governance model. Do the math on that gap: half or more of enterprise AI deployment is happening with no governance structure sturdy enough to call mature Kellton / AI Governance and Security. That's the backdrop every scoring model has to account for.

Why can't the old tools handle this? Firewalls and endpoint protection were built to catch known signatures and suspicious network traffic, not to interpret a natural-language response or judge whether an LLM just said something it shouldn't have. LLMs process unstructured data, produce probabilistic rather than deterministic outputs, and can reconstruct or infer sensitive information even when that information was never stored anywhere retrievable. A traditional DLP tool watches for a social security number pattern moving across a wire. It has no way to watch for a model quietly inferring someone's salary band from context clues across three separate documents OWASP GenAI LLM Top 10 2026.

CVSS runs into a similar wall. It was built to score known exploits against known software behavior, but LLM findings need to account for probabilistic behavior, blast radius, and the sensitivity classification of whatever content got exposed, not just a signature match. So what fills the gap? That's what the rest of this piece is for: a structured, factor-by-factor scoring model that TPRM and InfoSec teams can run consistently across every LLM-based vendor product in the portfolio.

What the evidence shows about how LLM data exfiltration happens

That number alone should end any debate about whether this is a theoretical risk worth scoring or a live one.

Cyberhaven's 2025 AI Adoption and Risk Report looked at actual usage across 7 million workers and found that 34.8% of corporate data fed into AI tools qualifies as sensitive, more than triple the rate from two years earlier. Source code makes up the largest share at 18.7%, followed by R&D material at 17.1% and sales and marketing data at 10.7%. Source code sitting at the top of that list should catch the attention of any security team that assumed the risk was mostly about customer PII.

Five mechanisms account for how this data actually gets out, and each one demands a different scoring lens. Prompt injection splits into direct and indirect forms: direct injection is a user overriding instructions in their own prompt, while indirect injection hides malicious instructions inside documents, emails, or web content the model reads on its own. Indirect injection is the harder one to catch, because the attack surface is whatever data the model happens to consume.

RAG pipelines introduce their own exposure. Without per-request access controls and sensitivity-label enforcement, a retrieval-augmented system functions as a high-throughput document retrieval tool with no real egress boundary. A USENIX Security 2025 study found that injecting just five poisoned texts per target question into a knowledge base of millions of documents achieved a 90% attack success rate across several benchmark datasets and models. There are five primary exfiltration mechanisms, each requiring a different scoring dimension.

Output channel leakage is the piece traditional security teams tend to miss entirely: data can leave through a rendered Markdown link, an auto-fetched image, or a downstream tool call triggered by the model's own response, none of which look anything like a traditional exfiltration path. The International AI Safety Report 2026 found sophisticated attackers bypass safeguards roughly half the time, a calibration point to consider before scoring any of this. So a control being present on paper tells a scorer almost nothing about whether that control actually holds up under pressure.

The incident record backs this up with specifics. EchoLeak (CVE-2025-32711, CVSS 9.3) triggered zero-click, remote data exfiltration from Microsoft 365 Copilot through a single crafted email, no user interaction required, bypassing the XPIA classifier and exploiting auto-fetched images along with a Teams proxy Top 5 LLM Security Tools for Enterprise AI Applications in 2026. Reprompt (CVE-2026-24307) pulled off single-click exfiltration from Microsoft Copilot through URL parameter injection, again with zero user-entered prompts. GitHub Copilot (CVE-2025-53773, CVSS 7.8) let prompt injection hidden in public repository code comments escalate to arbitrary code execution on developer machines Top AI Security Vulnerabilities to Watch out for in 2026 - Cycode. In August 2025, Cursor IDE users ran into hidden malicious text embedded in public GitHub README files that made the AI assistant execute hidden commands, stealing API keys and SSH credentials and running commands that should have been blocked, all triggered by asking it to clone or set up a repository.

MCP Tool Poisoning showed that even a tool description field inside a Model Context Protocol server can carry malicious instructions that cause an agent to exfiltrate files, and a separate OAuth proxy flaw (CVE-2025-6514, disclosed July 2025) put roughly 437,000 downloads at risk before it got patched scalacode.com. ServiceNow's Now Assist carried its own AI agent vulnerability (CVE-2025-12420, nicknamed "BodySnatcher"), patched October 30, 2025. And in one fintech RAG breach, attackers used reconstruction attacks to reverse-engineer embeddings back into millions of original client investment portfolios, while a related access-control bypass in Pinecone exposed over 200,000 healthcare records csoonline.com.

Line these incidents up and a pattern appears in how they occurred: none of them required a user to click a phishing link or fall for social engineering in the traditional sense. The exploit runs through the model's own processing pipeline. That's the pattern any scoring model has to be built to catch. According to Kellton / AI Governance and Security (cotacapital.com), 45% of enterprises surveyed suffered data leakage involving AI tools, showing that this is not a theoretical risk.

Why the OWASP LLM Top 10 and existing frameworks provide a starting taxonomy but not a score

OWASP's Top 10 for LLM Applications is the most widely cited practitioner framework for identifying and prioritizing LLM security risk. The 2026 edition, published August 4, 2026, ranks Prompt Injection and Sensitive Information Disclosure as co-leading risks, with a blended incident-data and expert-ranking analysis putting them at near-certain top-three probability, 0.99 and 0.95 respectively OWASP GenAI LLM Top 10 2026. Sensitive Information Disclosure sat at #6 in the earlier v1.1 version and rose to #2 by the 2025 edition, a jump that tracks pretty closely with where real-world exploits have been landing. Newer entries like System Prompt Leakage and Vector and Embedding Weaknesses (LLM09:2026) capture threats that only fully matured after the original list came out.

NIST's AI RMF and ISO/IEC 42001 round out the governance side, offering structured approaches to identifying and mitigating AI risk. MITRE ATLAS extends the familiar ATT&CK framework into AI-specific attack techniques, which is genuinely useful for mapping out how an attack might unfold. All of these share one thing in common: they hand a security team categories and controls to check against, not a number. None of them produces a comparable, defensible score that a TPRM analyst can drop into a vendor comparison spreadsheet and track over time.

That's the actual gap. TPRM teams need something they can compare across vendors, track quarter over quarter, and point to when justifying an accept, remediate, or reject decision.

None of this holds still. A point-in-time assessment gets stale fast, since model updates, system prompt changes, and new tool integrations can quietly undo remediations that passed review six months earlier. Any scoring model built from these factors needs to assume re-evaluation on a schedule, not a one-time signoff.

Factor 1: Data access scope (what the product can reach, and whether that access is bounded)

Data access scope acts as the multiplier that determines the severity of every other factor in this model. A product with a real prompt injection weakness but read-only access to low-sensitivity data scores nowhere near as high as a product with that same weakness and write access to unreleased M&A documents. Same vulnerability, wildly different consequence, because consequence is what scope actually measures.

Start by sorting the data itself into tiers: PII, PHI, financial records, source code, legal documents, strategic plans, each one carrying a different exfiltration consequence if it walks out the door. Then look at breadth. Does the product touch one defined, narrow corpus, or does it have simultaneous open access to email, CRM records, document stores, and internal databases? Broader access means a broader blast radius no matter what else is true about the product's controls.

Least-privilege enforcement matters just as much as breadth. Is the model's retrieval scope tied to the permissions of the specific user asking the question, or does it quietly have more reach than any individual employee would ever be granted on their own? In multi-tenant SaaS products, cross-tenant isolation becomes its own line item: has separation between customer environments actually been verified, or is it just assumed?

RAG systems deserve their own close look here, because access control is consistently the hardest problem in RAG deployments. A model that can summarize executive payroll data or unannounced M&A details for any employee who happens to ask the right question is a high-scope product, full stop, regardless of how good its other defenses look on paper.

Shadow AI throws a wrench into all of this scoring. Employees connect unsanctioned AI tools to corporate data sources all the time, often with zero visibility for IT or security teams watching from the outside Kellton / AI Governance and Security. Breaches involving shadow AI run more than $670,000 higher on average than breaches with low or no shadow AI involvement, reflecting the added complexity and regulatory scrutiny that comes with data flows nobody signed off on Kellton / AI Governance and Security.

Factor 2: Prompt injection exposure (how vulnerable the product's input surface is to instruction hijacking)

Prompt injection works because the model has no reliable way to tell the difference between a system instruction and content someone slipped in through a user prompt or an external document. It is at the top of OWASP's risk list, and it's the mechanism behind several of the highest-severity CVEs already covered here.

Score this factor across two distinct surfaces, because they behave differently and fail differently. Direct injection is the input a user types straight into the box. Indirect injection covers everything the model reads on its own initiative: documents, emails, web pages, database records, tool descriptions, MCP server metadata. Indirect injection tends to be the more dangerous of the two in agentic systems, since it can lead to data exfiltration, account takeover, or unauthorized action without the user ever making a mistake.

MCP Tool Poisoning is worth returning to here, because it proved that even a tool description field inside a Model Context Protocol server counts as a genuine injection surface. Any scoring exercise needs to ask a vendor directly whether MCP tool descriptions get validated and sandboxed before an agent ever touches them.

Agentic capability changes the math on all of this. A product where the LLM can call tools, fire off requests, or kick off workflows on its own multiplies the blast radius of any successful injection, so agentic products should score higher on this factor by default. That's just an honest accounting of what a successful attack can now reach.

When gathering evidence from a vendor, ask whether they run a cross-prompt injection attack classifier, and don't stop there. EchoLeak bypassed Microsoft's own XPIA classifier, which is the clearest possible proof that a control existing on paper isn't the same thing as a control that actually holds. Ask whether output channels are restricted to block Markdown-rendered link injection or auto-fetched image exfiltration, ask about red-team testing cadence, and ask whether tool access in agentic deployments runs through per-tool allow-lists. Weight the final sub-score by injection surface breadth multiplied by agentic capability multiplied by whatever evidence exists of classifier effectiveness, and treat a document-ingesting or email-connected product with no indirect injection defenses as a high-risk finding no matter what other controls it happens to have.

Factor 3: Output channel controls (whether data can leave through the model's responses)

Exfiltration in these systems rarely needs a network-layer breach at all. It happens through the model's own output, a rendered response, a Markdown link, an auto-fetched image, a downstream tool call, an API response that gets piped straight into another system.

Start with the rendering environment itself. Does the client render Markdown or HTML, and does it execute links automatically? EchoLeak is the case study here: it exploited reference-style Markdown formatting and auto-fetched images to exfiltrate data with zero clicks from the user. That's a documented, CVSS 9.3 incident against one of the most widely deployed Copilot products on the market Top 5 LLM Security Tools for Enterprise AI Applications in 2026.

Downstream tool calls deserve equal scrutiny. If a product's output feeds automatically into another system, whether that's a ticketing tool, a CRM update, or a workflow trigger, each one of those handoffs is its own potential exfiltration channel. Ask whether outbound API responses get inspected for sensitive content before delivery, or whether that inspection simply doesn't exist. Ask, too, whether the product writes summaries, logs, or embeddings out to storage sitting outside the enterprise boundary.

On the content side, look at whether PII and sensitive-data redaction gets enforced at the gateway layer or gets left entirely to the application built on top of the model. That distinction affects how much a control actually catches: gateway-layer enforcement catches every request that passes through, while application-layer enforcement only catches what that one application's developers thought to build. Check for output validation against sensitivity labels, meaning does the product verify that a response wouldn't expose content the specific querying user isn't cleared to see. Rate limiting and anomaly detection on output volume matter too, since bulk data extraction through repeated queries is detectable once a baseline exists to compare against. And immutable audit logs of model outputs aren't optional if there's ever a need to reconstruct what happened after an incident.

A residual risk remains even when every content control above is working: side-channel leakage. Response timing, confidence scores, and output structure can all leak information on their own, bypassing content-level controls entirely. Ask a vendor directly whether side-channel disclosure is even part of their threat model, because a lot of them haven't gotten there yet.

Factor 4: RAG pipeline integrity (whether the retrieval layer introduces its own exfiltration surface)

A RAG pipeline is, at the same time, a high-throughput document retrieval system and a summarization layer sitting on top of an LLM. Strip away the controls and that description doubles as an accurate description of a data exfiltration tool. That's a statement about what happens when RAG is built as an architecture without safeguards, or when the retrieval layer gets built without access boundaries baked in from the start.

OWASP formally recognized this with LLM08:2025, Vector and Embedding Weaknesses, which calls out vector databases and embedding pipelines as their own distinct vulnerability class, separate from prompt injection and separate from whatever the base model does on its own. A product can pass every prompt injection test in the book and still carry serious exfiltration risk if its retrieval layer has no per-request access enforcement, which affects how the product should be scored.

The fintech and healthcare incidents already covered here make the stakes concrete rather than theoretical. Reconstruction attacks reverse-engineered embeddings back into millions of original client investment portfolios in one case, and a Pinecone access-control bypass exposed over 200,000 healthcare records in another csoonline.com. Neither of those required breaking the model itself. Both required nothing more than a retrieval layer that couldn't tell one user's clearance level from another's.

Scoring this factor means asking pointed, specific questions rather than accepting a vendor's general assurance that "access is controlled." Does the vector database enforce access control per request and per user, or does a query return results pulled from across the entire knowledge base regardless of who's asking? Are sensitivity labels attached to embeddings themselves, or only to the source documents sitting somewhere upstream of them? And does the poisoning risk demonstrated in the USENIX study, where five poisoned documents hit a 90% attack success rate against a corpus of millions, get addressed anywhere in the vendor's threat model, or does it simply go unmentioned?

That's the whole point of building the model this way. Not a checklist to satisfy once, but a live measurement that moves as fast as the products it's scoring. There are two primary RAG.

Sources

  1. Top 5 LLM Security Tools for Enterprise AI Applications in 2026
  2. AI Governance and Security: A Quick Guide on LLM Data Leaks Protection
  3. Top AI Security Vulnerabilities to Watch out for in 2026 - Cycode
  4. Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus
  5. What Is Prompt Injection? Attacks, Types, & Prevention

More in Evaluation Criteria