Disclosed Data Exfiltration Vulnerabilities in RAG-Based Enterprise Products

Researchers discover how attackers exfiltrate data from enterprise AI without user interaction.

Staff Writer, Product & Vulnerability Coverage · · 10 min read
Cover illustration for “Disclosed Data Exfiltration Vulnerabilities in RAG-Based Enterprise Products”
Vulnerability Disclosure · October 6, 2026 · 10 min read · 2,297 words

The core vulnerability in RAG-based enterprise products is built into the architecture itself. Retrieval-augmented generation connects a live knowledge base, internal wikis, CRM records, SharePoint, code repositories, to a language model that treats retrieved content with the same weight it gives system instructions. The model treats a legitimate directive and an adversarial one buried inside a document it just pulled back as the same kind of instruction. Untrusted content and privileged data end up in the same context window, with nothing to keep them apart.

That single design choice is why standard security tools miss the problem. A browser-based DLP tool can see that someone opened an AI application. It cannot see what the retrieval pipeline pulled in, whether a document carried hidden instructions, or whether the model's response quietly reassembled sensitive content inside an otherwise normal-looking answer. Traditional DLP and CASB tools work at the network and endpoint layer, watching traffic and file movement, not the internal logic of a retrieval call. Enterprise AI risk platforms, Promptarmor among them, were built to look at this problem from a different angle: assessing the structural risk created when a retrieval pipeline wires privileged data sources straight into a language model, and surfacing the parts of a vendor's AI attack surface that legacy frameworks were never designed to see. RAG has become the standard way enterprises stand up knowledge-aware AI tools, partly because it updates automatically as source documents change, with no retraining required. That convenience is what makes the risk so hard to separate from the product. The attack surface runs across four layers, ingestion, vector storage, retrieval, and generation, and a weakness at any one of them undermines the security of the other three.

Indirect prompt injection as an attack channel through the retrieval pipeline

Indirect prompt injection does the most damage to RAG systems because it turns the system's own retrieval function into the weapon. An attacker plants instructions inside a document. The RAG system retrieves that document because it matches a user's query. The model then follows the embedded instructions, believing they came from a trusted source, because nothing in the architecture tells it otherwise.

No network access is needed. No stolen credentials, no phishing click, no user action of any kind. A crafted document only has to enter the corpus and get retrieved in response to a query that happens to match it semantically. Once those instructions sit inside the model's context, they can be exfiltrated in two ways. One is response-level: the model can be told to fold retrieved sensitive content into its output, or to direct an agentic workflow to send data to an outside server. The other is embedding-level: in some architectures, embedding inversion techniques can reconstruct the original text from vectors stored in the vector database. OWASP's LLM08:2025 category formally recognizes that embedding-inversion techniques can recover a significant share of source information by exploiting weaknesses in how embeddings are built.

Academic work has already mapped this out in formal terms. The EDEA framework, short for External Data Extraction Attacks, breaks the adversarial query into three parts: an extraction instruction, a jailbreak operator, and a retrieval trigger, and shows that earlier, more ad hoc attacks are really just instances of this same structure. That formalization followed earlier research, "Follow My Instruction and Spill the Beans," which demonstrated that data extraction from RAG systems could be done at scale, setting a feasibility baseline that later vulnerabilities in deployed products would go on to confirm.

Indirect prompt injection succeeds because it exploits the trust boundary, or rather the missing one, between retrieval and generation: a model cannot distinguish attacker-controlled instructions embedded in retrieved content from legitimate system directives. So if you want to catch this kind of AI-specific exposure, you need to keep watching how a vendor's AI systems behave in production and how untrusted content flows into trusted contexts, before the gap becomes a disclosed vulnerability. Most enterprise RAG deployments compound the problem by running retrieval under a single service account with broad access across the repository and no per-user authorization check at that layer. Every query touches the entire document corpus, including files the person asking the question could never open through any other route. What had been a theoretical mechanism turned into a confirmed, zero-click exfiltration path inside a Microsoft product.

CVE-2025-32711 (EchoLeak): how indirect prompt injection became a confirmed zero-click exfiltration path in Microsoft 365 Copilot

EchoLeak is the first documented case where indirect prompt injection was used for actual data theft inside a production enterprise AI system, and no user had to click anything. The vulnerability, tracked as CVE-2025-32711, was disclosed by Aim Security researchers Pavan Reddy and Aditya Sanjay Gujral and rated critical by Microsoft at a CVSS score of 9.3. It reached Copilot integrations across Word, Excel, PowerPoint, Outlook, and Teams, touching nearly every surface where Microsoft 365 users interact with the assistant day to day.

The attack began with a crafted email that carried hidden prompt injection instructions. Copilot retrieved that email as part of its normal RAG context, the same way it would retrieve any other relevant document, and in doing so it executed instructions written by the attacker. From there, the payload reached chat logs, OneDrive files, SharePoint content, and Teams messages, all pulled out to a server the attacker controlled. None of this required the victim to open the email, click a link, or approve anything. It ran as a zero-click attack from start to finish.

What made EchoLeak technically remarkable was the chain of defenses it slipped past, one after another. It bypassed Microsoft's XPIA classifier, the system built specifically to catch cross-prompt injection attempts. It got around link redaction. It then cleared Content Security Policy because it routed through an allowlisted Teams image proxy, an already-trusted domain that ended up carrying the stolen data to the attacker's server on Copilot's behalf. A Microsoft service, one explicitly permitted by the CSP, completed the exfiltration step for the attacker. All four of those steps happened without a single moment of user interaction.

Aim Security disclosed the issue privately to Microsoft in January 2025. Microsoft developed and deployed server-side fixes by May 2025, and customers did not need to take any action to be protected. There is no evidence the flaw was exploited before the patch went out. That timeline is a genuine point in Microsoft's favor: a fix rolled out entirely on the server side, with nothing for IT teams to configure or push. It also means that for the months between disclosure and patch, enterprise customers had no way to see the exposure themselves and no tool that would have flagged it.

EchoLeak is not a one-off quirk of how Microsoft built Copilot. The specific bypass techniques, the XPIA classifier, the image proxy, the particular CSP allowlist, belong to Microsoft's implementation. The underlying condition they exploited belongs to the architecture of any LLM-based assistant with broad reach into internal data. Any product built the same way, connecting a model to a wide pool of privileged data through retrieval, carries some version of the same exposure.

CVE-2025-53773: the same injection class reaching GitHub Copilot and enabling remote code execution

The pattern did not stop at Microsoft 365. CVE-2025-53773, rated 7.8 on the CVSS scale, showed the same injection class reaching GitHub Copilot, where hidden instructions embedded in source code files, code comments, GitHub issues, web pages, or other repository files could trigger the attack. Copilot read these files as part of its normal workflow, the same way Microsoft 365 Copilot read the crafted email in EchoLeak, and the hidden instructions executed as if they came from a trusted source.

What changed here was not the mechanism but the outcome. EchoLeak exfiltrated data. This vulnerability enabled remote code execution. What the underlying agent is allowed to do sets the outcome: an assistant that can read and summarize documents can leak data, and an assistant that can write and execute code can be made to run arbitrary commands. The blast radius of an injection attack scales with the permissions of the system it reaches, not with the sophistication of the injection itself. The same vulnerability class, confirmed twice now in two unrelated products built by two different engineering teams, points to a shared architectural condition that both engineering teams independently inherited.

CVE-2025-69286: how insecure infrastructure primitives in RAGFlow create equivalent exfiltration exposure below the prompt layer

Some exfiltration paths bypass the model and reach the infrastructure beneath the pipeline directly. CVE-2025-69286, found in RAGFlow, the open-source RAG engine built by Infiniflow, carries a CVSS score of 9.8 and shows the exposure can sit one layer down, in the infrastructure that supports the pipeline.

The root cause was a weak key generation scheme. Both the API key and the beta token were generated using the same URLSafeTimedSerializer, fed with predictable inputs, which meant one could be derived from the other. An attacker who got hold of a shared assistant or agent URL could work backward to the personal API key tied to it, and from there gain full control over that account's assistant and agent, no model interaction required at any point.

This matters for how defenders think about the overall risk. A team that spends its energy on prompt-layer defenses, injection detection, output filtering, careful handling of retrieved text, can still be fully exposed through a completely different door if the token generation underneath the application produces weak, derivable credentials. The exfiltration surface in a RAG system does not stop at the model's context window. Academic researchers have since taken this same pattern, injection in production software on one side, infrastructure weakness on the other, and generalized it into frameworks that test how far extraction attacks can actually reach.

The SECRET framework and academic extraction research on the attack class's true ceiling

EchoLeak proved the mechanism could work against one real product. Academic research since then has shown it generalizes well beyond that single case, to the point where researchers can extract knowledge base content verbatim from RAG systems that already have defenses built in.

The SECRET attack, short for Scalable and EffeCtive exteRnal data Extraction aTtack, is the first comprehensive attempt to formalize External Data Extraction Attacks against retrieval-augmented language models. It breaks the design of an adversarial query into the same three parts named earlier: an extraction instruction, a jailbreak operator, and a retrieval trigger. SECRET adds an adaptive layer on top of that structure: it uses language models themselves as optimizers to generate specialized jailbreak prompts, paired with a technique called cluster-focused triggering, which alternates between broad exploration and narrow, targeted refinement so the retrieval triggers actually work at scale.

Tested against four models, including three leading commercial systems, SECRET succeeded against all sixteen RAG instances researchers evaluated. It pulled a meaningful share of data out of systems running defended commercial models, where earlier extraction attempts had come away with nothing. That result builds on the earlier Harvard and Carnegie Mellon study, "Follow My Instruction and Spill the Beans," which first showed that an instruction-following model could be pushed into disclosing retrieved content word for word, at scale.

The finding that matters most for anyone defending one of these systems: standard detection does not catch well-built extraction attacks. SECRET's triggers read as coherent, ordinary natural language rather than the garbled token sequences that anomaly detectors are built to flag. Sentence-level similarity detectors caught none of it, so you get a 0% detection rate across every model tested, and the attack held up against perplexity-based detection too. If the attack text looks completely normal, filtering out what looks strange will not work. Patching EchoLeak closed one bypass chain in one product. It did not close the attack class, because the attack class does not depend on any single implementation detail that a patch can remove.

Knowledge base poisoning as a parallel exfiltration enabler

Every mechanism covered so far assumes the attacker's payload has already made it into the corpus, whether through a crafted email, a poisoned code file, or a manipulated document in SharePoint. That assumption points to a layer of the problem that sits even earlier than retrieval: ingestion. If a RAG system pulls content from wikis, ticketing systems, shared drives, or external repositories that multiple people can edit, every one of those entry points is a place an attacker can plant instructions long before any query ever touches them.

Enterprises often try to handle this with content sanitization: scanning documents for suspicious patterns before they ever reach the vector store. That approach runs into the same wall the SECRET research exposes at the detection stage. A sanitization filter built to catch obviously malicious text will miss injected instructions written in plain, coherent language that reads like any other internal document. Scale makes the problem worse. A knowledge base with a handful of curated documents can be reviewed by hand. One with tens of thousands of files pulled continuously from wikis, tickets, and shared repositories cannot be checked that way, and automated filters face the same blind spot that sentence-level similarity detectors hit against SECRET's triggers: a 0% catch rate against content built to look ordinary.

The pattern running through EchoLeak, the GitHub Copilot vulnerability, RAGFlow's token weakness, and the SECRET research is the same pattern from four different angles. Untrusted content and privileged access sit inside the same pipeline, with no real boundary separating them, and that weak boundary appears in whichever layer happens to be weakest in a given product, prompt handling, infrastructure, or the corpus itself. Closing one path does not close the others. Each one needs its own scrutiny, because the architecture that makes RAG so useful for enterprise knowledge work is the same architecture that keeps generating new ways for that knowledge to leak out.

More in Vulnerability Disclosure