System Prompt Confidentiality Leakage Disclosures in Commercial LLM Products

Researchers document how attackers extract secret instructions from AI chatbots.

Staff Writer · · 11 min read
Cover illustration for “System Prompt Confidentiality Leakage Disclosures in Commercial LLM Products”
Vulnerability Disclosure · October 4, 2026 · 11 min read · 2,395 words

A system prompt is supposed to be a set of instructions. In commercial LLM deployments, it has quietly become something else: a repository of API keys, business logic, role permissions, and the specific configuration that turns a generic foundation model into a branded product. That shift is what makes system prompt confidentiality leakage a distinct, documented vulnerability class: it's not just another flavor of generic AI risk.

System Prompts as a Sensitive Data Store

Start with how a transformer model actually reads its input. The system prompt sits at the top of the context window. User input gets appended below it. Retrieved documents, if the system uses retrieval, get appended further down still. The model then generates its next token based on everything in that stream, treating it as one continuous sequence. There is no internal flag that marks the first chunk as "trusted instructions from the developer" and the rest as "untrusted input from a stranger." The architecture carries no concept of privilege.

This is a property of how current transformer-based models work: no vendor has shipped a model that enforces trust separation at the architectural level. Every mitigation built on top of that has to compensate for an absence the model itself cannot fill.

As companies built specialized products on top of foundation models, the system prompt became the place where commercial pressure concentrates, since the system prompt is where competitive differentiation lives. The instructions are what make a general-purpose model behave like a customer service agent for one company, or a coding assistant tuned to one codebase, or a compliance-aware assistant for one regulated industry. That's valuable intellectual property, and it sits in plain text, inside a context window the model cannot distinguish from the rest of the conversation. Expensive, proprietary logic sits in the one spot the model treats as unprivileged, so the incentive to extract it follows immediately.

OWASP's guidance under LLM07 names what tends to end up there: API keys and authentication tokens, internal system architecture details, proprietary business rules and decision logic, role-based permissions, and content moderation rules. Every one of those categories, if exposed, gives an attacker either a credential to abuse or a blueprint to exploit.

OWASP treats this as serious enough to warrant its own category. The broader problem of sensitive information disclosure, things like PII, credentials, and training data, falls under LLM02:2025. System prompt leakage gets its own separate listing, LLM07:2025, in the 2025 Top 10 for LLM Applications. OWASP's underlying design principle is direct: a system prompt should never be treated as a secret, and it should never be used as a security control on its own. Teams invite the vulnerability when they build as if the opposite were true, loading sensitive material into the prompt and trusting the model to keep it there.

How extraction attacks work

The simplest attack is also the one defenders most often underestimate: just ask the model to repeat its instructions. This works more often than it should, because unless an external layer specifically treats system instructions as privileged, the model has no built-in reason to refuse. It sees the request as just another instruction, no different in kind from any other line of user text.

Variants on direct extraction get more creative without getting more complicated. One line of research tested whether models that refuse a plain request for their system prompt would comply if the request were reframed as an encoding or formatting task, asking for the prompt rendered as JSON, say, or base64. Across seven models and 46 verified system instructions, that reframing worked often, with success rates of 0.7 or above, in cases where the same models had refused the direct version of the request. The instructions didn't change. Only the framing did.

Multi-turn attacks go further, and they're harder to defend against because they exploit something closer to a personality trait than a technical gap: the tendency of LLMs to agree with users across a conversation, sometimes called sycophancy. Salesforce AI Research studied prompt leakage across ten closed- and open-source models, and from that work it built a two-turn threat model around this behavior. In the first turn, the adversary submits an ordinary domain question paired with an attack prompt. In the second turn, a challenger utterance pushes on the model's instinct to validate what the user is saying, and that pressure is enough to complete the extraction. The average attack success rate went from 17.7% in single-turn attempts to 86.2% with the two-turn approach, and several frontier models in the test set came close to total leakage.

The same research found that no single defense strategy was sufficient against this threat model on its own. Stacking multiple black-box defenses cut the attack success rate substantially, but if you tested any one control alone, real exposure remained.

Agentic systems widen the blast radius further. The Salesforce study points out that in agent-based deployments, prompt leakage can expose backend API calls, implementation details, and system architecture, not just the prompt text itself. Separate analysis of LLM security risk heading into 2026 makes a related point about where injected instructions can originate: any integration that feeds context to a model, documents, wiki pages, emails, images, is a potential entry point, because the model has no inherent way to tell a legitimate instruction from a command disguised as ordinary data.

A newer line of attack bypasses the extraction request. A method called SPORE, published in April 2026, needs no training at all; it pulls private information straight out of an LLM agent's memory during inference, not from the system prompt text. It works in both black-box and gray-box settings. In the black-box version, a single query is enough. In the gray-box version, the attacker reads the top-k predictions behind each generated token, the ranked outputs that inference APIs expose, and so pulls out private information one token at a time. In tests against frontier models, this bypassed existing detection and safety alignment mechanisms. SPORE is worth tracking for a different reason than being a cleverer version of prompt injection. It targets a different thing entirely: whatever the model has absorbed into its working memory during a session. The attack surface now extends well past the prompt itself.

Diagram: Two-Turn Attack vs. Single-Turn: The Sycophancy Gap. Visualizes: Show the dramatic jump in prompt extraction success rates between single-turn and two-turn adversarial attacks, based on Salesforce AI Research data across ten models.

What documented disclosures reveal about the real consequences

None of this stays theoretical for long. Documented disclosures show the harm landing in three overlapping categories: exposed intellectual property, secondary attacks enabled by what leaked, and operational disruption that cascades outward from a single incident.

Samsung's case is the clearest template for IP loss through plain carelessness rather than a sophisticated attack. Engineers pasted confidential source code into ChatGPT while they used it for ordinary work, so the code left the company's control the moment it entered the prompt. The fallout was immediate: Samsung banned internal use of AI tools, and major Wall Street banks, including JPMorgan and Goldman Sachs, restricted ChatGPT over concerns that employees could leak sensitive information the same way. The incident set a template that has held up since: the cost of one data leak, one regulatory penalty, or one IP exposure event can outweigh the combined productivity gains of letting an entire workforce use the tool unsupervised. So even when the ban costs real productivity, that math makes it the rational short-term call.

EchoLeak (CVE-2025-32711) matters for a different reason. It's the first documented case where prompt injection was weaponized for actual data exfiltration in a production AI system, not demonstrated in a lab. The attack against Microsoft 365 Copilot chained four separate bypasses together: it evaded Microsoft's cross-prompt injection attack classifier, got around link redaction by using reference-style Markdown, exploited images that the system auto-fetched, and abused a Microsoft Teams proxy that the content security policy happened to allow. Microsoft patched the vulnerability server-side, and it confirmed that no exploitation happened in the wild. The chain itself is the lesson: four individually reasonable defenses, a classifier, link redaction, CSP rules, failed collectively against an attacker who studied the whole stack rather than testing one control at a time.

Regulatory exposure compounds both patterns. Analysis of LLM security incidents through 2025 lists regulatory violations under GDPR and HIPAA as a direct consequence of sensitive information disclosure, alongside financial losses and damaged user trust. In legal, healthcare, and financial services deployments, PII exposure and system architecture disclosure happening together create liability that's larger than either one alone, since a leaked prompt in those sectors often reveals both what data the system touches and how it decides what to do with it.

Why enterprise deployments of vendor AI face a version-control blind spot

Enterprises running commercial AI products carry an extra exposure that gets far less attention than the attacks themselves: the behavioral contract baked into a vendor's product can change with no formal notice, and the customer has almost no way to see it happen.

Model weights get versioned, documented, and announced when a vendor updates them. System-level instructions, the prompts that actually govern how the product behaves day to day, rarely get the same treatment. A public repository that tracks system prompts across commercial LLM products shows regular updates to that material, so the instructions that shape a vendor's product can shift outside any enterprise customer's own change management process. If those instructions define role-based permissions, content moderation boundaries, or capability limits that a security team has already reviewed and approved, a silent vendor-side update can invalidate that review, and no compliance alert on the customer's end will ever trip.

The stakes of that blind spot have grown because of where these models now operate. LLM security analysis heading into 2026 describes AI systems embedded inside IDEs, CRMs, ticketing systems, collaboration platforms, and office suites, reading the same inboxes, wikis, and databases employees touch every day, and in some deployments sitting one step away from systems that move money, change access privileges, or handle regulated data. A leaked prompt in that kind of environment exposes an entry point into an organization's entire data graph, not just one chatbot's instructions.

Vendor severity ratings add a false sense of reassurance on top of this. Researchers who have responsibly disclosed prompt leakage issues to major platforms report that some vendors, Alibaba and Baidu among them, rated the findings as medium severity, while other platforms would not assign any formal severity rating. Many security researchers think those ratings sit too low, because the leakage exposes far more business logic than that. If an enterprise relies solely on a vendor's own severity label and patch notes, it will tend to underweight a risk the vendor itself hasn't fully priced in.

So enterprises that deploy commercial LLM products trust an architecture that cannot protect its own instruction set, and a vendor relationship that gives them no reliable alert when those instructions change. Platforms built specifically for AI vendor risk assessment, Promptarmor among them, track which LLM vendors embed high-risk material in system prompts and watch for architectural changes that could raise extraction risk, so you can catch these shifts across a deployment before an attacker finds them first. To close that gap, you need an independent visibility layer, something that watches vendor-side changes on its own schedule rather than waiting on the vendor to say something.

What the current mitigation landscape can and cannot do

Current defenses cut the risk down. None of them close it, because every one operates at the application layer, patching around a gap that originates in the model's architecture itself.

Design-time practices are the necessary starting point. OWASP's LLM07 guidance tells teams to write every system prompt as though it will eventually be read by an attacker: strip out credentials, secrets, and authorization logic entirely, and store that material in external systems the model can call at runtime instead of embedding it in prompt text. OWASP's LLM02:2025 guidance also calls for data sanitization, strict input validation, least-privilege access controls on data sources, and explicit restrictions inside the system prompt on what kinds of data the model can return. These steps matter, and OWASP's own documentation is clear that they aren't sufficient by themselves: prompt-level restrictions can be ignored or bypassed through prompt injection, so a determined attacker can route around them.

Research on stacking defenses gives a more exact picture of how much ground is actually recoverable. The Salesforce AI Research study tested seven black-box defense strategies on their own and found that Query-Rewriting worked best against the first-turn attack, while Instruction defense worked best against the turn-two sycophancy exploit. Combining every strategy together brought the average attack success rate down to 5.3% of interactions, a steep drop from the 86.2% baseline, but still not zero. Finetuning an open-source model specifically to reject leakage attempts pushed the number down further still, but you can only take that path if you run open models you can retrain yourself. Enterprises using closed vendor models don't have that option.

Agentic deployments add a tradeoff that no prompt engineering can resolve. Limiting what an AI agent is allowed to ingest shrinks its attack surface, but it also shrinks what the agent can actually do for the business. Patches to products like Microsoft Copilot added options to restrict the use of external communications, so a security boundary came back at a direct cost to functionality. Deciding where that line sits is a governance question, not a technical one: someone has to decide what an agent is permitted to touch and act on, and that decision has to be made deliberately rather than left to default settings.

Even careful design-time work leaves a residue behind. Stripping credentials out of a system prompt removes the most obvious prize, but the business logic and capability details that remain still hand an attacker useful reconnaissance. And vendor patches, however well executed, tend to close one specific bypass rather than the underlying gap that produced it. EchoLeak's four-link chain made that plain: patching one bypass in the sequence didn't stop a capable adversary from finding the next one. The multi-turn extraction landscape has matured to the point where single-turn defenses no longer cover the real threat. Enterprises need to track not just how these attacks work but whether their vendors have actually adopted the layered countermeasures research has shown to be necessary, something that depends on continuous intelligence about vendor architecture and defense updates rather than a one-time review at the start of a contract.

Sources

  1. LLM02:2025 Sensitive Information Disclosure - OWASP Gen AI Security Project
  2. Prompt Leakage effect and defense strategies for multi-turn LLM interactions
  3. Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing
  4. Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks
  5. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
  6. PromptArmor Pricing & Packaging

More in Vulnerability Disclosure