AI Vendor Bug Bounty Program Scope Gaps and Exclusion Analysis
Major AI vendors exclude model behavior from bug bounties, treating jailbreaks as unfixable.

A bug bounty works on a simple chain: someone finds a flaw, the vendor fixes it, a patch goes out, and an advisory tells the rest of the world what happened. That chain depends on the flaw being a single, reproducible thing, a bad line of code, a missing check, a broken permission. Jailbreaks and model-response issues don't behave that way. A model that can be coaxed into saying something harmful today might refuse the same prompt tomorrow and fail again next week under a slightly different phrasing, and there's no single line to patch because the "bug" lives in the statistical behavior of the whole system. Vendor policy language now says this directly: jailbreaks and model-response issues are not treated as discrete bugs fixable with a patch, and addressing them calls for broader research instead. The exclusion of "all content and model response issues" is a standing position in program language, not a placeholder waiting to be revised, because these issues are described as unfixable through the normal security-patch process. If you read that carefully, you see the exclusion categories stop looking like gaps vendors forgot to close. They're decisions about what counts as a security bug in an AI system at all, and those decisions carry weight for everything that follows.
What the major programs cover
Across the major programs, the same boundary repeats: infrastructure bugs are in, model behavior is out. By name, Google excludes prompt injection, jailbreaks, and alignment issues from its AI Vulnerability Reward Program. Submit a genuinely clever jailbreak to that program and the payout is zero, no matter how much work went into it. Standard jailbreaks, hallucinations, and prompt extraction that doesn't expose sensitive data also fall outside what Google will reward.
Microsoft's Copilot Bounty program draws a similar line but with different wording. It excludes AI prompt injection attacks if they carry no security impact on anyone besides the attacker, and it excludes attacks that aim to leak part of the system or meta prompt. Awards on that program run from $250 to $30,000, with Microsoft reserving the right to pay more at its own discretion. Both figures sound generous, but so much of the attack surface never gets to the scoring table.
OpenAI's program on Bugcrowd and Anthropic's program on HackerOne scope in model and product vulnerabilities, but both apply hard limits. General content-policy bypasses with no demonstrable safety or abuse impact don't qualify. Jailbreaks that just produce rude language, or that surface information a search engine would hand over anyway, are out. OpenAI's Safety Bug Bounty does carve out one real exception: third-party prompt injection and data exfiltration count when attacker text can reliably hijack a victim's agent. But "reliably" comes with a number attached, the behavior has to reproduce at least fifty percent of the time. Critics have pushed back hard on that bar, pointing out that a vulnerability doesn't need to fire every other try to cause real damage. One working exploit chain, triggered once in a high-value context, can do more harm than a flaky bug that fires on command.
The pattern across all four programs is the same doctrine from the first section, applied with slightly different thresholds each time.
The partial exception: how Anthropic's model safety track tries to bridge the gap
Anthropic runs a Model Safety track that does something the other programs don't: it pays out for jailbreak techniques. That alone makes it worth separating from the rest of the field. But the track is application-gated: a researcher has to apply and qualify before submitting, and the harm categories it rewards are set in advance rather than open to whatever a researcher happens to find.
The target is narrow by design. The track looks for universal jailbreaks, ones that consistently bypass safety measures across many different queries within specific high-risk domains like CBRN and cybersecurity. If a bypass only works in one narrow framing, against one refusal category, it earns nothing. The finding has to generalize across multiple high-stakes categories, or it doesn't count as in scope.
That narrowness is what makes the whole thing work financially for Anthropic. A universal, repeatable bypass of a named safety classifier behaves almost like a conventional bug: it's reproducible, it has a defined trigger, and fixing the classifier addresses the underlying issue in a way that resembles a patch. That's a very different animal from the open-ended "does this model ever say something bad" question that every other program refuses to touch. The track proves that jailbreaks can be made bounty-compatible, but only after being narrowed down to something that stops looking like a model-behavior problem and starts looking like a discrete, classifier-level bug.
Mozilla's 0din as a structural alternative to the infrastructure-only default
Mozilla's 0din program, launched in mid-2024, takes a different approach from the infrastructure-first default that governs Google, Microsoft, and OpenAI's main bounty tracks. It's built from the ground up around large language model behavior; that behavior isn't bolted on as an afterthought to a conventional program.
0din's scope reads like a direct answer to everything excluded elsewhere: prompt injection, guardrail jailbreaks, training data poisoning, denial of service, and the categories listed in the OWASP LLM Top 10. It casts a wider net than any single model-provider program does on its own. The submission process reflects the same intent. Researchers submit an abstract before a full exploit writeup, which lowers the cost of exploring a novel AI-specific finding that might otherwise get closed out as out-of-scope somewhere else before anyone takes a serious look at it.
The limit appears once a finding actually lands. 0din sits outside the major platform ecosystems, so if a report gets accepted, no patch automatically follows inside the model providers' production systems. Breadth of scope and reach into the systems people actually use are pulling in opposite directions here. A researcher can get a finding validated and rewarded through 0din while the underlying model a million people use every day stays the same, because the component responsible for the vulnerability belongs to someone else.
Where indirect prompt injection and MCP supply-chain attacks fall between program boundaries
Some of the clearest evidence for all of this comes from attacks that are already happening in the wild, not hypothetical ones dreamed up in a research paper. Indirect prompt injection now shows up in production environments, not just as a theoretical concern, but most bounty programs either exclude it outright or set the reproducibility bar so high that real attack chains can't clear it.
In late April 2026, Google's Security blog and Forcepoint X-Labs each confirmed that attackers are planting hidden instructions across the open web to hijack AI agents that happen to read the wrong page. The attacks seen so far stayed relatively unsophisticated and experimental, not fully industrialized, but they are live exploitation, not a lab demonstration. Google reported a measurable rise in malicious indirect prompt injection content across the billions of pages its crawlers process each month between November 2025 and February 2026.
The program rules layered on top of that make the gap obvious. Microsoft's Copilot program excludes prompt injection without cross-user impact. OpenAI's safety track requires fifty percent reproducibility. Both thresholds rule out a large share of real indirect injection attacks, the ones where the harm is genuine but confined to a single user, or where success depends on conditions that don't repeat cleanly every time.
The MCP supply chain sharpens the same point from a different angle. In the opening months of 2026, researchers filed dozens of CVEs against MCP servers, clients, and the infrastructure around them. The root causes read like a greatest-hits list from package ecosystem security: missing input validation, absent authentication, blind trust placed in tool descriptions that turned out not to deserve it. These are the same failure classes that plagued software package registries for years, now present in a new layer of the AI stack. Agent-mode sandboxing exclusions make the picture worse: code execution inside Agent Mode is formally out of scope when it happens in a contained environment, even when that contained environment is exactly where a real attack chain starts and ends.
The Flowise case offers a useful parallel, even though it's not a prompt injection story. CVE-2025-59528 was a CVSS 10.0 remote code execution vulnerability in one of the most widely deployed open-source AI agent builders, and nearly seven months passed between the patch and confirmed disclosure of active exploitation. When bounty scope is ambiguous, you get timeline gaps like that one.
How scope exclusions break the patch-and-advisory sequence
A vendor can accept a report, pay the researcher, and still never file a CVE or publish an advisory. When that happens, the sequence that normally spreads risk awareness through the wider software ecosystem stops working, and AI agent programs have started to do just that. The conventional sequence runs: report accepted, bounty paid, patch released, advisory published, and a CVE comes out the other end that downstream vendors and enterprise security teams can act on. Break any link in that chain and the vulnerability stays live in production while disappearing from view for everyone except the vendor who found it.
Enterprises running the affected product have no way to learn a risk exists. Downstream vendors who build on the same component can't patch something they don't know is broken. The vulnerability keeps running in deployed systems while vanishing from the collective record that the rest of the industry relies on. Researchers and enterprise teams have already had to deal with this pattern directly in the AI agent ecosystem.
CISA leadership has said as much in public, suggesting AI companies should be better represented in the CVE program. That framing understates what's actually broken. Getting more vendors to participate leaves the attribution problem for components built by loose communities of open-source contributors unsolved, leaves the taxonomy problem for failure modes that don't map cleanly onto existing vulnerability categories unsolved, and leaves the reproducibility problem for bugs that only show up sometimes unsolved. A CVSS 10.0 remote code execution flaw in a widely deployed open-source AI agent builder had nearly seven months separating public disclosure from confirmed active exploitation. That's seven months during which any security team or third-party risk team using the product had no signal that anything was wrong, because the normal channels that would have told them never fired.
Triage Bottleneck, AI-Generated Report Volume, and Program Economics
Bug bounty programs face a squeeze from two directions at once. Bug bounty programs respond to those two pressures by tightening scope, raising the bar for what counts as evidence before a finding gets taken seriously.
The real bottleneck is expertise. Reports get closed as informational because the person triaging them doesn't understand how prompt injection actually gets exploited, and that gap in understanding hits hardest at the programs most willing to consider AI-specific findings. Google has temporarily closed its Open Source Software Vulnerability Reward Program to product vulnerability submissions, because it says a growing flood of automated reports, most of them invalid, is the direct cause. You can already see what AI-generated report volume is doing to program economics elsewhere.
The direction all of this points is toward higher proof requirements and narrower, more precisely defined scope, not broader coverage of model-behavior findings. Most programs nominally accept indirect prompt injection as in scope, but in practice triage consistently undervalues it, especially when the actual harm depends on chaining several steps together. Even if a finding is technically in scope, it often goes unrewarded because the person reviewing it can't see the chain.
Headline bounty numbers and the median payout tell different stories, but the reason is the same. The headline figures come from critical findings against core infrastructure. The realistic middle band for AI-specific findings sits well below that, and the gap between the two is a direct readout of how triage is actually handling this category of report.
Why scope exclusions require continuous vendor-side detection
None of the exclusions covered so far are gaps waiting for a better-designed program to fill them in. They mark the edge of what a discrete-fix model can absorb, and the attack surface sitting past that edge needs continuous detection rather than the occasional researcher who happens to stumble onto it.
Because the advisory sequence is broken, you can't lean on CVE feeds or vendor disclosures to find out about AI-specific vulnerabilities sitting inside products you've already deployed. The signal doesn't travel through the channels that used to carry it. Jailbreaks, indirect prompt injection, model-behavior bypasses, and MCP supply-chain weaknesses all share the same awkward property: they're emergent, they depend on context, and they're often stochastic. Those traits make something a poor fit for a bounty program and make a strong case for continuous monitoring of deployed AI systems.
This surface is not what traditional security and governance frameworks were built for. A vendor AI product that cleared a point-in-time assessment last quarter can pick up new prompt injection exposure the moment its underlying model gets updated, a new plugin gets added, or a dependency changes, and no bounty program is scoped to catch any of that as it happens. That's the space no disclosure norm currently covers: the stretch of time between a vulnerability coming into existence and anyone outside the vendor finding out about it, if they ever do.
For security teams, third-party risk teams, privacy teams, and legal teams working with vendor AI products, that stretch of time is the actual exposure they're carrying. Continuous monitoring of the vendor AI surface is what corresponds to the part bug bounty programs were never built to cover, and understanding why the programs stop where they do is the first step toward building a detection approach that doesn't stop in the same place.


