Vulnerability Disclosure Timelines and SLA Compliance Across AI Vendors

AI systems have shattered the 90-day vulnerability disclosure window.

Staff Writer, Product & Vulnerability Coverage · · 10 min read
Cover illustration for “Vulnerability Disclosure Timelines and SLA Compliance Across AI Vendors”
Vulnerability Disclosure · October 10, 2026 · 10 min read · 2,273 words

Vulnerability disclosure timelines and SLA compliance across AI vendors have become a measurable risk category of their own, and the reason starts with a framework built for a different era. The coordinated disclosure standard that most of the software industry still leans on, report the bug privately, give the vendor a bounded window (typically 90 days) to fix it, then go public, crystallized in the early 2000s. It rests on an assumption that made sense at the time: a single researcher finds a single bug, and it takes real time and real skill for anyone else to turn that bug into a working exploit.

That assumption was scarcity. Bugs were hard to find. Exploits were hard to build. The 90-day window wasn't arbitrary, it was calibrated to how long human researchers needed to locate flaws and how long human exploit developers needed to weaponize them.

AI tooling didn't invent the speedup. The trend toward faster discovery and faster weaponization was already underway before any of today's AI systems existed. What changed is that AI removed whatever cushion remained in the system. Discovery no longer requires one researcher working one bug at a time. An AI system can scan code, verify findings, and route them to hundreds of projects within a matter of weeks, at a volume no human team can match. The one-researcher-one-bug model that the entire disclosure compact was built around has turned into something closer to a mass-submission coordination problem, and nothing about the 90-day clock was designed to handle that.

What AI-speed discovery and exploitation look like in practice

The shift isn't just that AI found bugs faster. Discovery and weaponization have become close to simultaneous, and that collapses the patch window that every enterprise remediation program is built around.

Consider the scale one disclosure program reported: verified vulnerabilities spread across hundreds of projects over roughly nine weeks, with discovery outpacing credited repairs by a ratio of roughly 16.5 to one. Researchers have a name for the widening gap this produces: a "vulnerability deficit," and it grows every day the ratio holds.

Exploitation moves on a similar curve. Working exploits for known flaws can now be produced in minutes at negligible cost, and in some categories, the time from disclosure to confirmed exploitation in the wild has fallen to less than a day. That timeline leaves almost no room for the validation, testing, and staged rollout that enterprise patch programs depend on.

This has a direct consequence for a signal many security teams still rely on: Microsoft's Exploitation Unlikely rating, used to decide which patches can wait. That rating is calibrated to what human exploit developers can do using traditional methods. It says nothing about what an AI system can do with the same public information, so treating it as a reason to deprioritize a patch no longer holds up. Microsoft Autopatch aims to get devices onto the latest quality update within a service level window that typically spans several days to a few weeks, depending on how an organization's update rings are configured. AI tooling can produce a working exploit from public patch details well before that rollout window closes. The gap between "patch released" and "patch applied everywhere" used to be a logistics problem. It's now an open door.

The patch-to-enterprise-fix gap

Diagram: The Patch Gap: From Disclosure to Enterprise Fix. Visualizes: Show the cumulative delay that stacks up between a vulnerability being reported and a fix actually running in enterprise environments.

Assume disclosure happens exactly as it should: a CVE gets assigned, a vendor ships a patch on schedule, no silent bounty, no missed advisory. Even in that best case, the path from vendor patch to a fix actually running in an enterprise environment takes months, not days.

The delay stacks up in stages. Advisory databases need time to ingest a new fix. Commercial vulnerability scanners need to refresh their detection signatures. Enterprise teams need to validate the patch against their own systems before it touches production, and most programs only start that work in earnest once a public advisory confirms the issue is real and urgent. One report put the structural estimate for this full chain, from private disclosure to a deployed enterprise fix, at somewhere between three and five months. At the snapshot date covered in that report, most of the disclosures tracked under the Mythos program had no public advisory.

Validation adds its own friction. A fix for a memory-safety bug can change the timing behavior of the surrounding code. Stricter input checks can break workflows that depended on the old, looser behavior. A dependency upgrade can force version changes across an entire chain of software that touches the patched component. For an ordinary package, validation commonly takes weeks. But if it's embedded, cryptographic, or subject to regulatory sign-off, it runs longer.

The Flowise case shows what this looks like end to end. CVE-2025-59528, a remote code execution flaw rated CVSS 10.0, the maximum severity score, got a patch in Flowise version 3.0.6 in September 2025. Confirmed exploitation in the wild wasn't documented until April 2026, about six months later, and by then thousands of publicly reachable Flowise instances were still running the vulnerable version. The disclosure process worked exactly as designed here. A CVE was assigned. A patch existed on time. The gap wasn't a failure of disclosure, it was a failure of patch management capacity across the enterprise population running the software.

That's the best-case outcome: disclosure functioning correctly still leaves a multi-month window open. The sections that follow cover what happens when disclosure itself breaks down, which is a different and deeper problem.

How agentic AI systems break the disclosure compact

The standard disclosure sequence, report, triage, patch, advise, assumes a few things about the software it's protecting: a clear owner, a flaw that can be reliably reproduced, and an impact that stays within some bounded scope. Agentic AI systems violate all three of those assumptions by design, not by accident [1].

A production agentic deployment typically stacks together a frontier model, a third-party orchestration framework, a collection of community-built MCP servers, and configuration choices made by whoever deployed the system. No single vendor in that stack has reviewed the whole thing. Accountability spreads across every layer, and that's a structural feature of how these systems get built, not a gap someone forgot to close.

Non-determinism compounds the problem. The same prompt can produce different outputs depending on conversation history, retrieved context, or what else is running at the time. CVE assignment and CVSS scoring both assume a flaw that can be demonstrated reliably on demand. An agentic vulnerability that only shows up under certain conversational conditions doesn't fit that mold, and the tooling built to score and track vulnerabilities has no good answer for it.

Agents also hold things that stateless software never did: credentials, session tokens, memory that persists across separate interactions. That creates a lateral movement surface with no precedent in traditional software security. A compromised agent isn't just a compromised piece of code, it's a compromised principal acting with whatever authority the user delegated to it.

The Model Context Protocol, introduced in late 2024, illustrates the speed of the problem. MCP created an entirely new category of software, the MCP server, that connects AI models to external tools, databases, and APIs. Thousands of community-built implementations spread across the ecosystem before any disclosure process had a chance to adapt to them. Between January and February 2026 alone, researchers filed more than thirty CVEs targeting MCP servers, clients, and infrastructure, and independent analysis suggests that count captures only a fraction of what's actually exploitable. One analysis of surveyed MCP servers found that the large majority use file operations vulnerable to path traversal. An AI model mediates requests between a client and a server it doesn't fully control, so the architecture itself creates exploitation scenarios that have no established disclosure pathway to route through.

The "silent bounty" problem: how accepted vulnerability reports disappear from the ecosystem

A security researcher reports a real vulnerability, the vendor's security team accepts the report, pays a bounty, and the story ends there. No CVE gets filed. No advisory gets published. The researcher got paid. Every enterprise running the affected product has no way of knowing the risk exists.

The sequence that's supposed to generate a signal for enterprise defenders, report accepted, bounty paid, CVE filed, advisory published, breaks at the last two steps across a documented pattern of cases in the AI agent ecosystem, with a vulnerability accepted and paid but never disclosed to the people running the affected product. The report gets accepted. The bounty gets paid. No notification reaches the population of users actually running the thing.

The Anthropic SQLite MCP Server case shows the downstream cost of that break. A researcher privately reported a SQL injection vulnerability in a reference implementation that had already been forked widely and folded into production workflows well beyond its original intent. Anthropic's response treated the repository as an archived demonstration implementation, placed it out of scope, assigned no CVE, and planned no patch. Every developer who had forked that code inherited the vulnerability without any way of knowing it was there.

Fourth-party risk makes the blind spot worse. When a SaaS vendor builds its product on top of a third-party model provider, the enterprise customer has no direct contractual relationship with that model provider. The SaaS vendor, in turn, often can't make binding guarantees about a model it doesn't control and can't fully audit.

The EchoLeak vulnerability shows how far the consequences can reach. CVE-2025-32711, rated CVSS 9.3, affected Microsoft 365 Copilot and required no direct user interaction with the malicious email that carried it. A prompt embedded in the email triggered silent exfiltration of sensitive organizational data the moment the victim later invoked Copilot for something unrelated. A major, trusted AI vendor's own model became the attack vector, through no action the victim could see or avoid.

Traditional third-party risk management has no mechanism built for catching a risk that never enters the CVE feed. Continuous monitoring exists to close that gap: the acceleration of discovery and exploitation timelines is why ongoing visibility into how AI vendors respond to emerging threats has become necessary, since the patch windows that conventional vendor risk reviews assume have collapsed from days to hours. A point-in-time vendor questionnaire can't catch a silent bounty. Only continuous visibility into what a vendor is actually doing, and what's actually running in production, has a chance of catching it. So you need to ask next whether a given vendor's disclosure posture makes that visibility possible, or makes it structurally impossible.

How AI vendor disclosure policies differ

Vendor disclosure policies vary enough that the differences amount to a measurable risk signal in their own right, not a matter of house style.

Anthropic's Project Glasswing publishes findings under a Coordinated Vulnerability Disclosure policy built around a 90-day default window, with defined exceptions: a 14-day extension for vendors making genuine progress, a 7-day timeline for critical vulnerabilities under active exploitation, and separate handling for patterns that span an entire ecosystem. As of October 2026, Anthropic had disclosed thousands of vulnerabilities across hundreds of open-source projects under this policy, with hundreds of identifiers issued as formal CVE records or GitHub Security Advisories. Anthropic also withheld Claude Mythos Preview from public release and built a disclosed-ledger framework to coordinate AI-generated discovery at scale.

OpenAI's approach differs in a specific and consequential way. Its Outbound Coordinated Disclosure Policy governs how OpenAI reports vulnerabilities it finds in third-party software, but the company takes an intentionally open-ended stance on timelines, and it won't commit to a fixed default window. On the inbound side, researchers who report vulnerabilities to OpenAI are required to keep details confidential until OpenAI's security team authorizes release, with public disclosure permitted afterward through a coordinated timeline. A researcher who believes OpenAI's remediation is moving too slowly has limited options for pushing the issue into public view. An enterprise watching from the outside has no upstream signal telling it a clock is even running.

Oracle offers a useful contrast from outside the AI-native vendor world. Oracle announced that starting May 28, 2026, it would issue Critical Security Patch Updates on a monthly cadence in the months between its existing quarterly Critical Patch Updates, and it named AI-driven acceleration of vulnerability discovery and remediation speed as the reason. A legacy enterprise software vendor changing its patch cadence specifically because of AI-driven discovery speed is itself an admission that the old quarterly rhythm no longer matches the threat.

These aren't stylistic choices. A 90-day default with narrow, published exceptions gives an enterprise something it can plan around. An open-ended default with a confidentiality requirement on inbound researchers gives an enterprise nothing to plan around. Evaluating an AI vendor's disclosure policy design, not just its existence, is something enterprises can reasonably demand as part of vendor risk review, and services built around continuous assessment of vendor AI security posture, Promptarmor among them, exist precisely because that demand has nowhere else to go once the questionnaire stage ends.

Why existing compliance frameworks cannot close this gap

Major compliance frameworks all specify patching and notification timelines, and none of them were built with AI-speed discovery in mind. They were calibrated for a world where critical vulnerabilities surfaced one or two at a time, found by skilled human researchers working at human pace.

That calibration is visible in the timelines themselves. A framework that expects a security team to triage, assess, and remediate a handful of findings a month runs into a different problem entirely when an AI system can surface vulnerabilities across an entire software stack at once. Passing an audit built on pre-AI assumptions doesn't mean the underlying exposure has closed; it means the measuring stick hasn't caught up yet.

Sources

  1. CSAI Foundation
  2. The AI Agent Disclosure Vacuum
  3. Measuring the Vulnerability Disclosure Policies of AI Vendors
  4. The Collapsing Exploit Window: AI-Speed Vulnerability Weaponization

More in Vulnerability Disclosure