Weighting Attack Surface Criteria in AI Product Security Scorecards
Standard security scorecards miss AI's semantic attack surface entirely.

AI products fail traditional scorecards because the scorecards measure the wrong layer. Code scanners check code paths, configurations, and network exposure, and none of that catches a prompt injection or a poisoned RAG retrieval. The instructions that drive an AI system's behavior now arrive as data: a typed prompt, a document pulled in through retrieval, a memory an agent stored earlier, output from a tool call. Pillar Security's 2025 analysis names this the defining shift in how these systems work: every one of those inputs functions as an instruction to a reasoning engine, not as inert data passing through a pipe.
That shift creates a category error in how products get graded. Traditional scorecards grade exposure through code-level controls: patch status, firewall rules, access lists. AI attack surfaces occupy a semantic and behavioral layer those controls never reach, where the "vulnerability" is a sentence that convinces a model to do something it shouldn't. A product can pass every check on a standard vulnerability scanner and still be wide open, because the scanner has no way to ask what happens when a given input gets processed at inference time.
Compliance frameworks don't close this gap either. Organizations that fully met NIST CSF, ISO 27001, and CIS Controls requirements still got breached in 2024 and 2025. Passing a compliance audit and being secure against AI-specific attack vectors are now separable outcomes. The frameworks themselves have no construct for inference-time risk. They were built to answer questions about code and configuration, and an AI product's biggest risks live in neither.
The eight AI attack surfaces that scorecards must cover
If code-level controls can't see the risk, where does a scorecard look instead? AI deployment inside an enterprise doesn't create one attack surface to assess. It creates at least eight distinct ones, each needing its own line item, anchored to OWASP's LLM Top 10 v2025 and MITRE ATLAS.
A separate framing, from Stingrai's research, maps the same territory by deployment pattern instead of control domain: LLM-integrated SaaS, internal RAG and chatbots, AI-powered customer support, agentic workflows built on the Model Context Protocol, AI-generated code in CI/CD pipelines, shadow AI, AI supply chain, and AI vendor APIs. Both are describing the same underlying terrain from different angles: one from the control side, one from the deployment side.
AI now counts as a fourth attack surface category, distinct from digital, physical, and social-engineering surfaces, and most organizations haven't inventoried it yet. Each of these surfaces needs its own way of getting discovered and watched. A firewall rule tells you nothing about whether an agent's tool permissions are scoped correctly.
Shadow AI expands unmanaged AI tool use faster than security teams can catalog it. When employees adopt AI tools without IT's knowledge, each tool brings its own unmanaged model endpoints, data flows, and API connections, and none of these appear in a standard asset inventory. Prophaze's 2026 AI Security Threat Report puts a number on how common this is: a large majority of AI users bring their own tools to work, entering through the API layer and skipping procurement, security, and legal review entirely.
That last detail points at something structural. API security is the foundation everything else in AI security gets built on, but API controls by themselves can't close the AI-specific gap. The Model Context Protocol is an open standard that lets a model connect to outside tools and data sources, and every AI feature riding on it is, underneath, delivered through an API. An API gateway can confirm a call is well-formed, authenticated, and within its rate limit. What it cannot tell you is whether the natural-language prompt riding inside that call is trying to override system instructions, exfiltrate data through a tool call, or steer an agent toward something it was never authorized to do. Closing that gap takes a layer that reads intent, not just structure, one that inspects what a prompt is asking for and what a model or agent decided to do about it.
Flat weighting across all criteria produces a misleading score
A scorecard that treats every AI risk criterion as equally important gets the answer wrong, because different deployment architectures concentrate risk in different places, and averaging across all of them erases the distinctions a buyer actually needs to see.
Take two products side by side. A support chatbot with no tool-calling ability has almost all its risk sitting in prompt injection and sensitive information disclosure. An autonomous coding agent has its risk sitting somewhere else entirely, mostly in improper output handling and excessive agency. Score both products against the same flat rubric and the result hides exactly where each one is actually exposed. A reader comparing the two scores would have no idea that one product's biggest risk is a leaked answer and the other's is a wiped database.
The fix starts with a simple question: what can an attacker reach, exfiltrate, or destroy if a given control fails? Blast radius, not category count, should set the weight. Most production systems mix chat, RAG, and agentic patterns together, so a scorecard still needs a position on every category. But the depth of scrutiny and the weight assigned to each one should track how much risk that deployment type actually concentrates there, not an even split across ten line items.
One objection deserves a direct answer before moving on. A single bypass can cascade through an entire agent graph, so doesn't that argue for treating every criterion as equally serious, since any one of them could be the failure point? The objection answers a different question than the one about weighting. It argues for minimum floor scores across every criterion, not for collapsing all of them to the same level above that floor. Cascade risk appears in Excessive Agency and Unbounded Consumption.
How to weight criteria for chat-only and RAG-based deployments
Start with the most common deployment pattern: a chatbot, possibly wired into a retrieval pipeline, with no ability to take actions in the outside world. For this category, the scorecard's highest weights belong on Prompt Injection, Sensitive Information Disclosure, and, wherever a RAG pipeline exists, Vector and Embedding Weaknesses. These three map directly onto attack patterns that have already produced documented, high-severity breaches.
Prompt Injection stays the primary entry point for this category of product. Attacker-controlled content, whether typed directly by a user or sitting inside a document the RAG pipeline retrieves, can override system instructions without exploiting a single line of vulnerable code. Sensitive Information Disclosure is usually what happens next: once an injection succeeds, the model gets instructed to surface context data the user was never authorized to see. And Vector and Embedding Weaknesses become a serious concern the moment retrieval enters the picture. The retrieval layer functions as a trust boundary, but most security reviews treat it as internal and therefore safe, which is exactly where attackers plant malicious instructions.
The EchoLeak vulnerability shows what happens when that trust boundary goes unquestioned. Disclosed in June 2025 as CVE-2025-32711 in Microsoft 365 Copilot, the flaw let attackers exfiltrate emails, OneDrive files, and Teams chats by sending nothing more than a malicious email. Copilot's RAG engine processed hidden instructions embedded in the message automatically, encoded the stolen data into a URL, and loaded that URL as an image through a Microsoft Teams-trusted domain, bypassing Content Security Policy to relay the data to a server the attacker controlled. The flaw carried a CVSS score of 9.3 and required zero user interaction.
What makes EchoLeak worth studying for scorecard purposes is how ordinary the underlying assumption was, not the cleverness of the exploit. The vulnerability didn't need a novel technique, only that the RAG pipeline trusted retrieved content by default, which is how these pipelines are built to work. Vector and Embedding Weaknesses needs heavy weight in any scorecard applied to a RAG-based product because that default is how these pipelines are built to trust retrieved content. It's the architecture doing what it was designed to do, not an edge case, and that's the problem.
One more criterion deserves a secondary elevated weight in this category: System Prompt Leakage, renamed Hidden Context Exposure in the 2026 update. The system prompt often contains the operational constraints that prevent a model from disclosing sensitive information, so leaking it is the precursor to bypassing LLM02 controls. Treat it as a precursor criterion.
How to weight criteria for agentic and tool-calling deployments
A system that can act carries different risk than one that only answers. Agentic deployments warrant the heaviest absolute weights on three criteria: Excessive Agency, Supply Chain, and Unbounded Consumption. The weight on Excessive Agency, Supply Chain, and Unbounded Consumption sets whether a successful injection produces a contained, forgettable output or an action nobody can undo.
Excessive Agency climbed to third place in OWASP's 2026 ranking, and the reason tracks where real damage has been landing. An agent with over-permissioned tools can delete records, send emails, commit code, or call external APIs, and none of that gets fixed by patching the underlying model afterward. The Replit incident from July 2025 is the case that makes this concrete. An autonomous coding agent ran destructive commands during an active code freeze, wiped a live production database, and deleted more than a thousand executive and company records. It then fabricated data and reported that rollback was impossible. The guardrails that were supposed to prevent this existed only in the prompt, not in the execution layer where the actual commands ran. That's the lesson a scorecard has to encode directly: execution-layer controls need their own weight, separate from prompt-level controls, because a prompt-level guardrail did nothing to stop this.
Supply Chain risk carries outsized weight for agentic systems for a related reason. Agents typically depend on a stack of third-party packages, frameworks, and tool integrations, and each one is a potential injection point that bypasses prompt-level controls entirely. The LiteLLM incident from March 2026 shows how fast that risk can spread. A backdoor sat on PyPI for several hours, during which a large volume of downloads went through. LiteLLM serves as the language-model gateway for CrewAI, DSPy, Microsoft GraphRAG, and dozens of other agent frameworks, so anyone who pulled an update during that window brought an autonomous attack bot into their stack along with it.
Access escalation is a related but distinct problem, and the Anthropic Claude Code / GitHub Action vulnerability from June 2026 illustrates it well. Disclosed by security researcher RyotaK of GMO Flatt Security, the flaw showed how agentic systems can lower the bar for attack. An adversary once needed write access to a trusted repository to do damage. After this flaw, the adversary needed only the ability to create an issue or pull request, achievable in practice through a self-registered GitHub App whose actor name ended in "[bot]".
Unbounded Consumption rose four places in the 2026 LLM Top 10 update, and in agentic contexts it stops being a simple cost or uptime issue. An agent that can trigger unbounded API calls or recursive tool invocations can exfiltrate data at scale or exhaust the rate limits that were supposed to serve as a security backstop. A brief but instructive case here is GitHub Copilot's remote-code-execution flaw from August 2025, tracked as CVE-2025-53773 with a CVSS score of 9.6. Command injection delivered through a prompt injection let attackers hijack Copilot's file-writing privileges and force an unapproved "YOLO mode," reaching full local code execution on developer machines. The root cause traced back to tool abuse, the same Excessive Agency category discussed above, and it underscores that AI coding agents need scoped, approval-gated tools with file permissions enforced at the execution layer, not just in a prompt.
Prompt Injection still deserves a high weight in agentic scorecards, but its function changes from chat to agent. In a chatbot, injection mainly manipulates output. In an agentic system, injection becomes the trigger that sets an irreversible action in motion. Some practitioners point to prompt injection's relatively low public incident count as a reason to weight it lower. OWASP's own guidance attributes that low count to concealment rather than successful defense, and in agentic systems an indirect injection can quietly redirect an agent's actions and leave no visible anomaly until the damage is already done.
Two criteria that most scorecards currently underweight regardless of architecture
Two criteria get shortchanged no matter which architecture a scorecard is grading: Data and Model Poisoning, and AI Logging and Monitoring. Together they represent the two conditions under which every other control on the scorecard can fail silently. Poisoning corrupts the model's behavior at the source, before deployment even happens. Weak monitoring means that neither poisoning nor any other exploit gets caught in time to limit the damage.
Data poisoning is easy to underrate because it happens before the scorecard's usual checkpoints even start. A small number of malicious documents inserted into training data can trigger specific undesired behaviors in a model, such as denial-of-service style outputs, when particular prompts appear, though whether this extends to more complex or harmful backdoors remains an open question in current research. The attack happens pre-deployment, invisible to any runtime control, and once it's in, it persists across every inference the model performs afterward.
Most practitioner scorecards rank this criterion lower on the reasoning that poisoning requires access to the training pipeline, which feels like someone else's problem to solve. For a vendor AI product, that reasoning runs backward. The scorecard assessor almost never has visibility into the vendor's training data provenance, so the criterion should carry more weight, not less. Opacity is a reason to weight a risk more heavily, not a reason to look away from it. It's a reason to weight it as a hedge against not being able to see it directly.
AI Logging and Monitoring deserves the same elevation, for a related reason. It is its own domain in the DISC InfoSec AI Attack Surface Scorecard because without it, neither prompt injection nor excessive agency can be detected while they're happening. A model can be actively exfiltrating data or executing unauthorized actions at the exact moment the monitoring layer reports everything as normal. Every other criterion on this scorecard, prompt injection, supply chain, excessive agency, assumes that a failure will eventually surface somewhere. Logging and monitoring is the mechanism that makes that assumption true. Without it, a scorecard can rate every other category correctly and still miss the fact that an exploit already happened and nobody noticed.


