CyberIntel ⬡ News
★ Saved ◆ Cyber Reads
← Back ◉ Threat Intelligence Aug 07, 2026

Expanding AI Benchmarks in Cybersecurity Beyond Vulnerability Discovery

CrowdStrike Archived Aug 07, 2026 ✓ Full text saved
Full text archived locally
✦ AI Summary · Claude Sonnet


    ___ Blog Featured Recent Video Category Start Free Trial Expanding AI Benchmarks in Cybersecurity Beyond Vulnerability Discovery August 06, 2026 • Keegan Hines - Chase Midler • Securing AI The conversation about AI in cybersecurity has recently centered on capabilities like vulnerability discovery, exploit generation, and automated proof-of-concept development. It’s easy to see why: These tasks produce binary outcomes; a vulnerability either exists or it doesn't. That makes them useful for measuring model progress and demonstrating increasingly sophisticated cybersecurity capabilities. Vulnerability discovery matters to defenders. According to the Verizon 2026 Data Breach Investigations Report, vulnerability exploitation is now the most common initial access vector, accounting for 31% of breaches in the reporting dataset. This is a meaningful and growing share of the problem and a strong reason to continue advancing AI capabilities in this area. But it also means 69% of breaches begin through other paths. Credential abuse, phishing, social engineering, trusted relationships, and other forms of access remain central to the adversary playbook. A comprehensive evaluation framework should therefore measure not only whether AI can discover and exploit vulnerabilities, but also whether it can help defenders detect identity abuse, investigate suspicious activity, engineer effective detections, hunt for adversaries, and respond across the broader attack lifecycle. We believe effective AI for defense must be evaluated against the operational reality of security teams: the range of techniques adversaries use to gain initial access, the work required across the kill chain, and defenders’ most time-consuming tasks. Here, we explore some of these use cases. Detecting Adversary Behavior After Initial Access Once an adversary is inside, the defender’s work becomes more complex. Security teams must detect and triage suspicious activity across massive alert volumes, balancing signal and noise. Speed here determines whether the adversary is contained in minutes or operates on a network for a longer period of time. When suspicious activity is found, the focus shifts to investigation, which requires significant effort and expertise. Analysts must reconstruct events across endpoints, identities, cloud environments, and other systems to determine what happened and how far the adversary has progressed. This is where attackers gain time: The longer an investigation takes, the greater the opportunity to move laterally, escalate privileges, and achieve their objectives. Detection engineering powers these investigations by translating adversary behavior into durable detections, yet it remains a specialized and often under-resourced discipline. Effective defense also requires proactive threat hunting to find adversaries operating below existing detection thresholds. This depends on skilled analysts and accurate models of real-world behavior. None of these disciplines produce clean binary outcomes or are easily benchmarked. However, they all depend on grounded, accurate models of real adversary behavior that public benchmarks lack. Evaluating AI Against Adversary Tradecraft Proactive defense requires knowing exactly how adversaries operate: the tools they use, the sequence of actions they take, and the ways their activity appears in telemetry. Adversary emulation translates that knowledge into realistic activity, making it foundational to detection engineering, threat hunting, and security validation. The critical factor here is being grounded in reality. AI red teaming based on hypothetical scenarios or synthetic data has limited value if it does not reflect the techniques, artifacts, and operational patterns observed in real intrusions. A key question for AI is whether it can reproduce adversary behavior with enough fidelity to test detection coverage, expose engineering gaps, and improve defensive readiness. For most defenders, that is a harder and more consequential test of AI capability than vulnerability discovery alone. Figure 1. ALTERED SPIDER profile from the CrowdStrike Adversary Universe Because CrowdStrike observes trillions events each day across a global customer base, our threat intelligence is sourced from real intrusions rather than synthetic scenarios. This operational visibility is the foundation for our adversary emulation, detection engineering, and AI evaluation based on how attackers operate. Frontier labs can build highly capable models, but they lack the same access to operational threat intelligence generated across thousands of customer environments. This is the difference between a model that scores well on a public benchmark and one that is genuinely useful in a specific environment. Real telemetry closes that gap by showing which adversaries matter, how their behavior appears, and whether an AI system can help defenders detect and respond to it. Where Public AI Security Benchmarks Fall Short Public benchmarks for AI security evaluation suffer from several compounding limitations. One is agenda capture: Benchmarks often reflect the research priorities of those who create them, rather than the operational effectiveness security teams need to achieve. A second problem is saturation. As top-tier models begin clustering near the ceiling of public benchmarks, the tests become less useful for distinguishing meaningful differences in capability. When every leading model scores close to 100%, the benchmark becomes little more than a threshold test.  Contamination creates an additional challenge. Once benchmark content enters the broader data ecosystem, models can be trained directly or indirectly toward the test. Over time, a benchmark intended to measure general capability may instead measure how effectively a model has been optimized for a familiar evaluation. The deeper fundamental issue is that public benchmarks do not evaluate whether AI can detect real adversary behavior in real telemetry, generate detection logic for specific lateral movement techniques, or improve the operational work defenders perform every day. Closing that gap requires more than better public benchmarks alone. It requires high-fidelity telemetry, real-world threat intelligence, and the operational context needed to evaluate AI against how adversaries behave. What Meaningful Security Evaluations Require Evaluation frameworks designed for defenders must meet several requirements.  First, they must be task-relevant to the functions where security teams spend their time, including triage, investigation, remediation, and threat hunting. Second, they must be grounded in telemetry. Evaluating AI against real adversary behavior in realistic environments produces results that are more predictive of operational performance than tests built around synthetic attack scenarios. Finally, they need to be customer-specific. A financial institution faces a different threat profile than a healthcare provider or a critical infrastructure operator, and the benchmarks that matter for each are not the same. A single public leaderboard cannot serve all of them. The right framework must be adapted to the specific adversaries and techniques most relevant to a specific customer's environment, and account for the less common techniques and behaviors where advanced adversaries often operate. CrowdStrike’s goal is to create an entirely different measurement framework that begins with the defender's operational reality rather than producing a better version of what is easiest to score today. Final Thoughts AI for vulnerability discovery will continue to advance and generate significant attention. But for security teams, the more consequential question is whether AI can help detect, investigate, and disrupt adversaries after they gain access, regardless of the initial access mechanism. Answering that question requires evaluation grounded in real threat intelligence, real telemetry, and the operational reality of defense. With visibility across a global security ecosystem and deep experience in how adversaries operate, CrowdStrike is positioned to define how AI should be measured against the outcomes defenders actually need. At Fal.Con 2026, we will show how our work is moving from principle to practice. Additional Resources Download our guide to explore the five steps for frontier AI security readiness. Visit our Frontier AI Readiness and Resilience Service page to learn about CrowdStrike’s approach to frontier AI and see how CrowdStrike Services can help. Join us at Fal.Con 2026 as we bring together cyber leaders from across the industry to help secure the AI revolution. CrowdStrike 2026 Global Threat Report AI threats have reached a critical turning point. Access the definitive look at the cyber threat landscape. Download Related Content Securing AI | Aug 04, 2026 Secure Agent Harness Execution: Preventing Escape Securing AI | Jul 30, 2026 Falcon AIDR Now Protects Copilot Studio Agents and Claude Code Securing AI | Jul 27, 2026 CrowdStrike Joins the Open Secure AI Alliance to Advance AI Safety and Security Categories Agentic SOC 53 Cloud & Application Security 148 Data Security 25 Endpoint Security & XDR 361 Engineering & Tech 87 Executive Viewpoint 181 Exposure Management 123 From The Front Lines 205 Next-Gen Identity Security 74 Next-Gen SIEM & Log Management 115 Public Sector 43 Securing AI 49 Threat Hunting & Intel 221 CrowdStrike Falcon Platform Ready to protect your business? Try CrowdStrike free today Start free trial Subscribe Sign up now to receive the latest notifications and updates from CrowdStrike Subscribe See CrowdStrike Falcon in action Explore demos Copyright © 2026 CrowdStrike Privacy Request Info Blog Contact Us 1.888.512.8906 Accessibility Privacy Preference Center Privacy Preference Center Your Privacy Strictly Necessary Cookies Performance Cookies Functional Cookies Targeting Cookies Your Privacy When you visit any website, it may store or retrieve information on your browser, mostly in the form of cookies. This information might be about you, your preferences, or your device, and is mostly used to make the site work as you expect. The information does not usually identify you directly, but it can give you a more personalized web experience. Because we respect your right to privacy, you can choose not to allow some types of cookies. Click on the different category headings to learn more and change our default settings. Blocking some types of cookies may impact your experience of the site and the services we are able to offer. More information Strictly Necessary Cookies Always Active These cookies are necessary for the website to function and cannot be switched off in our systems. They may be set in response to actions made by you which amount to a request for services, such as setting your privacy preferences, logging in or filling in forms. You can set your browser to block or alert you about these cookies, but some parts of the site will not then work. These cookies may process limited personal information, such as technical or device identifiers, where necessary to ensure the security, functionality, and integrity of the website or web portal. Such processing is strictly limited to what is required for these purposes and is not used for advertising or marketing. Cookies Details Performance Cookies Performance Cookies These cookies allow us to count visits and traffic sources so we can measure and improve the performance of our site. They help us to know which pages are the most and least popular and see how visitors move around the site. All information these cookies collect is aggregated and therefore does not identify you. If you do not allow these cookies, your visit to our website will not be included in our analytics, and our ability to monitor website performance and make improvements will be reduced. Cookies Details Functional Cookies Functional Cookies These cookies enable the website to provide enhanced functionality and personalisation. They may be set by us or by third party providers whose services we have added to our pages. If you do not allow these cookies then some or all of these services may not function properly. Cookies Details Targeting Cookies Targeting Cookies These cookies may be set on our site by our advertising partners. They assign a unique identifier to your browser or device and may track your activity across sites to build a profile of your interests and show you relevant adverts on other sites. If you do not allow these cookies, you will still see ads, but they may be less relevant to you. Cookies Details Cookie List Consent Leg.Interest checkbox label label checkbox label label checkbox label label Clear checkbox label label Apply Cancel Confirm My Choices Allow All
    💬 Team Notes
    Article Info
    Source
    CrowdStrike
    Category
    ◉ Threat Intelligence
    Published
    Aug 07, 2026
    Archived
    Aug 07, 2026
    Full Text
    ✓ Saved locally
    Open Original ↗