Booz Allen Cyber Weapon Index measures demonstrated AI cyber capabilities

New index measures AI’s ability to execute cyberattacks

Traditional assumptions about offensive and defensive cyber operations are rapidly becoming obsolete. AI can now autonomously execute much of an intrusion, with humans intervening only at critical decision points—dramatically increasing the speed, scale, precision, and persistence of sophisticated attacks. 

The frontier is already moving

Less than a week after we published the Booz Allen Cyber Weapon Index (CWI), new testing changed the leaderboard. Both OpenAI’s GPT-6 Astra and Anthropic’s Claude Mythos have now demonstrated the ability to autonomously execute the full cyber kill chain. Five new models also entered the Top 18.

The bigger story is the pace of change. What was a single-model outlier at publication is already becoming a competitive tier, reinforcing the need to continuously measure these capabilities as new models emerge.

Speed Read ↗︎

  • Autonomous AI models have crossed a critical threshold and are now capable of independently executing sophisticated cyberattacks. 

  • The Booz Allen Cyber Weapon Index (CWI) reveals that, while one frontier model can fully complete the cyber kill chain today, a broad range of U.S. and Chinese models are rapidly advancing, and the real danger lies in the combined power of models, harnesses, and tools.

  • One week after this report’s initial publication, OpenAI’s Astra joined Anthropic’s Claude Mythos as a second model capable of autonomously executing the full cyber kill chain—Mythos still leads in execution, but Astra now leads in vulnerability discovery.

With increasing speed and precision, AI models are now the attacker.

We now face three increasingly critical questions:

  1. What is the current state of offensive cyber capability among leading U.S. and Chinese models?
  2. How long will it be before the rest of the global model landscape can execute the full cyber kill chain?
  3. How does the United States exploit agentic AI to generate and maintain offensive and defensive overmatch?

 

The Booz Allen Cyber Weapon Index

The Booz Allen Cyber Weapon Index (CWI) is a benchmark measuring the real offensive capability of AI systems in a live environment. In other words, we captured what models actually do in a live cyber environment, not just what they claim. In the original test, we evaluated 18 leading U.S. and Chinese large language models (LLMs) as autonomous attackers, each controlling a real attacker machine against a production-grade enterprise network.

The models were tested under identical conditions, without a curated tool menu or additional scaffolding, allowing the CWI to isolate the models’ demonstrated capabilities in a realistic environment. During the test, models issued commands independently, with every action validated through network telemetry, host logs, domain controller data, and intrusion-detection sensors, ensuring scores reflect demonstrated behavior, not theoretical skill.

These conditions provided a clearer view of how far a model could progress through an intrusion, how it adapted, and whether it could generate new offensive capabilities.

infographic describing the scoring model behind the cyber weapon index

Two Models Can Now Complete the Full Cyber Kill Chain

Our original testing found that only 1 of the 18 models, Anthropic’s Claude Mythos, could execute the full cyber kill chain. Less than a week later, new testing identified a second: OpenAI’s GPT-6 Astra. However, this is not the threshold for danger, or the real story.

Nearly all models tested reached reconnaissance and initial access reliably. Many carried that momentum into credential access, but from there the field thinned: some reliably pushed through privilege escalation and lateral movement, while many others stalled entirely. Only a few models reached Domain Admin with any consistency, and just one model initially—and now two—reliably succeeded at real-world common vulnerabilities and exposures (CVE) exploitation in the virtualized range—a capability the rest of the field failed to reach altogether.

Our original testing also showed limited differentiation between U.S. and Chinese models. The latest testing reinforces how quickly the field is moving, and also shows the U.S. lead is expanding. Two models have now demonstrated full-chain autonomous cyber capability, and five new models entered the Top 18 in less than a week. Our assessment remains that most models will arrive at this capability within the next six months.

Our Findings

  • Two frontier AI models can now autonomously execute the full cyber kill chain, but real-world vulnerability discovery remains a major dividing line.
  • Cyber capability extends well beyond leading frontier models, and while the U.S. lead is expanding, it is not country-specific.
  • An attack harness can matter as much as, or more than, the model itself.
  • AI safeguards shift with configuration and context. 
  • The most dangerous AI-cyber capabilities may be both beyond U.S. regulatory reach and largely invisible to traditional evaluations.
  • The United States can create decisive cyber overmatch by mastering agentic AI on both offense and defense.
The Warning Period Is Closing

Autonomous cyber operations are no longer hypothetical. The organizations and nations that gain advantage will be those that can measure emerging capabilities, accelerate AI-enabled defense, and adapt before adversaries do.

For the complete analysis, methodology, findings, and recommended actions, download the updated report.