Traditional assumptions about offensive and defensive cyber operations are rapidly becoming obsolete. AI can now autonomously execute much of an intrusion, with humans intervening only at critical decision points—dramatically increasing the speed, scale, precision, and persistence of sophisticated attacks.
Less than a week after we published the Booz Allen Cyber Weapon Index (CWI), new testing changed the leaderboard. Both OpenAI’s GPT-6 Astra and Anthropic’s Claude Mythos have now demonstrated the ability to autonomously execute the full cyber kill chain. Five new models also entered the Top 18.
The bigger story is the pace of change. What was a single-model outlier at publication is already becoming a competitive tier, reinforcing the need to continuously measure these capabilities as new models emerge.
Autonomous AI models have crossed a critical threshold and are now capable of independently executing sophisticated cyberattacks.
The Booz Allen Cyber Weapon Index (CWI) reveals that, while one frontier model can fully complete the cyber kill chain today, a broad range of U.S. and Chinese models are rapidly advancing, and the real danger lies in the combined power of models, harnesses, and tools.
One week after this report’s initial publication, OpenAI’s Astra joined Anthropic’s Claude Mythos as a second model capable of autonomously executing the full cyber kill chain—Mythos still leads in execution, but Astra now leads in vulnerability discovery.
We now face three increasingly critical questions:
The Booz Allen Cyber Weapon Index (CWI) is a benchmark measuring the real offensive capability of AI systems in a live environment. In other words, we captured what models actually do in a live cyber environment, not just what they claim. In the original test, we evaluated 18 leading U.S. and Chinese large language models (LLMs) as autonomous attackers, each controlling a real attacker machine against a production-grade enterprise network.
The models were tested under identical conditions, without a curated tool menu or additional scaffolding, allowing the CWI to isolate the models’ demonstrated capabilities in a realistic environment. During the test, models issued commands independently, with every action validated through network telemetry, host logs, domain controller data, and intrusion-detection sensors, ensuring scores reflect demonstrated behavior, not theoretical skill.
These conditions provided a clearer view of how far a model could progress through an intrusion, how it adapted, and whether it could generate new offensive capabilities.
Our original testing found that only 1 of the 18 models, Anthropic’s Claude Mythos, could execute the full cyber kill chain. Less than a week later, new testing identified a second: OpenAI’s GPT-6 Astra. However, this is not the threshold for danger, or the real story.
Nearly all models tested reached reconnaissance and initial access reliably. Many carried that momentum into credential access, but from there the field thinned: some reliably pushed through privilege escalation and lateral movement, while many others stalled entirely. Only a few models reached Domain Admin with any consistency, and just one model initially—and now two—reliably succeeded at real-world common vulnerabilities and exposures (CVE) exploitation in the virtualized range—a capability the rest of the field failed to reach altogether.
Our original testing also showed limited differentiation between U.S. and Chinese models. The latest testing reinforces how quickly the field is moving, and also shows the U.S. lead is expanding. Two models have now demonstrated full-chain autonomous cyber capability, and five new models entered the Top 18 in less than a week. Our assessment remains that most models will arrive at this capability within the next six months.
Autonomous cyber operations are no longer hypothetical. The organizations and nations that gain advantage will be those that can measure emerging capabilities, accelerate AI-enabled defense, and adapt before adversaries do.
For the complete analysis, methodology, findings, and recommended actions, download the updated report.