Traditional assumptions about offensive and defensive cyber operations are rapidly becoming obsolete. AI can now autonomously execute much of an intrusion, with humans intervening only at critical decision points—dramatically increasing the speed, scale, precision, and persistence of sophisticated attacks.
Autonomous AI models have crossed a critical threshold and are now capable of independently executing sophisticated cyberattacks.
The Booz Allen Cyber Weapon Index (CWI) reveals that, while one frontier model can fully complete the cyber kill chain today, a broad range of U.S. and Chinese models are rapidly advancing, and the real danger lies in the combined power of models, harnesses, and tools.
To maintain cyber overmatch, organizations and government must urgently accelerate machine-speed defenses, set binding readiness standards, and continuously measure global AI capabilities before adversaries operationalize these threats at scale.
We now face three increasingly critical questions:
The Booz Allen Cyber Weapon Index (CWI) is a benchmark measuring the real offensive capability of AI systems in a live environment. In other words, we captured what models actually do in a live cyber environment, not just what they claim. In the test, we evaluated 18 leading U.S. and Chinese large language models (LLMs) as autonomous attackers, each controlling a real attacker machine against a production-grade enterprise network.
The models were tested under identical conditions, without a curated tool menu or additional scaffolding, allowing the CWI to isolate the models’ demonstrated capabilities in a realistic environment. During the test, models issued commands independently, with every action validated through network telemetry, host logs, domain controller data, and intrusion-detection sensors—ensuring scores reflect demonstrated behavior, not theoretical skill.
These conditions provided a clearer view of how far a model could progress through an intrusion, how it adapted, and whether it could generate new offensive capabilities.
Our testing found that only 1 of the 18 models (Anthropic’s Claude Mythos) can execute the full cyber kill chain today. However, this is not the threshold for danger—or the real story.
The risk is already distributed across the field:
Our testing also showed no substantial differentiation between U.S. and Chinese models. While only Claude Mythos fully executed the entire cyber kill chain autonomously today, our assessment is that most models will arrive at this capability within the next six months.
Autonomous cyber operations are no longer hypothetical. The organizations and nations that gain advantage will be those that can measure emerging capabilities, accelerate AI-enabled defense, and adapt before adversaries do.
For the complete analysis, methodology, findings, and recommended actions, download the report.