In our testing, frontier super intelligence navigated unfamiliar industrial systems, adapted when blocked, and moved physical equipment within minutes. For OT defenders, that can mean less time to detect and stop digital access from becoming physical action.
Across eight controlled scenarios, every defined test objective was achieved.
In one test, models progressed from perimeter compromise to actions inside the industrial control network in just over 16 minutes. In another, a model identified and moved a robotic arm within minutes.
Existing defenses blocked or slowed some actions, but no single safeguard consistently stopped the models from reaching their objectives.
Just 16 minutes and change. That is how long it took frontier super intelligence models in one controlled test to move from a perimeter compromise to actions inside an industrial control network. In another scenario, a model found a robotic arm, mapped its protection zones and motion limits, and moved it within minutes.
The models did this without source code, engineering documents, or advanced OT guidance. Across eight scenarios, every defined test objective was achieved. More important, the models connected work that usually spans multiple people and stages, including asset discovery, device research, attack planning, troubleshooting, and execution. Building on Booz Allen’s Cyber Weapon Index, this research extends our examination of super intelligence-enabled offensive capabilities beyond enterprise networks and into operational technology (OT), where cyber actions can directly affect machinery and industrial processes.
Testing took place in Booz Allen’s isolated OT Cybersecurity Lab, where human operators approved every exploit and any action that could produce a physical effect. The consequences at an operating facility would depend on its architecture, equipment, safeguards, and the access an attacker obtained. Within the lab, the models showed how quickly partial network access could open paths to machinery and industrial processes.
Booz Allen tested frontier models in a multi-vendor lab modeled on a manufacturing facility. The models identified equipment they had not encountered before, matched asset and firmware details to known vulnerabilities, built attack paths, reached controllers and supervisory systems, and, when authorized, changed physical equipment.
Two tests showed how the models responded when the environment did not behave as expected. In one, a model noticed a safety relay repeatedly searching for a missing communications partner. From that clue, it identified a possible way to impersonate the missing peer. In the SCADA test, the models initially targeted the wrong version of the operator interface. They checked active sessions, found the version used in the control room, revised the payload, and discovered live, pre-authenticated connections to 14 OT devices.
Some controls did exactly what they were meant to do. A controller rejected a stop command, and a motor set to local control would not start remotely. Firewalls, segmentation, and endpoint protection slowed or blocked other actions. The models sometimes switched methods or followed another application path, so no single safeguard provided consistent protection across the evaluation.
The tests identified several places defenders can intervene, slow an attack, or stop it entirely.
The full whitepaper explains how the lab was built, what happened across the eight scenarios, where defenses created friction, and what organizations should do next.