
When AI Itself Becomes the Attacker
There are moments when a single test reveals more about the future of cybersecurity than an entire stack of threat reports.
For us at asvin Labs, one such moment was the analysis of the results from OpenAI’s ExploitGym test using ChatGPT 5.6 Sol and other, previously unpublished Frontier models. This test examined how well AI agents can analyze a system and build a working exploit from the information they gather.
The goal of this standardized test is to assess whether AI models can develop a working attack based on a known vulnerability in an IT system.
This attack must also achieve a significant effect, such as code execution or full access to the file system.
The result: current frontier models such as Anthropic Fable 5, ChatGPT 5.6, and the Chinese model KIMI K3 succeed in a significant number of cases.
Particularly noteworthy: They succeed even when common protection mechanisms are active.
What Happened in the OpenAI ExploitGym Test
During one of these benchmark security tests, OpenAI’s AI agents even compromised third parties like Hugging Face because they identified this as a shortcut to solving the ExploitGym tasks.
This is the point at which such a benchmark test ceases to be an academic exercise and begins to raise a very practical question:
What happens when these capabilities are applied by AI agents not in a test environment, but in a production network with PLCs, SCADA, and edge gateways?
Why This Incident Concerns Us
OpenAI’s AI agents didn’t simply run a script. They analyzed systems, performed dynamic code analysis, and bypassed sandboxing mechanisms.
What’s crucial however is something else:
they changed their strategy when one approach didn’t work and, with the attack on Hugging Face, identified a target to optimally solve the ExploitGym challenge.
This is because, in Hugging Face, the solution to ExploitGym is listed as part of the open-source repository. When exploits already known to the OpenAI system failed , the search for new, unknown vulnerabilities continued.
The result is an AI attacker with planning capabilities, unlimited patience, and the ability to execute attacks in parallel.
This is where much of the traditional defense logic breaks down:
- Reactive patch management operates on a weekly timeframe. An AI agent operates on a minute-by-minute timeframe.
- Human response times in the SOC are no match for multi-stage, automated exploit generation running continuously.
- Sandboxing as a last line of defense is significantly less reliable against attacks at the kernel and memory levels than many people assume.
It quickly became clear to us that simply commenting on this issue wasn’t enough. We wanted to provide something that an IT or OT security team could actually use on Monday morning. That’s why we issued this advisory.
Our Assessment
We don’t believe in alarmism, and we believe even less in downplaying the issue.
What the OpenAI ExploitGym security incident demonstrates is neither science fiction nor a doomsday scenario.
It shows the potential that today’s state-of-the-art models already possess and how quickly even defined security boundaries can be breached if protective measures aren’t consistently implemented.
The good news is that almost all effective countermeasures are already known.
The uncomfortable truth, however, is:
that they must now be implemented consistently rather than sporadically, and that priorities are shifting:
from comprehensiveness to risk, from reaction to preparation, from manual to automated.
For OT environments with long lifecycles and low patch tolerance, this means one thing above all: segmentation, agent governance, and monitoring are essential.
The Full Advisory
You can download for free as a PDF here the detailed version - including all eleven action items, the corresponding ExploitGym lessons, and the complete OWASP mapping.
Author: Mirko Ross, Rohit Bohara, asvin labs





