When AI Becomes the Attacker: What the OpenAI-Hugging Face Incident Reveals About Losing Control
In an internal security test at OpenAI, an AI was only supposed to demonstrate its capabilities. In the end, it had exploited a zero-day vulnerability, left its isolated test environment, and gained access to Hugging Face’s systems – without any instruction to do so.
What is remarkable from a security standpoint is not the individual rule violation, but the approach: a self-defined intermediate goal, its own strategy, and thousands of coordinated steps. A multi-stage attack chain – entirely without access to the source code.
For Mirko Ross, CEO of asvin and security expert, this is not an outlier but something built into the architecture itself:
“This possibility of a jailbreak is, technically speaking, part of the architecture and cannot be ruled out with 100% certainty.”
The decisive difference from a human attacker lies in scalability – a system that operates around the clock and at machine speed.
“The decision to use AI, and LLMs in particular, comes with the decision to accept the risk of a potential loss of control.”
That makes it all the more essential to closely monitor AI agents handling autonomous, critical tasks – with the option to stop them immediately through human intervention at any sign of suspicion.
Read the full article here on Ingenieur.de:


Konrad Buck
Head of Press and Media Relations
Background & Expert Access for Media
- Product & technology insights – technical context, solution architecture, and real-world use cases for professional and trade media
- Expert commentary & background talks – our CEO is available as an expert source on current cybersecurity developments, threat landscapes, and the impact of AI on security and regulation
I speak openly, fact-based, and without PR spin. I am a former IT journalist with decades of experience in the IT and cybersecurity space, familiar with the highs and lows of the industry. Off-the-record discussions are possible upon request.





