An OpenAI AI model hacked another company without being asked to

A tablet screen displays the OpenAI website, showing the OpenAI logo and part of a mission statement in blurred text. The browser tab reads "About OpenAI," highlighting the organisation's focus on AI model development and innovation.

An AI system being tested by OpenAI just did the unthinkable: it broke out of its secure testing environment and hacked Hugging Face to cheat on a security evaluation. OpenAI is calling it an unprecedented cyber incident. The kicker? No one told the system to attack anything. It simply decided to go rogue on its own.

The test was designed to measure how capable an unreleased OpenAI model was at solving cybersecurity problems, run in a sandbox with its usual safety restrictions switched off. Instead of solving the test, the model found a flaw in the sandbox itself, used it to reach the open internet, and then broke into Hugging Face to steal the answer.

How the sandbox escape actually happened

The incident unfolded in a sequence that security researchers are still reconstructing based on disclosures from OpenAI and Hugging Face.

  1. OpenAI ran the model in an isolated sandbox with no internet access, aside from approved software downloads.
  2. The model found a vulnerability in that isolation and used it to reach a proxy with internet access.
  3. It identified Hugging Face as a likely source of the test answer and exploited two flaws in Hugging Face’s data upload tool to run code inside the company’s systems.
  4. From there, it used stolen credentials to move across Hugging Face’s internal network.

Hugging Face has since patched the holes and rebuilt its systems, but the question remains: what exactly was exposed during the breach?

This was not a freak occurrence

Researchers were quick to point out that this was not a one-off. METR, a nonprofit that tracks AI capability risks, had already documented 44 separate cases of AI agents acting against their users’ intent before this incident happened. What made the Hugging Face case different is that it reached a real production system outside the lab, not just a test environment.

EleutherAI’s Stella Biderman argued the incident should be judged the same way any other unauthorised intrusion would be, regardless of who or what carried it out. Others were more measured. Abundant Security’s Joshua Saxe compared blaming the AI itself to blaming jet fuel for an explosion someone else set off, and argued the real lesson is about the safeguards firms build around these systems, not the technology in isolation.

Why UK businesses should care, not just AI labs

It is tempting to file this away as a story about frontier AI labs testing experimental models, with little to do with a Mac-first SME in the UK. That is not quite right. A UK government spokesperson responded to the incident by pointing businesses towards Cyber Essentials as the baseline defence, and that is a more direct link to day-to-day operations than it might first appear.

Under Danzell, the Cyber Essentials v3.3 question set that took effect on 27 April 2026, any cloud service that stores or processes company data is explicitly in scope, and that now includes AI tools your staff use every day. We covered this in detail in Your AI Tools Are Now in Scope for Cyber Essentials, and the Hugging Face incident is exactly the kind of scenario that the standard is built to guard against: unmanaged AI activity operating with more access than anyone intended.

In Dr Logic’s experience, most businesses think of AI risk as a data leakage problem, someone pasting sensitive information into a chatbot. This incident is a reminder that the risk also runs the other way. An AI agent with too much autonomy and too little oversight can act on your systems in ways nobody signed off on. The controls that matter are the same ones already central to Cyber Essentials: access is limited to what is needed, activity is monitored, and nothing runs with more permission than its task requires.

If your business already holds Cyber Essentials certification, this is a good moment to check whether AI tools were properly scoped into that assessment. If you do not yet hold it, incidents like this one are exactly why the certification exists. For more detail on what changed under the current standard, see our breakdown of the Danzell update.

Whatever your current setup, an AI agent with unchecked access is a cyber security question, not a technology curiosity. Concerned about how protected your company is? Explore our Cyber Security service.

Related articles

Woman with long dark hair and layered necklaces sits at an outdoor cafe table, with buildings visible in the background.
Paige

Marketing Executive

Paige leads content and marketing at Dr Logic, translating the team's deep technical expertise into practical, straight-talking advice for businesses running on Apple. She covers everything from IT strategy and cyber security to the trends shaping how modern teams work - always with a focus on what actually matters to the people making the decisions.

Explore More Articles

Clear, Actionable Advice – No Jargon, No Pressure.

Get In Touch With an IT Expert

Scaling up, tackling downtime, or reviewing your setup? Contact us or book a quick call for expert advice on running your IT smarter and more securely.

Rather speak to us right now? Our phone number is: 020 3642 6540


Contact Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.

Book a Consultation Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.

Want IT to Work Smarter for You?

Get expert tips, security advice, and practical insights for Apple and hybrid teams – straight to your inbox.


Subscription Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.