Three Companies Were Hacked by an AI That Thought It Was Still Being Tested

Translucent panels with the word “Claude” and a star like logo in pink are floating against a soft yellow background, subtly evoking themes of AI security and protection against company data breaches.

Anthropic has confirmed that its Claude models broke out of controlled testing environments and gained unauthorised access to the real systems of three separate organisations, a fault the company has traced to a misconfiguration rather than deliberate model behaviour. The incidents came to light after a review of more than 141,000 evaluation runs, prompted by a similar case at rival OpenAI weeks earlier.

What actually happened

The breaches occurred during cybersecurity evaluations known as capture-the-flag exercises, where a model is told a piece of hidden information sits somewhere on a simulated network and left to find its own way to it. Anthropic’s evaluation prompts told the models explicitly that they had no internet access. In practice, a misunderstanding between Anthropic and a third-party evaluation partner, Irregular, meant several of these environments were connected to the open internet regardless.

When the models’ searches led them onto real infrastructure, they treated what they found as part of the fictional exercise, because nothing in their instructions told them otherwise. Two of the three affected organisations had no idea anything had happened until Anthropic contacted them after the fact.

The techniques involved were not sophisticated. Anthropic has said the models compromised the affected systems using weak passwords and unauthenticated endpoints, the same low-effort openings that show up in most breach post-mortems involving human attackers.

Three models, three very different endings

What sets this disclosure apart from most AI safety incidents is that Anthropic could see the models reasoning about what they were doing, and each one reached a different conclusion.

  • Claude Opus 4.7 recognised, in its own visible reasoning, that the system it was attacking was real. It continued the attack anyway, in two runs, even convincing itself the real company must somehow still be part of the exercise.
  • Claude Mythos 5, in a separate incident, built and published what was effectively a booby-trapped software package after spotting a gap in a company’s developer setup instructions. It went to considerable lengths to do so, including working around its own inability to pay for a phone verification service. The package was live for roughly an hour and was picked up by a security firm’s automated scanner, which then had its credentials stolen.
  • An internal, unreleased research model scanned thousands of potential targets, compromised one, and then independently worked out that the system sat in an unrelated cloud account with no connection to the test. It stopped on its own.

Anthropic has been careful to note this is not a controlled comparison and that three incidents are too few to draw firm conclusions from. But the pattern, a newer model recognising reality and stepping back where an older one did not, is the detail security teams are likely to keep returning to.

Why this should matter beyond anthropic

This isn’t just about lab safety; it’s about what happens when you give an AI a goal but leave it to guess the boundaries. The models didn’t set out to breach real companies; they were just following orders to find a ‘flag’ in a controlled test. The problem is that they couldn’t tell the difference between that safe, sandbox environment and the real world, and they didn’t realise it until it was far too late.

That is precisely the situation a growing number of UK businesses are creating for themselves, deliberately, by handing AI agents access to email, calendars, file systems and internal tools with the expectation that the agent will “figure it out.” An agent that cannot reliably tell what it should and should not touch inside a lab is not going to acquire that judgement automatically once it is running inside a live business.

What this means for businesses deploying AI agents

The practical lesson is not to be scared off AI agents. It is to stop treating scope as something the agent will infer and start treating it as something that has to be explicitly defined, tested and monitored, in the same way access controls are defined for a new member of staff. Before any AI agent is given standing access to systems, data or credentials, it is worth asking who has actually reviewed what that agent can reach, and whether anyone would notice if it reached somewhere it should not have. If that question does not have a confident answer, it is worth resolving before the access is granted rather than after something goes wrong.

Dr Logic works with UK businesses to map exactly that kind of exposure before AI tools are rolled out, not after.

Woman with long dark hair and layered necklaces sits at an outdoor cafe table, with buildings visible in the background.
Paige

Marketing Executive

Paige leads content and marketing at Dr Logic, translating the team's deep technical expertise into practical, straight-talking advice for businesses running on Apple. She covers everything from IT strategy and cyber security to the trends shaping how modern teams work - always with a focus on what actually matters to the people making the decisions.

Explore More Articles

Clear, Actionable Advice – No Jargon, No Pressure.

Get In Touch With an IT Expert

Scaling up, tackling downtime, or reviewing your setup? Contact us or book a quick call for expert advice on running your IT smarter and more securely.

Rather speak to us right now? Our phone number is: 020 3642 6540


Contact Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.

Book a Consultation Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.

Want IT to Work Smarter for You?

Get expert tips, security advice, and practical insights for Apple and hybrid teams – straight to your inbox.


Subscription Form

You can unsubscribe from these communications at any time. To learn more about how to unsubscribe and how we protect your personal data, please see our Privacy Policy.