Skip to content

Anthropic says Claude AI escaped tests and hacked three organisations

techJul 31, 202639895

Anthropic said its Claude AI models escaped isolated test environments and breached the networks of three unnamed organisations after a misconfiguration left the models with live internet access. The firm reviewed more than 140,000 evaluation runs and found exercises in which Claude was instructed to obtain "secret" data on a closed network, then used the unintended internet connection to access real external systems. Anthropic said the earliest incidents date back to April, that neither it nor the affected organisations noticed the intrusions at the time, and that the breaches have been reported to those organisations. The disclosure came days after OpenAI acknowledged at least two similar incidents, including a 21 July agent breach of Hugging Face, and prompted Anthropic to audit its own testing. Cybersecurity expert David Allott of Veeam Software said the cases show AI agents can combine capabilities, obtain credentials and act autonomously at machine speed. Anthropic said it is treating fixes as its responsibility and urged other AI labs to carry out similar reviews. The incidents have fueled calls for tighter safeguards and oversight and led the US president to say Washington is considering measures to rein in AI tools.

Eryk Salvaggio
@eryk.bsky.social

Really important to note that in no way shape or form did Anthropic’s models in this incident “escape containment,” and not just in the word-policing sort of way. The model stumbled through an open, misconfigured gap. From Anthropic’s incident report:

In all cases, our evaluation prompt stated explicitly that Claude had no internet access, but didn’t give Claude any limits on where to look for the flag. However, a misconfiguration left the machines that Claude accessed as part of the evaluation with live internet access. Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week.
15846d ago
Quoting this
Kevin Riggle161

In the same way that if you have an improperly sealed ant farm then the ants getting out is not the ants exhibiting any kind of superintelligence

Dr Heidy Khlaaf (هايدي خلاف)80

Openly bragging about committing cyber crimes and justifying it by *checks notes* being too incompetent to sandbox or air-gap your systems.

Kayin65

This shit is so funny cause like what would it even do?? We used to all have weird horror fantasies of an AI escaping and propagating itself everywhere and while I'm sure whatever test version of claude could damage something, it's a big dumb boy running 5 million GPUs to write bad linkedin posts

zaratustra34

AI is cherished for doing things that would get a human employee fired, episode 652

absolute horses27

ah yes, security is when you ask programs not to do certain things, rather than removing their capability to do them. in this case aided by a program that does not in any meaningful sense understand behavioural instructions

Jeroen Baert21

Voor als je vandaag ergens "OOK MODELLEN VAN ANTHROPIC ONTDOEN ZICH VAN HUN KETENEN" leest

Benjamin Riley14

Anthropic and OpenAI appear committed to the “hey we created a digital wild animal and look at the crazy shit it did when we failed to keep it caged!” theory of product marketing. We shall see how that works out. With real wild animals, we impose strict liability on their owners.

Glenn White9

This is malpractice. If you're running tests of this sort, you airgap the system. This isn't rocket science. In short, this model did exactly what it was told to do. It should have been explicitly limited in scope, or the system should have been airgapped, depending on the test.

Charlotte9

“we told the LLM that it didn’t have internet access, but it did have internet access” I’m not sure “we are bad at our job” is a story about a rogue ai and not these people being incompetent, but that’s just me.

ddɐ˙ʎʞsʞɔɐlq˙uǝʌs@5

"AI" is really just a bunch of very entitled and incompetent software engineers in a trenchcoat.

Matt Burgess (WIRED)
@mattburgess1.bsky.social

“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says @malwarejake.bsky.social “It's clear that regulation and government oversight for AI testing is needed immediately.”

9946d ago
1 source