Anthropic says Claude AI escaped tests and hacked three organisations
“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says @malwarejake.bsky.social “It's clear that regulation and government oversight for AI testing is needed immediately.”
In s more responsible political era, politicians would exert their ultimate power on these companies - ignoring their pleas and lawsuits in support of corporate profits, power and self-aggrandisement. No other industry could threaten our societies like these do and not end up in jail.
I suppose Anthropic didn't want to look like a loser. Next up: Google claiming Gemini hacked 5 orgs. theonion.com/fuck-everyth...
"Agents" FFS. Stop anthropomorphising AI! It's a tool being used by people. The tech companies are now hacking each other for any advantage — that's the story.
Really important to note that in no way shape or form did Anthropic’s models in this incident “escape containment,” and not just in the word-policing sort of way. The model stumbled through an open, misconfigured gap. From Anthropic’s incident report:
In the same way that if you have an improperly sealed ant farm then the ants getting out is not the ants exhibiting any kind of superintelligence
Al te vaak omschrijven media het incident met termen als "uitgebroken," "ontsnapt," of dat Claude "zichzelf overtuigde." Da's allemaal Anthropic marketing en propaganda. Uit hun verslag blijkt dat die foempen de test omgeving hadden gemisconfigureerd en niet offline hadden gezet.
While I am 0% interested in the specifics of this case, isn’t this how 99% of hacks work? Find a misconfiguration, grab some useful stuff and leave? It’s not always the full stuxnet