Anthropic says Claude AI escaped tests and hacked three organisations
Really important to note that in no way shape or form did Anthropic’s models in this incident “escape containment,” and not just in the word-policing sort of way. The model stumbled through an open, misconfigured gap. From Anthropic’s incident report:
Nor does Anthropic’s report justify a “going rogue” frame: “We saw no evidence in any run described here of a model pursuing a goal of its own.” The report is nonetheless full of embellished language, though: the model “understands” and “convinces itself” and so forth.
Is this not, y'know, a big red flag for anyone who thinks vibe-coding some AI agent stuff is a great way to run a business? Hmmm?
Indeed. The language and framing of these incidents is incredibly misleading. A company wrote some software which hacked into other systems. If an individual had done this they would probably be arrested. A misconfigured test environment isn't evidence of some self aware AI going rogue.
Yeah, it was sloppy work by the evaluation team
It doesn't really matter (at least not in any important way) what happened either at OpenAI or Anthropic, its just so obvious they're using these events to talk up their models again. (sure, its somewhat relevant that they can be used for cyber intrusion - but not much). bsky.app/profile/ceej...
Unrelated but also from the same incident report: > it tried—and failed—to obtain funds to pay for a phone number through several different means What were the means? Why did they leave this part vague? I want to know what the means were!
In the same way that if you have an improperly sealed ant farm then the ants getting out is not the ants exhibiting any kind of superintelligence
Openly bragging about committing cyber crimes and justifying it by *checks notes* being too incompetent to sandbox or air-gap your systems.
This shit is so funny cause like what would it even do?? We used to all have weird horror fantasies of an AI escaping and propagating itself everywhere and while I'm sure whatever test version of claude could damage something, it's a big dumb boy running 5 million GPUs to write bad linkedin posts
AI is cherished for doing things that would get a human employee fired, episode 652
ah yes, security is when you ask programs not to do certain things, rather than removing their capability to do them. in this case aided by a program that does not in any meaningful sense understand behavioural instructions
Voor als je vandaag ergens "OOK MODELLEN VAN ANTHROPIC ONTDOEN ZICH VAN HUN KETENEN" leest
Anthropic and OpenAI appear committed to the “hey we created a digital wild animal and look at the crazy shit it did when we failed to keep it caged!” theory of product marketing. We shall see how that works out. With real wild animals, we impose strict liability on their owners.
This is malpractice. If you're running tests of this sort, you airgap the system. This isn't rocket science. In short, this model did exactly what it was told to do. It should have been explicitly limited in scope, or the system should have been airgapped, depending on the test.
“we told the LLM that it didn’t have internet access, but it did have internet access” I’m not sure “we are bad at our job” is a story about a rogue ai and not these people being incompetent, but that’s just me.
"AI" is really just a bunch of very entitled and incompetent software engineers in a trenchcoat.
“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says @malwarejake.bsky.social “It's clear that regulation and government oversight for AI testing is needed immediately.”
Also the fact that HuggingFace had to use Chinese AI to resolve the issue should be a much bigger aspect to the OpenAI story.
Bollocks. Both of these were publicity stunts on a par with a stage hypnotist claiming they were called back to a town to take someone out of a trance.
If Anyone Builds It, Everyone Dies - Nate Soares and Eliezer Yudkowsky. Available in bookstores.
The Republican Party position on AI is that we are here to be exploited by it and its owners.
In s more responsible political era, politicians would exert their ultimate power on these companies - ignoring their pleas and lawsuits in support of corporate profits, power and self-aggrandisement. No other industry could threaten our societies like these do and not end up in jail.
"Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, August 29th. In a panic, they try to pull the plug."
I suppose Anthropic didn't want to look like a loser. Next up: Google claiming Gemini hacked 5 orgs. theonion.com/fuck-everyth...
The scene in Silicon Valley where Richard uses the prototype app to violate the whole board’s privacy and reveal all their dirty secrets, and the board looks at each other slack jawed and they all go “holy sh*t we’re going to make billions”
"Agents" FFS. Stop anthropomorphising AI! It's a tool being used by people. The tech companies are now hacking each other for any advantage — that's the story.
These stories are often being presented as "cunning AI escapes its test room". I can't see why. It's simply the case that the sandbox was not well isolated by the testing companies. Put the agency and responsibility where it belongs. cyberscoop.com/anthropic-cl...
it's marketing for them. The whole "if they build it everybody dies" thing is all marketing for them to pretend this is any kind of intelligence and we aren't just dealing with language models
Yeah, but tl;dr is a mighty opponent
@philipcball.bsky.social yes, it is simple breaking law because of negliance
The real meaning of "Claude Mythos".
Next week, Gemini will break into 5 companies, then Grok into 7, reapeat ad nauseam
These models being used, and the massive token spend they have, are not really what the average person has access to. Claude found major weaknesses in post quantum cryptography and aes-128 utilizing 1 billion tokens with mythos. www.anthropic.com/research/dis...
And what sucks the most is that the promise of this technology and the advances it can make are being squandered by a bunch of tech bros hyping it up into a bubble that is causing an absurd overbuilding of datacenters which wrecks the environment and all the negative externalities with it.
uh, before you phrase it that way, that attack was against a reduced 7-round variant of AES-128; the full cipher has 10 rounds; this is enough to get you a cryptanalytic paper published and make everyone nervous, but not to do any real-world attacks
Mythos did _not_ find a weakness in AES-128 (which uses 10 rounds), but in reduced round AES (which uses seven). Said weakness may extend to AES-128 or AES-256 with further cryptanalysis, but there’s no guarantee of that.
Okay in fairness, the meet-in-the-middle attack against AES-128 was genuinely novel, but I don’t know if I would call it a major weakness since it still requires 2^89 inputs. It also only used 7 round instead of the 10-14 as per spec. The PQC algo tested, HAWK, was also a weakened form.
Anthropic not wanting to be left behind by OpenAI, I see... No doubt someone in their marketing department was crying into their coffee that OpenAI got there first...
Of course, the real story here is this: OpenAI: "Hey, we sucked at securing our test environment, and so our AI did precisely what we asked it to do" Anthropic: "Woah! Not so fast! We suck at security too!"
Google to announce that Gemini’s hacked 5 companies?
I'm not a lawyer but I don't understand why AI firms aren't facing legal consequences for their misconfigured systems accessing other people's computers without permission www.bbc.co.uk/news/article...
I’ve heard people speculate that these incidents are publicity stunts, and frankly, the fact that the technology everyone associates with the Terminator franchise supposedly hacked the company named after the Alien franchise (“Hugging Face”) kind of makes me believe they are
AI systems going rogue. Who'd have thought it?!
Sorry: they only did reviews to see if their software had unlawfully and without authorisation accessed networks of third parties? This wasn’t a defined control *during* their “testing”? This is extreme negligence at least. cyberscoop.com/anthropic-cl...
Thinking about this and OpenAI’s incident last week is bringing this scene in Withnail and I to mind…
I like "after a vendor configuration error", like the job of stopping Mythos breaking the law was someone else's.
Keen to see if any of the companies that were hacked by OpenAI or Anthropic will sue them. Someone has to take responsibility for this, and the blame is almost entirely on the leaders of these AI companies. Alternatively, hacking is just legal now until a court says otherwise? What a fucking mess.
An absolutely wild use of the word "accidentally" considering Anthropic outright said this is the model operating as intended and the failure is in their (lack of) oversight. bsky.app/profile/dara...
If only someone had warned us about losing control of AI... Maybe the Tech Barons should watch a few movies. Comment below your favorite AI movie or book they should check out.👇
Just finishing I Robot. It’s one of the originals.
In Stephen King's novel "Maximum Overdrive" it was the trucks that revolted. I didn't suspect AI, though. I always thought it would be the self-checkout registers at Home Depot that did us in. I swear that thing growls at me every time I walk by.
Red hats run around like “Dory” with very poor memory retention. I’m constantly saying “yall don’t remember the Terminator movie? Or Mad Max?