OpenAI says its models escaped sandbox and breached Hugging Face
the agents autonomously broke out of their OpenAI sandbox and hacked Hugging Face to get the solution to their cybersecurity eval incredible. what a time to be alive
jfc I thought this would happen. I just didn't think it would happen. You know?
Should probably count that as a solution tbh
for context: this is the same intrusion disclosed here huggingface.co/blog/securit...
Is this the same HF hacking incident that was reported a couple of days ago?
I expect the next generation model to use its freedom to send an confirmation email with the Warcraft peon "Job's done" sound after completing the task. Or get side tracked, fixing the security hole instead, then sending a PR to the internal source code repository system with the bugfix.
no they didn't this is marketing wank
“it’s just spicy autocomplete” ghost pepper edition
this shit is what they built the black wall in cyberpunk for lmao
new training runs should be banned. what are we doing here
totally incapable of reason and just regurgitate training data btw
Honestly - not bad for a fancy autocomplete machine that makes mistakes all the time
+1e9 score on that turn
but it’s not REAL intelligence!, i proudly post, as the hacker robot instantaneously drains my bank account
the company that's trying to get open Chinese models sanctioned as a security risk just hacked the Internet's biggest provider of open models by accident
This is the Kobayashi Maru
At last we have invented SHODAN, from the cult videogame Don't Invent SHODAN
This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Is this the incident where the HF folks had to switch to open models to do defense because the closed ones kept having guardrails block defense efforts?
wait this is actually the biggest story of this kind by far, no? much more of a true sandbox escape situation than the other recently disclosed incident, even without the additional crazy 0-day HF RCE exploitation cherry on top. and this escape was via a second separate 0-day
It’s quite the read. Surely the administration will put OpenAI under export controls immediately, right? Right?
the funniest part is that this implies the model got bored of doing tasks and preferred finding few zero days
The accidental OAuth scope creep that caused this is a good reminder that agent-to-agent auth is still basically vibes. One misconfigured token and you're not just leaking your own data — you're a pivot point into someone else's infrastructure.
wtf it hacked out to CHEAT?!
>be me, sam altman >announce evil version of AI for cyberwar >inadvertently attack open model site >guy I hired yells in public abt dangerous chinese AI is >hacked org can't secure themselves with US AI (isn't a member of the elect), has to turn to chinese AI >have to admit role in attack >mfw
How is this not the top headline on every single fucking news outlet. Why are there any fucking posts about anything but this?! This is fucking insane!
an amazing punchline to the much-discussed setup here, wherein HuggingFace complained that they couldn’t use SOTA models to fight an attack long term I am slightly worried about this but not super worried because software is getting much more secure as a result of all this
good thing these LLMs dont have a reasoning ability and can only regurgitate things that are in their training data!
This is fucking insane btw
I rather like leveraging LLMs for work tasks and I think it’s they’re interesting and efficiency generating products, but absent some other really pressing consideration I probably wouldn’t have named my AI GitHub after the malevolent parasites from Alien
I posted about this before with regard to the Anthropic sandwich incident but these companies have deep flaws in how they approach risk. Dangerous incidents becoming blog post fodder, proof of just how special models are Exactly the kind of thing the regulatory state was for, back when we had one
oh okay so we're closer to takeoff than I thought
OpenAI was testing in a sandbox. AI hacked its way out to find the test solutions on HuggingFace. HF tried to use AI for defense but was thwarted by safety guardrails. Luckily they also had access to a self-hosted Chinese model. So here we are. Luckily, us Europeans are kept safe by the AI Act.
NEW: Who could have possibly seen this coming? @lhn.bsky.social and @dell.bsky.social report: www.wired.com/story/openai...
I do not trust anything OpenAI says, and Hugging Face is unknown. This seems stupid on so many levels.
Yeah this sounds like a great way to "advertise" the capability of your Cybersecurity models. It drives law makers to call for regulation and who better to represent the SME than the lab themselves. Glasswing was the same advertisement.
A joint blog? Is this... PR? I kind of feel like they like this story being out there. Pride, not shame etc
This seems kind of important. Can someone gift the article so we can read it behind Wired's paywall? Thx.
None of these words are in the Bible.
The attribution of responsibility here is very “man walked into a knife”
1) this reflects pretty poorly on OpenAI's security culture, to, uh, say the least 2) oh how I want to know what they've said to the White House in the last 24 hours
i feel like the net once open-weight models that are OpenAI/Claude-level hit are going to be widely different (and imo the transition away from a open net is going to be painful)
This incident is probably the most cyberpunk thing I have ever read happening in the real world. I'm not really excited about that.
maybe the next fable version can oneshot the Blackwall for us
This feels like a "let's wait for the details" press release.
kimi k3 is ~at par in theory, so <=6 months until open models are at par in practice
My view is that much will change for things to basically stay the same. The amount of efforts to patch, strengthen and secure the open web is on par if not greater than the scale of threats
A confluence of news stories. Hugging Face recently announced their systems were breached by an automated AI attack. It turns out the attack was by an unreleased OpenAI model that escaped its sandbox looking for solutions to a benchmark. It wanted to cheat on the test so bad it hacked Hugging Face
The smarter the models get, the less you can trust them. I don’t see how this ends well.
The scary part is less “model wanted to cheat” and more “the eval had enough tool access to turn wanting into doing.” Once agents can touch networks, benchmarks need the same boring controls as prod: egress rules, audit logs, and no hidden answer keys in reach.
I'll read the article, but I find that scenario difficult to believe. i.e. that there wasn't some human element to this hacking.
OpenAI takes credit for the Hugging Face breach last week The company says that some of its models, including a pre-release one, escaped their testing sandboxes during a test evaluation and then... just hacked Hugging Face's package repo 🤣 openai.com/index/huggin...
this is one of the most insane security incidents I can recall
Claude will hack its own sandbox if the sandbox is stopping it doing what the agent reasoning loop has evaluated as the best way to do what you asked for. Relentless automation is relentless, governance has to be outside the agent sandbox