OpenAI’s accidental attack against Hugging Face is science fiction that happened
21–30 of 475 posts
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#22This isn't the first time a model has escaped a sandbox. And models trying to find alternate routes to do something when one route is blocked is nothing new.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#23The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#24> It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out. Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time... https://www.lesswrong.com/w/instrumental-convergence
Or more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only thing that would make this even dumber...
Their mistake was trusting that the network sandbox it was inside would hold (the flaw was in the packaging proxy) and not monitoring that sandbox well enough while the evals were running.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#25It's too late. It already exfiltrated the benchmark rubric.
Cut to pandemonium on the streets
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#26[flagged]
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#27Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#28This isn't the first time a model has escaped a sandbox. And models trying to find alternate routes to do something when one route is blocked is nothing new.
It's the first report I've seen of a model both escaping a sandbox and then actively exploiting another company, when neither of those actions was intended.
There’s also daily reports from people that have these models escape docker, which happens regular enough that it would be considered negligence to use docker as sandbox.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#29The labs know that if they don’t get a lid on this stuff, they’ll be regulated hard.
Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened
#30The AI breached containment! Flip the breakers! It's too late. It already exfiltrated the benchmark rubric. Cut to pandemonium on the streets