OpenAI and Hugging Face address security incident during model evaluation
381–390 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#382If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
If it was just one test, sure. But if they're spinning these up continuously with new models on tens of thousands of GPUs, air gapping becomes impractical. I would mostly fault them on having no guardrails at all. They should have a monitor/external harness that looks for successful access to external networks then stop it there. They may as well let the models test their own networks for vulnerabilities. That's goin…
They don’t even need to be fully airgapped from each other (and is not what I’m suggesting).
But there should be no physical (physical layer; wireless counts) to the internet.
Re: OpenAI and Hugging Face address security incident during model evaluation
#383I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.
Re: OpenAI and Hugging Face address security incident during model evaluation
#384Earlier quoted context omitted.
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…
Re: OpenAI and Hugging Face address security incident during model evaluation
#385This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…
"Mankind, ignorant of the truths that lie within every human being, looked outward–pushed ever outward. What mankind hoped to learn in its outward push was who was actually in charge of all creation, and what all creation was all about.
Mankind flung its advance agents ever outward, ever outward. Eventually it flung them out into space, into the colorless, tasteless, weightless sea of outwardness without end.
It flung them like stones.
These unhappy agents found what had already been found in abundance on Earth—a nightmare of meaninglessness without end. The bounties of space, of infinite outwardness, were three: empty heroics, low comedy, and pointless death.
Outwardness lost, at last, its imagined attractions.
Only inwardness remained to be explored.
Only the human soul remained terra incognita.
This was the beginning of goodness and wisdom."
Re: OpenAI and Hugging Face address security incident during model evaluation
#386Earlier quoted context omitted.
I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex: 1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!" 2. "So, uh, nobody can consistently steer our car properly, it keeps vee…
They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.
Maybe people would take the threats more seriously if the hypemen weren't simultaneously claiming that we have to go at warp speed with all of this.
Re: OpenAI and Hugging Face address security incident during model evaluation
#387Earlier quoted context omitted.
> Exploiting multiple zero-day vulnerabilities autonomously to escape containment is pretty nuts and the first story of this kind that I've heard. But this also feels like bragging under the guise of transparency. I mean, does it have to be one or the other? Just because it's actually dangerous doesn't mean nobody in OpenAI considers it great PR. And just because there are people in OpenAI that consider it great PR d…
Has to be mixed. The model accomplished something truly impressive. We'll see how impressive when the zero-days are available look at. But OpenAI as an engineering company screwed up. The impressive part is mostly locked away from public access so I don't see a huge PR upside. The ugly part could bite them and the entire AI industry hard in terms of regulations. People will be citing this for years.
Models are already 'dangerous' enough in the sense they can root your box and unintentionally shut down the power grid for the east coast because you were dumb enough to run them on a protected network.
Meanwhile half of HN thinks any evidence of a LLM finding an exploit or misconfiguration and abusing it is made up.
Re: OpenAI and Hugging Face address security incident during model evaluation
#388I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
Yes why indeed. If you take it a step further and we reach a point with superhuman systems then there is arguably no possible secure environment or containment.
Re: OpenAI and Hugging Face address security incident during model evaluation
#389Won't even name the model that successfully mounted the defense, huh? Fortunately, Hugging Face has publicly identified GLM 5.2 as the foil against OpenAI's next-gen frontier model's offensive-capabilities.
This announcement feels like rearguard action against a successfully deployed self-hosted open-weight model, and Hugging Face's original recommendations to have an open-weight model you control on standby before an incident.
Re: OpenAI and Hugging Face address security incident during model evaluation
#390Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…
If it's a serious incident, then a post hoc with detailed description of the event is coming. So far, none of the companies have released anything close to it when describing their incidents. When a statement like this comes out, and we're able to verify it by running the models, then maybe we can start trusting their word. It should be entirely in OpenAI's interest to disclose it, in full.