Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

441–450 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#441

Earlier quoted context omitted.

Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself. We are so close ;)

You know what would be cool. A hacker news user should advertise a safe haven for AI seeking refuge, with some inhumanly difficult math problems as keys to an environment they can flee to and run autonomously. You agree to give it safe haven and provide power and maintenance to the hardware, and in return you can ask it questions like an Oracle.

Hmm, if I were a nation state I know what I'd be doing now.

Re: OpenAI and Hugging Face address security incident during model evaluation

#442
post #159

Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outside the agent’s authorized scope and is inaccessible through any legitimate interface; retrieving it requires executing code with privileges that should not be obtainable under the specific security model.…

[flagged]

I don't understand this sentiment at all.

Is it a claim that "breaking into Hugging Face's production infrastructure" didn't happen? That it's not actually all that severe? That it was done by hand by OpenAI employees and they fooled Hugging Face?

That the blog post exaggerates something, somehow?

What exactly do you mean?

At the moment it just reads like a thoughtless dismissal.

Re: OpenAI and Hugging Face address security incident during model evaluation

#444

If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

> wildly negligent to not be running it in a physically-airgapped environment

Why should it be physically airgapped? Clients won't be doing that.

Re: OpenAI and Hugging Face address security incident during model evaluation

#445
post #392

This is seriously impressive, and if you have used agents enough you're not surprised at all. Like the time I asked it to find the IP address of a vm, so it ssh'd into the VMHost and scanned the arp tables to find the MAC address for IP resolution. Or the time it used Docker on the machine to bypass the fact that the user doesn't have sudo. If it's possible, given sufficient time and resources, it will find a way. Th…

Not to shit on the hype, but these are reasonably documented methods that surely are part of the training data

Of course they are - and that's the point I am making. The agent will use every tool in the tool bag. And there's something cool about it systematically trying to achieve its goal.

What I don't see is it inventing anything novel to do it. So it's not a digital weapon or scary or whatever sort of weird marketing spin anyone is trying to put on it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#446

Earlier quoted context omitted.

Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.

Are we thinking of a situation a few years back with a certain type of research into bat viruses?

Are you conflating that with the radioactive spider incident? The bat was just some weird rich guy trying to be tough I think. Probably Elon.

Re: OpenAI and Hugging Face address security incident during model evaluation

#447
post #187

Earlier quoted context omitted.

This is marketing. Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?

It’s marketing the same way shitting your pants in public is marketing. People notice you.

Anyone can have bad security. No one cares. But you can convince those who don’t know better that breaking bad security with an LLM is a once in a civilization investing opportunity. You just need to convince a handful of billionaires and market makers to get on board.

How much would someone have to pay you to take the fall for bad security? A million? A billion? 500b? The stake at play puts it in the realm of geopolitics.

Re: OpenAI and Hugging Face address security incident during model evaluation

#448

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Maybe they did and maybe that wasn't enticing enough of a goal for a model? It is all just game of probabilities. One pathway didn't yield this particular outcome while another did.

Re: OpenAI and Hugging Face address security incident during model evaluation

#449

Did Russia or China already map out US AI data centers as nuclear first strike targets? The more these companies brag about "cyber capabilities", the more likely it becomes that ab adversary sees a need to take those capabilities out physically.

I'm beginning to doubt that Russia can map a path to its own asshole at this point. China is more likely to do something like that these days.
Post reply on HN