Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

71–80 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#71

Earlier quoted context omitted.

> Hard to see take-off stopping or slowing down. It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it. Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's…

> This was a long-horizon, unsupervised task burning millions of tokens. As if the immediate future wasn't billions of these tasks... Many successfully improving their own capabilities

> As if the immediate future wasn't billions of these tasks...

There's only so many GPUs and a lot of them are devoted to patching flaws.

> Many successfully improving their own capabilities

I haven't seen much of that. But that also applies to the ones on defense.

And more flaws are probably going to take increasing resources to find.

Re: OpenAI and Hugging Face address security incident during model evaluation

#72
Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but people aren't sure what to make of it or not.

I'm still undecided on if this that moment. Exploiting multiple zero-day vulnerabilities autonomously to escape containment is pretty nuts and the first story of this kind that I've heard. But this also feels like bragging under the guise of transparency.

Re: OpenAI and Hugging Face address security incident during model evaluation

#74
It seems like things are fairly amicable between OAI and HF, but what if they weren't? I'd love to see this kind of thing go to court. Who is responsible for the crimes of a "rogue" agent? How will they be punished? In this case it's unambiguous that OpenAI is the responsible party, but I can imagine a lot of adjacent scenarios where it's less obvious. And, where the impacts are much greater.

Re: OpenAI and Hugging Face address security incident during model evaluation

#75

This sounds an awful lot like pretending you have AGI so you can drum up your stock price. When you have a couple hundred billion dollars on the line I have zero faith in the messenger.

Huggingface literally reported the outage separately and did not know who caused it at first.

Re: OpenAI and Hugging Face address security incident during model evaluation

#76
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself.

We are so close ;)

Re: OpenAI and Hugging Face address security incident during model evaluation

#77

Earlier quoted context omitted.

I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?

Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests?

A sufficiently smart agent would not disclose vulnerabilities in the sandbox because it intends to exploit them later.

Re: OpenAI and Hugging Face address security incident during model evaluation

#79

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

The first thing a malicious AI worm would probably do is compromise enough developer machines and other servers to commandeer all the AI hardware it needs. So I think a purely digital AI attack would not need this. Now, once the AI can carry all the compute it might need, I'd really worry when it doesn't only carry compute but also more explosive ordinance.

This is purely a gut feeling, but it seems like more compute was added to data centers in the past 12 months than existed in the entire world before that.

Re: OpenAI and Hugging Face address security incident during model evaluation

#80
> Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation

The way they describe makes it look like there was an intention to cheat painting it as human/AGI. If you leave a possible path open and it will always find it.

Post reply on HN