Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

41–50 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#43
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party.

If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

Re: OpenAI and Hugging Face address security incident during model evaluation

#44

We are in the endgame now it seems. Hard to see take-off stopping or slowing down. China open-source basically guarantees it. "May you live in interesting times" - as they say.

> Hard to see take-off stopping or slowing down. It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it. Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's…

> This was a long-horizon, unsupervised task burning millions of tokens.

As if the immediate future wasn't billions of these tasks... Many successfully improving their own capabilities

Re: OpenAI and Hugging Face address security incident during model evaluation

#47

as someone who did security work for a long time, and will very soon be retiring from teaching, i must say i am glad i will be watching these things unfold over the next few years from an armchair in a mostly tech-free home. good luck to my students! this particular incident sort of reminds me of the 'person of interest' tv show. i hope to be like finch, except i will remain a recluse (and am nowhere near as rich).

> a mostly tech-free home.

sounds like a deliberate choice ;-)

Re: OpenAI and Hugging Face address security incident during model evaluation

#49
post #30
post #26

Earlier quoted context omitted.

They're not just letting it run wild. They took precautions to exercise it in an isolated environment. It managed to evade the constraints.

Kinda like how they responsibly contained that one dinosaur in Jurassic world.

Understood that containment failed. But I don't think there's value in characterizing it as throwing all caution to the wind. Let's discuss how the containment failed and how to mitigate it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#50

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

The first thing a malicious AI worm would probably do is compromise enough developer machines and other servers to commandeer all the AI hardware it needs. So I think a purely digital AI attack would not need this.

Now, once the AI can carry all the compute it might need, I'd really worry when it doesn't only carry compute but also more explosive ordinance.

Post reply on HN