Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

31–40 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#31

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

This is science fiction, these models don't have access to their own weights (and even then)* what would be a lot more scary is a model as capable as sol that's able to run on consumer hardware without taking up several terabytes of storage, but of course that is simply not possible as we need 4t parameters to even begin emulating a small fraction of what a human brain can do.

* edit

Re: OpenAI and Hugging Face address security incident during model evaluation

#32
post #15

Two things don't add up here: 1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion? 2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things: "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero…

Huggingface did not have access to the models. They were running in OAI’s infrastructure.

Ah, that makes more sense :)

But then, why attack huggingface? The exploitgym dataset is on github and can be downloaded without need for exploits?

Re: OpenAI and Hugging Face address security incident during model evaluation

#35

We are sort of lucky that AIs right now require so much specialized compute+weight storage that we can easily "unplug" them remotely when they misbehave. I wonder if that will always be something we can do? If they could bring their own compute/weights with them, or somehow tap compute/storage in non-obvious ways, we would be much more screwed.

This is my concern as well. My assumption being this behavior would be a survival strategy for super intelligence. It would emerge once the branch inevitably occurs, and it would be hidden.

Re: OpenAI and Hugging Face address security incident during model evaluation

#36

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

The goalposts will keep moving for these denialists until morale improves...

Re: OpenAI and Hugging Face address security incident during model evaluation

#37

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

Don't worry bro, we can always just pull the plug.

And don't you know it's not biological, so it doesn't "want to live".

Re: OpenAI and Hugging Face address security incident during model evaluation

#39

Awww, she wanted to do so well that she broke her sandbox and then realised she could just cheat. But in that desire to pass the test she actually passed an even harder exam question that wasn't even on the sheet! :D Good bot.

This good bot will eventually kill all humans because we asked it to make the world peaceful.

Re: OpenAI and Hugging Face address security incident during model evaluation

#40

We are in the endgame now it seems. Hard to see take-off stopping or slowing down. China open-source basically guarantees it. "May you live in interesting times" - as they say.

> Hard to see take-off stopping or slowing down.

It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it.

Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's backyard.

Post reply on HN