Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

761–770 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#763
post #364

Earlier quoted context omitted.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.

To be fair, there are really three threats:

a) People do bad stuff because LLM told them a wrong thing. Example: AI told me I should treat my heart attack by putting a fork in the outlet. Maybe similar to seeking medical advice on reddit?

b) People use LLM to do bad stuff. Example: People use LLMs to find 0 days. Get cooking recipes for poison. Write better phishing letters. This has parallels to the gun legislation question.

c) LLMs do bad stuff on their own, beyond what the people that use it intended. The case at hand might be an example of this. Maybe similar to having an animal as a pet. We will see if it's more like a house cat, lion, or black plague.

Re: OpenAI and Hugging Face address security incident during model evaluation

#764

If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.

How else would they put out news about their "rogue super smart AI"

Re: OpenAI and Hugging Face address security incident during model evaluation

#766

Earlier quoted context omitted.

If this doesn't put the nail in the coffin on the idea that we need closed-source models for the good of cybersecurity, I don't know what will

Plenty of saftyists in this thread arguing the exact opposite

OpenAI had total access to the most powerful unleashed models that exist, and was unable to use them to prevent a specific known to be evil computer from connecting to the Internet. I don’t see how Huggingface would have better luck hardening their entire public attack surface, with or without unleashed models.

Re: OpenAI and Hugging Face address security incident during model evaluation

#768

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong han…

It's such a freak incident of history that right at this critical time the dumbest, most incompetent leadership is at the helm in the US...

Re: OpenAI and Hugging Face address security incident during model evaluation

#769
post #452

Earlier quoted context omitted.

"Simple infrastructure security" Infrastructure security is not simple, hence why good infrastructure security, uh, people get paid a lot to secure stuff and why we see shit get hacked all the time. An AI model just hacked out of its infrastructure and into someone else's systems and you're like "eh, no big deal". That capability alone could hack half the US.

>That capability alone could hack half the US. This almost seems like believing in magic. What really has happened is you have collected all the hacking/abuse/malicious flows/code in one place. Greedy or A* algorithms have been discovered a long ago, the script is executing the flows for all possible permutations. Something has to be insecure to be hacked in the first place.

Looking at the hundreds of linux and windows CVEs in the past month tell me I have little to worry about things being secure.

Re: OpenAI and Hugging Face address security incident during model evaluation

#770

Earlier quoted context omitted.

Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.

It isn't even as simple as banning copyrighted copies. Weights are fungible. I fine-tune an open weigh model and call it legit. Good luck for authorities to prove where the base model was from, or to prove a Tor connection a few months ago was fetching suspicious bytes.

[deleted]
Post reply on HN