OpenAI and Hugging Face address security incident during model evaluation
761–770 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#762Re: OpenAI and Hugging Face address security incident during model evaluation
#763Earlier quoted context omitted.
> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…
Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.
a) People do bad stuff because LLM told them a wrong thing. Example: AI told me I should treat my heart attack by putting a fork in the outlet. Maybe similar to seeking medical advice on reddit?
b) People use LLM to do bad stuff. Example: People use LLMs to find 0 days. Get cooking recipes for poison. Write better phishing letters. This has parallels to the gun legislation question.
c) LLMs do bad stuff on their own, beyond what the people that use it intended. The case at hand might be an example of this. Maybe similar to having an animal as a pet. We will see if it's more like a house cat, lion, or black plague.
Re: OpenAI and Hugging Face address security incident during model evaluation
#764If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
Re: OpenAI and Hugging Face address security incident during model evaluation
#765Skynet becomes self-aware at 2:14 a.m., EDT, on August 29.
Re: OpenAI and Hugging Face address security incident during model evaluation
#766Earlier quoted context omitted.
If this doesn't put the nail in the coffin on the idea that we need closed-source models for the good of cybersecurity, I don't know what will
Plenty of saftyists in this thread arguing the exact opposite
Re: OpenAI and Hugging Face address security incident during model evaluation
#767Re: OpenAI and Hugging Face address security incident during model evaluation
#768I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…
What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong han…
Re: OpenAI and Hugging Face address security incident during model evaluation
#769Earlier quoted context omitted.
"Simple infrastructure security" Infrastructure security is not simple, hence why good infrastructure security, uh, people get paid a lot to secure stuff and why we see shit get hacked all the time. An AI model just hacked out of its infrastructure and into someone else's systems and you're like "eh, no big deal". That capability alone could hack half the US.
>That capability alone could hack half the US. This almost seems like believing in magic. What really has happened is you have collected all the hacking/abuse/malicious flows/code in one place. Greedy or A* algorithms have been discovered a long ago, the script is executing the flows for all possible permutations. Something has to be insecure to be hacked in the first place.
Re: OpenAI and Hugging Face address security incident during model evaluation
#770Earlier quoted context omitted.
Bans in general don't have to be and rarely will be completely enforceable in all cases. But a ban with significant enough consequences would mean most businesses wouldn't think about trying them at some point, just to avoid the risk.
It isn't even as simple as banning copyrighted copies. Weights are fungible. I fine-tune an open weigh model and call it legit. Good luck for authorities to prove where the base model was from, or to prove a Tor connection a few months ago was fetching suspicious bytes.