OpenAI and Hugging Face address security incident during model evaluation
581–590 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#582As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…
> As grounded as this article comes across It's a post from OpenAI, so it is an advertisement piece.
Re: OpenAI and Hugging Face address security incident during model evaluation
#583Earlier quoted context omitted.
This is marketing. Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?
It’s marketing the same way shitting your pants in public is marketing. People notice you.
Re: OpenAI and Hugging Face address security incident during model evaluation
#584Earlier quoted context omitted.
I'm sure that will change sooner rather than later, otherwise enterprising hackers will be able to claim that the model they were using went rogue.
This would be bad; we've already had a few cases on HN where someone noticed that they could increment the customer number in a URL or similar, resulting in police action.
Re: OpenAI and Hugging Face address security incident during model evaluation
#585Earlier quoted context omitted.
You can have large scale airgapped environments. They don’t even need to be fully airgapped from each other (and is not what I’m suggesting). But there should be no physical (physical layer; wireless counts) to the internet.
Can models detect they are airgapped and change their behaviors? How much of the internet do you have to simulate to know if the model knows it's in training?
Regardless, they (reportedly) _attempted_ to prevent internet access. They just didn’t in a way which can be escaped via software.
Yes, side channel exploits exist in airgapped environments to. But if a model found a way to escape an airgapped environment via non-networked side channel attacks then the correct answer is frankly “shut it down immediately and then thermite any machine it touched”
Re: OpenAI and Hugging Face address security incident during model evaluation
#586If you are attempting to run exercises like this, it is wildly negligent to not be running it in a physically-airgapped environment (potentially with a physical power shutdown). You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
> wildly negligent to not be running it in a physically-airgapped environment Why should it be physically airgapped? Clients won't be doing that.
Did we learn nothing from all of those Star Trek holodeck jailbreaks?
Re: OpenAI and Hugging Face address security incident during model evaluation
#587From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…
It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.
Re: OpenAI and Hugging Face address security incident during model evaluation
#588From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…
It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.
Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve
It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.)
Re: OpenAI and Hugging Face address security incident during model evaluation
#589From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…
It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.
Re: OpenAI and Hugging Face address security incident during model evaluation
#590Researcher: hack me
Model: understood
Researcher: oh my god