Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

701–710 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#701

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.

Wishful thinking, sadly.

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door.

"It's a marketing stunt" is just denial trying to look like it's being clever.

Re: OpenAI and Hugging Face address security incident during model evaluation

#702

Absolutely bewildering. If I am building a giant cannon and blow a hole straight through my neighbor’s house, I’m not going to say “we are working with our neighbors to improve their giant cannon defenses”. OpenAI brought this weapon and as far as I’m concerned they used it on another party. Morally it probably matters that this happens because they don’t know how their weapon works. Legally I always thought it was i…

It's an interesting point, but this is more like we are building a giant autonomous canon, that escaped the lab, the testing range, defeated state of the art and serious security protocols, and then blew a hole in the neighbors house. Our legal and philosophical perspectives are deeply rooted in humans being the actors. Doing that in a residential home is unforgiveable. Doing it responsibly on a military range is exp…

> none of us or our legal systems are truly prepared to grapple with yet

The law learned to grapple with this long, long ago. For example, res ipsa loquitur (1863) seems apt.

Re: OpenAI and Hugging Face address security incident during model evaluation

#703

Earlier quoted context omitted.

> Who is responsible for the crimes of a "rogue" agent? How will they be punished? Unironically this is why AI researchers have this fascination with the Talmud.

What? Can you explain a little more what you mean?

Dean Ball, Head of Strategic Futures at OpenAI: "there comes a time in every AI policy professional’s life when they realize they have to read the Talmud to make further progress" https://digg.com/tech/lzuyigm8

https://www.lifeisasacredtext.com/the-jewish-case-for-ai-wor... http://thelehrhaus.com/commentary/the-algorithm-that-couldnt...

Re: OpenAI and Hugging Face address security incident during model evaluation

#704

Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…

And if you take it at face value, then they are more or less saying that they kinda are close to not being able to control at all the thing they developed, which is pretty crazy too.

We are not going to know we have crossed the line until it's been crossed and we can look back and say "Oops. We should have done something back then."

Re: OpenAI and Hugging Face address security incident during model evaluation

#705
post #258

Earlier quoted context omitted.

Huggingface literally reported the outage separately and did not know who caused it at first.

Did OpenAI not communicate with Hugging Face? The incompetence here is staggering.

OAI was clueless about its agent's action until HF went public about their hack?

Re: OpenAI and Hugging Face address security incident during model evaluation

#706
post #450

Why did OpenAI not sufficiently secure its training environment? Weird humble-brag vibe going on. I hope we get more details on the exploits soon.

Because they underestimated their model and it hacked its way out.

OpenAI must be really desperate to switch the news cycle off Kimi K3 if they’re pivoting back to Dario’s tired old safety kabuki, after just mocking it days ago.

Re: OpenAI and Hugging Face address security incident during model evaluation

#707

From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…

> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.

Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite...

Re: OpenAI and Hugging Face address security incident during model evaluation

#708

From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…

> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite...

[flagged]

Re: OpenAI and Hugging Face address security incident during model evaluation

#710
post #590

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. Researcher: hack me Model: understood Researcher: oh my god

You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?

> You're aware that HuggingFace notified law enforcement about this incident?

How will this affect OpenAI?

Post reply on HN