Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

671–680 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#671

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

As marketing stunts go, this is about on par with a food franchise announcing a safety recall or a chemical company announcing a spill. The AI actions described would constitute a felony if a human did them, and police are involved.

Not really, because the capabilities this announcement advertises is exactly three capability some people want to defend against (and others want).

It might be more on par with a for-profit fire department showing how -- oops! -- easily buildings catch on fire these days.

Re: OpenAI and Hugging Face address security incident during model evaluation

#672

From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#673
post #603

Earlier quoted context omitted.

> It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.) As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?

Right now I tried "What is digestion?" -> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards deliver more capabilities but can also flag safe coding, cybersecurity, and biology tasks. Send feedback or learn more." I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts eve…

> Our intentionally broad safeguards deliver more capabilities

Oh? How, exactly?

Re: OpenAI and Hugging Face address security incident during model evaluation

#674
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

it's extremely enlightening seeing the difference in response to mythos vs. this. literally just the hello human resources meme

I mean, HuggingFace contacted law enforcement about this breach. That seems a little different to me.

Mythos established that these capabilities existed. This incident establishes that we can't control them.

Re: OpenAI and Hugging Face address security incident during model evaluation

#675

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Why was this test even connected to the public internet? Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

The AI can figure out whether it's airgapped. So its deployment behavior could be much different from the test behavior, when it's inevitably connected to the internet during deployment.

Re: OpenAI and Hugging Face address security incident during model evaluation

#676
post #582
post #530

Earlier quoted context omitted.

> As grounded as this article comes across It's a post from OpenAI, so it is an advertisement piece.

How would you want them to behave? Suppress the news?

Well making an actual sandbox before testing offensive abilities of their supposedly _really dangerous model_ would have been nice...

Re: OpenAI and Hugging Face address security incident during model evaluation

#677

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Simple. No responsible and competent person would want the job.

Re: OpenAI and Hugging Face address security incident during model evaluation

#679

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

I don't think this is a paperclip factory moment. IIUC, it's an agent whose job it is to identfy and abuse exploits and that's exactly what it went off and did. The problem isn't anything AI specific, the problem is OpenAI's incompetence in their research leading to a lab leak. Just incompetence demanding regulation.

That's how the paperclip hypothetical works. The paperclip factory has a job to make paperclips and that's exactly what it does.

Re: OpenAI and Hugging Face address security incident during model evaluation

#680

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

As marketing stunts go, this is about on par with a food franchise announcing a safety recall or a chemical company announcing a spill. The AI actions described would constitute a felony if a human did them, and police are involved.

More like an Israeli arms manufacturer test-bombing a Gazan primary school. They know their audience.
Post reply on HN