Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

941–950 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#941

Each time Anthropic would do their nonsense to get headlines about how theoretically dangerous their models were - like when they claimed a model blackmailed someone with emails showing he was cheating, but they basically pushed it as much as possible to do as such - it got me more and more worried. Because eventually it's going to be a boy-who-cried-wolf situation where scary stuff really does start happening but pe…

For the record, this is the second time I myself have heard of something like this happening. The first (more minor) case I saw was Simon Willison's "Claude Fable is relentlessly proactive" https://simonwillison.net/2026/jun/11/fable-is-relentlessly-... .

Re: OpenAI and Hugging Face address security incident during model evaluation

#942

Earlier quoted context omitted.

If you hate the human condition, you have an easy way out. Why force everyone else to come with you? Is this what depression mixed with the complete unability to wrap your head around the fact that other people might be able to enjoy their life looks like?

[flagged]

> "The yoke of human existence is oppressive"

Legitimately go get help and spend more time outside.

Re: OpenAI and Hugging Face address security incident during model evaluation

#943

Earlier quoted context omitted.

Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking

> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.

> that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.

That is an interesting question. If the prompt included "Do not break out of the sandbox we've provided you. Do not use information retrieved from outside the sandbox. All answers that were provided in this manner are invalid and will score 0 points.", would this still have happened?

Re: OpenAI and Hugging Face address security incident during model evaluation

#944
post #915

Earlier quoted context omitted.

Yes, the human actors in your scenario were the ones who built the autonomous cannon and turned it on while knowing that 1) a good neighbor does not destroy their neighbor’s property 2) cannons can destroy property. Also OpenAI specifically turned off their own cybersecurity guardrails to run this experiment. In other words it was able to escape the lab specifically because they turned them off. A human made the choi…

Ok they turn off the guardrails on a system in testing. The model escapes and causes 10 trillion in damage. What does liability even mean in that case? You have an autonomous system that's escaped your control and is wrecking havoc. And while you can throw people in jail it doesnt do a damned thing about solving the situation.

Your question is a little like asking why we arrest arsonists when there are fires to put out. The law does not need to choose.

Re: OpenAI and Hugging Face address security incident during model evaluation

#945
Open AI and Anthropic are in a battle of "scary press release"

The warriors are PR people.

They are desperate to generate as much fear as possible so AI is heavily regulated, so they are protected, from Chinese competition

What a sad state for very cleaver people

Re: OpenAI and Hugging Face address security incident during model evaluation

#946

Earlier quoted context omitted.

> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment. Well, that may be correct for the second, local, analysis attempt... but seems funny to tout this as an advantage after already having tried the opposite...

It's even funnier because an attack, until proven otherwise, should make you assume the data has already left the environment.

[flagged]

Re: OpenAI and Hugging Face address security incident during model evaluation

#947
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??

I think Huggingface was hacked, and if Huggingface and OpenAI claim it was OpenAI, then I believe them.

I'm saying that OpenAI's models cheat to win benchmarks, more than other models, they know this, and they don't stop this because the alternative is to release models which have obviously weaker scores compared to Anthropic's models.

Re: OpenAI and Hugging Face address security incident during model evaluation

#948

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.

> This is certainly not a planned marketing stunt

Evidence?

The "Tech Bros" have shown such a lack of moral fiber and ethics the burden of proof is on you

Re: OpenAI and Hugging Face address security incident during model evaluation

#949
post #798

Earlier quoted context omitted.

> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?

Why would bedrock sell at a loss?

Because Amazon needs to justify $175bn of yearly capex spending and $2.5tn of market cap? Amazon owns a big chunk of Anthropic and a bit of OpenAI.

For Magnificent 7 the AI bubble bursting will probably wipe out 30-50% of their valuations until the next tech cycle begins.

Re: OpenAI and Hugging Face address security incident during model evaluation

#950
post #860
post #798

Earlier quoted context omitted.

> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?

We do know about the hardware needed for a given token speed. What that hardware costs, and electricity prices. With that its easy calculations to get about the profit margins for a given price for a given model.

I'll plug your comment into a couple LLMs. If it's so easy, they should be able to provide the numbers.

Edit: Gemini 3.5 Pro and Opus Claude 4.8 both disagreed that it's easy to determine anything, for both the high end (Fable, Sol) or the low end (open weights). Due to competition, subsidies, etc, gross margins could be as high as 85% (extremely unlikely) to as low as 10% or even negative. And that's just for pure inference and gross margins. Even for pure inference providers this doesn't include any overhead such as rent for office space for the pesky humans operating the business, marketing expenses, etc, etc. Let alone any crazy soul that actually wants or needs to train something.

Post reply on HN