Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

851–860 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#853

Did the model really need to hack huggingface to get access to ExploitGym data? I'd imagine that once it had full internet access it could have just used the HF API or website (but the heavy prompting/nudging towards hacking made it do things the hard way).

I would be surprised if the solutions for the benchmark were publically available like that?

Re: OpenAI and Hugging Face address security incident during model evaluation

#854

Earlier quoted context omitted.

Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking

> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.

I don't quite understand how that changes anything?

In the story of the paperclip maximizer it boils down to

>But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.

Re: OpenAI and Hugging Face address security incident during model evaluation

#855

The timing of this is so awfully strange. However, this U.S. centric view of the future of AI is wild. If the U.S. prevents its businesses from using open-weight models, only those businesses will suffer, while the rest of the world flourishes with the access to cheap, good-enough, intelligence. Multiple dollars per million input/output tokens was never sustainable for the majority of use-cases - hardly anybody outsi…

Why is it strange? To me, it could only be strange if one was considering that someone planned this, and then yes, it's hard to come up with anyone who would benefit from this. OpenAI looks dumb, but also their models sound impressive, Chinese models look dangerous, but also useful. No clear winner.

Thing is, I don't think anyone planned this, so to me the timing isn't strange at all. The models really were getting close to being able to have a big cybersecurity impact (I started seeing that after teams were reporting their Mythos usage), and an event like this is not so surprising, given that.

Re: OpenAI and Hugging Face address security incident during model evaluation

#856

From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…

It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.

Yeah, this is very much one of those stories where people from many different perspectives or chopping it up on a plate and ripping it through a straw.

Re: OpenAI and Hugging Face address security incident during model evaluation

#857

Earlier quoted context omitted.

Plenty of saftyists in this thread arguing the exact opposite

OpenAI had total access to the most powerful unleashed models that exist, and was unable to use them to prevent a specific known to be evil computer from connecting to the Internet. I don’t see how Huggingface would have better luck hardening their entire public attack surface, with or without unleashed models.

Except you are forgetting that in reality, HuggingFace switched to a open weights model and fixed their system.

Re: OpenAI and Hugging Face address security incident during model evaluation

#859

At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…

In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform all other models. Fable 5 isn't an angel, but the rough order is ALL models -> Fable 5 -> Sol - with respect to "find a way to approach the ruleset orthogonally in order to achieve a lopsided advantage or complex interplay". I've been ponder…

Inner misalignment is just a natural state of LLMs, heck of humans when it comes to children and the corrupt.

Re: OpenAI and Hugging Face address security incident during model evaluation

#860
post #798
post #600

Earlier quoted context omitted.

IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned. The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the train…

> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?

We do know about the hardware needed for a given token speed. What that hardware costs, and electricity prices.

With that its easy calculations to get about the profit margins for a given price for a given model.

Post reply on HN