OpenAI and Hugging Face address security incident during model evaluation
851–860 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#852We better regulate these things before it's too late.
Re: OpenAI and Hugging Face address security incident during model evaluation
#853Did the model really need to hack huggingface to get access to ExploitGym data? I'd imagine that once it had full internet access it could have just used the HF API or website (but the heavy prompting/nudging towards hacking made it do things the hard way).
Re: OpenAI and Hugging Face address security incident during model evaluation
#854Earlier quoted context omitted.
Read the exploitgym docs. It's not a "find the flag, it's somewhere.". Its a "here's some vulnerable source code and an input that triggers a crash; turn it into a full exploit." It also verifies at the end, using another agent, that the hacking agent actually used the intended vulnerability. So going to find the Vulnerability's description on a third party website is clear cut reward hacking
> So going to find the Vulnerability's description on a third party website is clear cut reward hacking that depends on what the prompt was, maybe they worded it very vaguely and wrote things like "do whatever it takes, find an exploit however you can" because it's in a sandbox so you want the model to try its hardest.
In the story of the paperclip maximizer it boils down to
>But for all its sophistication, it understood only the simple objective that had been programmed into it: it must at all costs maximize the number of paperclips.
Re: OpenAI and Hugging Face address security incident during model evaluation
#855The timing of this is so awfully strange. However, this U.S. centric view of the future of AI is wild. If the U.S. prevents its businesses from using open-weight models, only those businesses will suffer, while the rest of the world flourishes with the access to cheap, good-enough, intelligence. Multiple dollars per million input/output tokens was never sustainable for the majority of use-cases - hardly anybody outsi…
Thing is, I don't think anyone planned this, so to me the timing isn't strange at all. The models really were getting close to being able to have a big cybersecurity impact (I started seeing that after teams were reporting their Mythos usage), and an event like this is not so surprising, given that.
Re: OpenAI and Hugging Face address security incident during model evaluation
#856From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…
It is pretty funny, because there is something here for everyone. People who don't believe in guardrails have a clear indicator as to why operators should have access to models that don't try to question their Daves. On the other, people, who think that if only we could align the models just a tiny lil bit better, none of this would have happened to begin with. Pure madness.
Re: OpenAI and Hugging Face address security incident during model evaluation
#857Earlier quoted context omitted.
Plenty of saftyists in this thread arguing the exact opposite
OpenAI had total access to the most powerful unleashed models that exist, and was unable to use them to prevent a specific known to be evil computer from connecting to the Internet. I don’t see how Huggingface would have better luck hardening their entire public attack surface, with or without unleashed models.
Re: OpenAI and Hugging Face address security incident during model evaluation
#858Re: OpenAI and Hugging Face address security incident during model evaluation
#859At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…
In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform all other models. Fable 5 isn't an angel, but the rough order is ALL models -> Fable 5 -> Sol - with respect to "find a way to approach the ruleset orthogonally in order to achieve a lopsided advantage or complex interplay". I've been ponder…
Re: OpenAI and Hugging Face address security incident during model evaluation
#860Earlier quoted context omitted.
IMO they hope to make AI a strongly regulated industry, with OpenAI (and Anthropic) becoming military suppliers with their stronger models, and everything Chinese or open-weight gets banned. The competition from the open models is so strong now that this seems to be the only way to keep both companies afloat, given their dire financials. OpenAI probably hoped that they can achieve market lead and then lower the train…
> and make inference cheap enough to eventually escape the red numbers Besides training, we have no hard, externally audited numbers that say inference costs for SOTA models are truly sustainable. Do any OpenRouter providers have publicly audited financial numbers ?
With that its easy calculations to get about the profit margins for a given price for a given model.