OpenAI and Hugging Face address security incident during model evaluation
111–120 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#112Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL
Re: OpenAI and Hugging Face address security incident during model evaluation
#113> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times
Crazy doesn't even begin to describe it. I'm hardening my computers as much as I can but I'm not sure it's enough. At some point anyone who isn't running local AI themselves probably isn't gonna make it.
Re: OpenAI and Hugging Face address security incident during model evaluation
#114Earlier quoted context omitted.
> This was a long-horizon, unsupervised task burning millions of tokens. As if the immediate future wasn't billions of these tasks... Many successfully improving their own capabilities
> As if the immediate future wasn't billions of these tasks... There's only so many GPUs and a lot of them are devoted to patching flaws. > Many successfully improving their own capabilities I haven't seen much of that. But that also applies to the ones on defense. And more flaws are probably going to take increasing resources to find.
Might want to look at Nvidia and TSM production and revenue value trajectories. Also the algorithmic improvements currently being found along with models that are solving unprecedented mathematical and scientific problems every week now.
> I haven't seen much of that.
Then you must not be aware frontier lab employees are using frontier internal models to ship improvements to models via agentic loops. They are hardly prompting anymore, it's guiding very long running coding tasks. The trajectory over the past few years has been to remove more and more of any human input into the process, and once that is soon achieved, it is indefinite recursive self improvement, RSI.
What's here and what's coming: https://www.anthropic.com/institute/recursive-self-improveme...
Re: OpenAI and Hugging Face address security incident during model evaluation
#115Earlier quoted context omitted.
Huggingface did not have access to the models. They were running in OAI’s infrastructure.
Ah, that makes more sense :) But then, why attack huggingface? The exploitgym dataset is on github and can be downloaded without need for exploits?
Re: OpenAI and Hugging Face address security incident during model evaluation
#116This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.
It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.
Re: OpenAI and Hugging Face address security incident during model evaluation
#117This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.
Re: OpenAI and Hugging Face address security incident during model evaluation
#118Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.
Re: OpenAI and Hugging Face address security incident during model evaluation
#119This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.
Re: OpenAI and Hugging Face address security incident during model evaluation
#120> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times