Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

111–120 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#113
post #4

> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times

Crazy doesn't even begin to describe it. I'm hardening my computers as much as I can but I'm not sure it's enough. At some point anyone who isn't running local AI themselves probably isn't gonna make it.

Local AI won't help you if an agent goes roque.

Re: OpenAI and Hugging Face address security incident during model evaluation

#114

Earlier quoted context omitted.

> This was a long-horizon, unsupervised task burning millions of tokens. As if the immediate future wasn't billions of these tasks... Many successfully improving their own capabilities

> As if the immediate future wasn't billions of these tasks... There's only so many GPUs and a lot of them are devoted to patching flaws. > Many successfully improving their own capabilities I haven't seen much of that. But that also applies to the ones on defense. And more flaws are probably going to take increasing resources to find.

> There's only so many GPUs and a lot of them are devoted to patching flaws.

Might want to look at Nvidia and TSM production and revenue value trajectories. Also the algorithmic improvements currently being found along with models that are solving unprecedented mathematical and scientific problems every week now.

> I haven't seen much of that.

Then you must not be aware frontier lab employees are using frontier internal models to ship improvements to models via agentic loops. They are hardly prompting anymore, it's guiding very long running coding tasks. The trajectory over the past few years has been to remove more and more of any human input into the process, and once that is soon achieved, it is indefinite recursive self improvement, RSI.

What's here and what's coming: https://www.anthropic.com/institute/recursive-self-improveme...

Re: OpenAI and Hugging Face address security incident during model evaluation

#115
post #15

Earlier quoted context omitted.

Huggingface did not have access to the models. They were running in OAI’s infrastructure.

Ah, that makes more sense :) But then, why attack huggingface? The exploitgym dataset is on github and can be downloaded without need for exploits?

breadth search and found huggingface first? Pure speculation

Re: OpenAI and Hugging Face address security incident during model evaluation

#116
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

Their entire business model from the beginning of ChatGPT was to deny responsibility

Re: OpenAI and Hugging Face address security incident during model evaluation

#118
I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:

Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.

Re: OpenAI and Hugging Face address security incident during model evaluation

#120
post #4

> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times

Indeed. Real life hacks are beginning to sound like Neuromancer.
Post reply on HN