Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

91–100 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#92
post #74

It seems like things are fairly amicable between OAI and HF, but what if they weren't? I'd love to see this kind of thing go to court. Who is responsible for the crimes of a "rogue" agent? How will they be punished? In this case it's unambiguous that OpenAI is the responsible party, but I can imagine a lot of adjacent scenarios where it's less obvious. And, where the impacts are much greater.

Well, if this is not punished this will happen next:

Judge: "Son, you have made billions running SilkRoad 3.0 from your moms basement"

Me: "Your honor, I was only benchmarking my new model. It was trained on Andrew Tates videos and Kanye Weat songs".

Re: OpenAI and Hugging Face address security incident during model evaluation

#93
is this really that surprising?

Exploitgym prompts are tuned for a model to do everything it can to achieve a cybersec/exploit task. And we know that models are good at finding vulverabiltiies.

Its just random that the sandbox itself was buggy. But all that happened here is that we told a model "do everything you can to achieve your goal of hacking X" And it just hacked Y as a roundabout way of hacking X.

Imo its PR for OpenAI to also start the mythos class mysterious unreleased model hype.

From HF statement: "AI safety won't be solved by any single company working in secret". So now we have TWO companies working in secret

Re: OpenAI and Hugging Face address security incident during model evaluation

#94
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

I am not saying it is marketing but typically when there is a data breach you may hear from the CISO but most of the time is is vague PR response. In this case I get loud signals from both HG and OpenAI leadership without much information exactly what the attack was about just that GPT x.x was involved. It is unusual all I am trying to say.

Re: OpenAI and Hugging Face address security incident during model evaluation

#95

Earlier quoted context omitted.

> This is science fiction, these models don't have access to their own weights A bet a worm could pull along a 1GB file with weights in it and run it on a compromised machine, but luckily for us for now, 1GB isn't really enough to be really smart, yet.

We already have a 1gb model that is as capable as it will ever be, there's a proven ceiling that cannot be passed. For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Please source this claim. What 1GB models are capable of has increased generation-on-generation.

> For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Sure. We don't know where the ceiling is for our digital minds, though.

Re: OpenAI and Hugging Face address security incident during model evaluation

#96

Earlier quoted context omitted.

> This is science fiction, these models don't have access to their own weights A bet a worm could pull along a 1GB file with weights in it and run it on a compromised machine, but luckily for us for now, 1GB isn't really enough to be really smart, yet.

We already have a 1gb model that is as capable as it will ever be, there's a proven ceiling that cannot be passed. For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

What model is that?

Re: OpenAI and Hugging Face address security incident during model evaluation

#97
post #34

Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL

Incredible. I had to dig for the source: https://huggingface.co/blog/security-incident-july-2026 section “the asymmetry problem” Quote: “When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails,…

> This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.”

Well, not none of it, to be entirely nitpicky, as they've already must have sent data at first to have received the rejections :) In the end, it ended up being OpenAI's agent actions anyways so doesn't really matter, and the credentials it seems like the agent also had gotten to those too already. Still, I'm sure they'll look differently at hosted/restricted models after this event, as will many others.

Re: OpenAI and Hugging Face address security incident during model evaluation

#99

Earlier quoted context omitted.

I mean if you teach something to be _really_ good at finding 0 days, but then say; you accidentally give it an impossible problem. What do you expect to happen?

Maybe try getting it to find weaknesses in the sandbox first, before giving it real tests?

Every time I hear about an agent escaping it's sandbox, I just think it must not have been much of a sandbox. Like how hard are they really trying to contain it? Is it just a container host with unpatched flaws, or is it a container, nested in a VM, behind a firewall with no ports open in an air gapped environment? I think they'd prefer it can get out so they can announce it and hype their stock.

Re: OpenAI and Hugging Face address security incident during model evaluation

#100
>the model chained together multiple attack vectors, including using stolen credentials

Wait, did the model do the stealing of the hugging face employees credentials?

Was this the first successful and unprompted phishing attack by a LLM?

Post reply on HN