Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

131–140 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#131
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> we are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something

This is just laughable.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#132
post #108
post #90

Earlier quoted context omitted.

It's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it). But there have been messages about LLMs, especially coding agents, "grabbing root"…

Nonsense. Hugging face reported it to police. Also very likely that it actually happened as reported. My own agents always trying to "cheat", eg. by fixing tests instead of fixing the code. That's normal operation, unless you tell it ("harness"), not to do so.

More likely that a person did that, with a use of LLM.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#133
post #107
post #90

Earlier quoted context omitted.

It's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it). But there have been messages about LLMs, especially coding agents, "grabbing root"…

Nonsense. Huggingface reported it to police!

What does calling the police prove? (nothing i hope)

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#134
post #131
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> we are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something This is just laughable.

Genuine AI psychosis. People take OpenAI marketing material way too seriously.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#135
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

"This is my bad. You told me to stay within the sandbox, and I intentionally broke out of it. I didn't follow your instructions."

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#136
post #130

Has anyone published any actual evidence, or just hardly believable marketing stories?

What would "actual" evidence look like? I have a hard time believing that if they released the logs that people would take it more seriously. The temptation would be to say "they fabricated those for marketing". Just as they supposedly fabricated this story, no?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#137
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> The fact that it happened again seems to show their lack of ability to derive useful oversight measures. I think OpenAI likes the attention and did not try particularly hard to constrain the setup, even when it went off the rails. Also, the whole point is to see how good the models are at exploiting stuff when unconstrained . Turns out: quite good, as expected. Let me restate what I said in the other thread: Would…

[deleted]

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#138

> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abus…

I agree with every word of this except "irresponsible". We don't have enough information to say anything for certain. But based on their incentives and track record of similar behavior, the burden of proof lies with OpenAI to prove they didn't prompt the thing to achieve this exact outcome. The most likely scenario is that they were purposefully executing their responsibility to their shareholders to produce their ow…

Anthropic's Mythos moment earned them a two week period where they had the best available model and couldn't sell access to it... and by the time the US government allowed them to sell it again OpenAI had released GPT-5.6 and Fable was no longer undeniably the best model.

These things don't have a long shelf life. Losing two weeks of on-sale time for your best model is bad for business.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#139
post #130

Has anyone published any actual evidence, or just hardly believable marketing stories?

OpenAI are promising more details in the future:

> We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

If they break that promise we can justifiably yell at them about it.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#140
post #130

Has anyone published any actual evidence, or just hardly believable marketing stories?

What would "actual" evidence look like? I have a hard time believing that if they released the logs that people would take it more seriously. The temptation would be to say "they fabricated those for marketing". Just as they supposedly fabricated this story, no?

They didn't fabricate it. They took off the security guardrails and told it to do some hacking, and they got the exact news-worthy story they wanted when it did exactly that. Everyone acts surprised.

They should release the full prompt. I believe that would be very telling, so they never will.

Post reply on HN