Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

11–20 of 259 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#11
post #7

This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended. Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

> where that behavior was never even intended.

Strongly doubt that. Did they even share the prompt?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#13
post #7

This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended. Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#14
Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next”.

What purpose could this behavior serve, other than cyber attacks and whatnot? Why train and optimize models for these things, if not for being used in cyber warfare?

Perhaps they envision a future where the DoD is going to be their biggest customer?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#15

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Being right _all the time_ for positive outcomes is difficult/expensive.

Being "right" just once for negative outcomes is achievable and rewarding.

And things are getting desperate.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#16

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

> did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose?

That's the point. It's like a pool hall with "NO GAMBLING" signs posted on the walls.

The message is that the hall is intended for gambling, but that the hall's patrons may be held liable if the situation becomes inconvenient for the proprietor.

In this case, the product is intended for hacking, but of course the user may be held liable if the situation becomes inconvenient for the model's proprietor.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#18
post #15

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Being right _all the time_ for positive outcomes is difficult/expensive. Being "right" just once for negative outcomes is achievable and rewarding. And things are getting desperate.

The very reason I have always felt a bit of undue loyalty to blue team. A red teamer just has to find one vuln, blue team needs to find _all_ vulns.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#19
post #7

This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended. Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

> This feels straight out of sci-fi.

Most AI marketing is straight up science fiction.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#20
post #3

Is it normal for these training/eval runs to go on for over a month?

The way I read it was different things happening over several runs, such as the agents comparing notes so to speak, using artifactory

I can’t get over how the process is exactly what a hacker hive does. Communicate leaving notes in some random file.
Post reply on HN