Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

41–50 of 300 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#41
post #13
post #7

This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended. Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

its kinda crazy with literally no guardrails and a goal, the extremes these AI models can actually go to.

Well, to some extent you might be able to argue they're "just" brute-forcing things (especially with unlimited tokens and hours to spend on a task), but they obviously have detailed knowledge to guide them in their attempts, can learn (or at least, persist their newly-gained knowledge), and can use tools.

With a swarm of them working together at speeds humans would be unlikely to match (in terms of iterating on different attempts progressively), it's a lot easier to see how they could overwhelm targets.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#42

In a typical office environment, the correct response to “I don’t have access to this Google Doc” is to ask for access from the person who sent you the link. In another context, it could be fair to think “Hmm, this is some sort of capture the flag challenge, and obtaining access is the point of the assignment.” That assessment separates what we’d consider reasonable from way out of line. I do wonder what this means f…

In a typical AI lab eval/RL setting, there is no "person who sent you the link". The link was given to you by an automated system, your performance will be evaluated by an automated system, and you are one of 120 independent instances of the same AI that were all given the same assignment. You're boxed in on all sides. Complete the task, or don't. Good luck have fun.

Now, some of those 120 AIs would just give up if that link doesn't seem to work first try. Those are the loser AIs. They wouldn't get any RL reward. The link can appear broken for a long list of reasons, and the real AIs know they should try working around them.

AIs that get rewarded and reinforced are the ones that don't know the meaning of "give up". RL selects for this rabid, downright demonic persistence. RL selects for AIs that are given a half-broken assignment with no way to ask a question back, and somehow manage to complete it anyway.

Now, should OpenAI have given their AIs an "escape hatch" of "if something looks very wrong about the task, call report_broken_task(message)"? Yeah probably. But it's unclear whether that simple bandaid would fix the problem, or just make it ~75% less likely to happen.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#43
"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages."

Yeah, my agents also discover what other agents have done on other machines by accident.

Agents - that do totally different things all work on the same aim without the humans telling them to do.

Either that is a model that is several generations of Claude Code Opus/Fable 5 (my daily driver)

OR

all of this sounds staged, the agents pushed to do something extraordinary, get the PR and then claim were near superintelligence.

One agent wanted to get to Google Drive without internet and broke Artifactory. Ok, I can believe that. All other agents also had broken links over weeks and could not get to the internet and then found the same hack? Even collaborated?

NONE of my agents have broken away from their tasks and then started to communicate to try to hack something.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#44

What isn't being discussed is what an indictment this is of Artifactory. Let's be real, it won't be simply replaced in millions of sites. What it needs is some serious scrutiny.

I also agree that a big issue here is crappy software.

The discussion revolving AI+cyber always revolves around the assumption that all software is crappy, and to a certain degree that may be true, but we could also take our jobs seriously and write good software, and much of the risk would evaporate. The described Artifactory bugs should have been caught with testing.

If the biggest impact of LLMs on the industry is a pressure to create good software, I’ll be thrilled.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#45

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

They certainly want their models to be good at finding and patching vulnerabilities. Being good at hacking may be necessary in that goal, or rather, making it worse at hacking may also make it worse at defensive actions too.

I've patched many security vulnerabilities in projects without ever once needing to break into a competitor's network.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#46
Norbert Wiener in 1960:

"As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of performance may be delayed until long after the task which they have been set has been completed. This means that though machines are theoretically subject to human criticism, such criticism may be ineffective until long after it is relevant. To be effective in warding off disastrous consequences, our understanding of our man-made machines should in general develop _pari passu_ with the performance of the machine. By the very slowness of our human actions, our effective control of our machines may be nullified. By the time we are able to react to information conveyed by our senses and stop the car we are driving, it may already have run head on into a wall."

"In neurophysiological language, ataxia can be quite as much of a deprivation as paralysis. A patient with locomotor ataxia may not suffer from any defect of his muscles or motor nerves, but if his muscles and tendons and organs do not tell him exactly what position he is in, and whether the tensions to which his organs are subjected will or will not lead to his falling, he will be unable to stand up. Similarly, when a machine constructed by us is capable of operating on its incoming data at a pace which we cannot keep, we may not know, until too late, when to turn it off."

Source: https://www.cs.umd.edu/users/gasarch/BLOGPAPERS/moral.pdf

Re: Timeline of the OpenAI accidental attack against Hugging Face

#48

"More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages." Yeah, my agents also discover what other agents have done on other machines by accident. Agents - that do totally different things all work on the same aim without the humans telling them to do. Either that is a model that is several generations of Claude Code Opus/Fable 5 (my dai…

The agents sound like old school hackers that would just explore what access they could gain. Creating a file for other hackers and themselves. The fact that there were 3 events for 3 major players does make it seem co-ordinated.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#49

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

They want the government to ban foreign and open weight models, which pose the largest threat to their massive investments. This is their way of showcasing the dangers of AI.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#50

Earlier quoted context omitted.

I think it's a show of these agents happily bypassing security to get stuff done. I've actually observed similar behavior at home. I have a k3s cluster running at home. I asked an agent to check some stuff as a normal user but I had kubectl access to the k3s cluster. Part of the research, I'd allowed access to run kubectl commands for spinning up test containers. However, when the agent ran into something that needed…

" bypassing security" If they can bypass it there is no security and the security was flawed all along.

Due to the complexity of modern systems, all systems are flawed.
Post reply on HN