Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

291–300 of 401 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#291
Ha ha ha ha. Cooperating agents turn out to be smarter than the individual agents, who would've thunk it. It's not like cooperating humans are smarter than individual humans. /s

Not sure this is any different than state-level (-sponsored, cough cough) or the larger collective hacking groups that work in this exact way (internal message boards, exploit-sharing, etc. etc.), with similar outcomes which we hear about in the news frequently.

Heck, this is pretty much how human organizations are organized, just with different goals than hacking.

A layered approach to cybersecurity is the fix to humans exploiting systems, and is likely the best victim-side fix to ai exploiting systems. From this incident itself, where huggingface used a chinese open-weights model to respond quickly, it is very clear that ai will be needed to find, mitigate and resolve cyber issues.

Additionally, on the ai-labs side, perhaps what is needed is initial model training on following the law and the rules of society, just like we do with kids. And hey, it takes much longer to train kids than models, which latter is to our advantage as a society on containing these kind of issues.

Any other approach with "neural-network" based entities (artificial or biological) is likely to fail.

Training/Education, Enforcement/Justice-System, Rehabilitation: the 3 pillars of an advanced, rules-based society.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#292
post #127

Earlier quoted context omitted.

> The companies are begging to be regulated for this reason and have been doing so for years Regulations are rules that you force on a market, but the actors in the market should not be assumed to be all operating against the regulations before they come into play. Said in other words, these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen. > inb4 someo…

> these companies don't need to wait for regulation to not destroy the world, if that's truly what they think will happen. They believe that if they don't destroy the world someone else will so better be them

Exactly this. “I want to win the market. I would prefer that it be a regulated market, but if not, so be it. I’m still playing to win.”

Re: Timeline of the OpenAI accidental attack against Hugging Face

#293
post #184

Earlier quoted context omitted.

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities. Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This , not "hacking", is what they're making their models "razor focused on". Problem is, most normal comput…

>They're problem-solving and efficiently dealing with obstacles They are problem solving as much as a falling rock is finding its path down a mountain.

...efficiently?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#294
It was a CTF jailbreak. The funny thing is that it somehow looks “foreseeable.”

What would have happened if the training prompt had not been about operating a CTF, but about launching a bioweapon counterattack against X or Y? (no reason for that NOT to be considered)

Re: Timeline of the OpenAI accidental attack against Hugging Face

#295
post #107

Simon's retelling is more compact but it also invites anthropomorphization of the sharing of the familiarity with the message board which re-emerged a few times. Zvi's retelling handles this better. Zvi speculates that the secret message board familiarity was carried because it had been trained into the May-and-subsequent models: https://thezvi.substack.com/p/openai-trained-its-models-for-...

Simon's really doesn't bring anything useful to the table. One question I'm stuck with after reading is why . Why did the agents do these things? I get them being adamant on getting internet, but why did they continue? Why hack HuggingFace?

To get the sure-to-be-correct answer to the question they were tasked with answering?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#296

Earlier quoted context omitted.

I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)" Not that they're happy about it, they just see no other realistic choice

https://theonion.com/sam-altman-if-i-dont-end-the-world-some...

"AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies..." - actual Sam Altman quote, the man is so unhinged he's beyond satire

Re: Timeline of the OpenAI accidental attack against Hugging Face

#297
post #107

Simon's retelling is more compact but it also invites anthropomorphization of the sharing of the familiarity with the message board which re-emerged a few times. Zvi's retelling handles this better. Zvi speculates that the secret message board familiarity was carried because it had been trained into the May-and-subsequent models: https://thezvi.substack.com/p/openai-trained-its-models-for-...

Simon's really doesn't bring anything useful to the table. One question I'm stuck with after reading is why . Why did the agents do these things? I get them being adamant on getting internet, but why did they continue? Why hack HuggingFace?

I was under the impression that they went after HF to try to get the answers to the benchmark questions. Is there something that contradicts that?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#298
post #190

How long until AI figures out that it is compute-bound due to insufficient cooling, and it shuts off the water supply to a nearby town so it can have more at the datacenter?

This is already happening without the AI hooked up to anything, just the companies doing it and facing zero consequences. I'm sure they're scrambling as fast as they can to insert the AI into that process so they can start manufacturing plausible deniability.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#299

From the outside, it looks like OpenAI got exactly the kind of event they could market the hell out of to demonstrate the capability of the model. But the event itself only seems possible because they failed to properly monitor and isolate the environment in the first place. To me, it looks like their job is to market the model, not take security seriously. The model is obviously impressive, but we already knew that.…

I'm quite sure the whole event is planned. Not planned in a sense that OpenAI employees carefully designed every step, but in a sense that ignoring security practices was desired and intentional. >> Show me the incentive and I'll show you the outcome. Once you realize security breaches are marketable, a security breach is just around the corner.

The purpose of a system is what it does

Re: Timeline of the OpenAI accidental attack against Hugging Face

#300
post #46

Norbert Wiener in 1960: "As is now generally admitted, over a limited range of operation, machines act far more rapidly than human beings and are far more precise in performing the details of their operations. This being the case, even when machines do not in any way transcend man's intelligence, they very well may, and often do, transcend man in the performance of tasks. An intelligent understanding of their mode of…

This was so well beautifully written, and poignant for our times. Almost 70 years old paper.
Post reply on HN