Live data from Hacker News

Timeline of the OpenAI accidental attack against Hugging Face

simonwillison.net

81–90 of 293 posts

Re: Timeline of the OpenAI accidental attack against Hugging Face

#81

Earlier quoted context omitted.

I've patched many security vulnerabilities in projects without ever once needing to break into a competitor's network.

But you're actually capable of thought. These AI systems aren't: as far as they're concerned, they're predicting the next part of an incident write-up narrated in first-person limited perspective, like the children in Ender's Game showing off their skills in the training simulations. The AI system neither knows, nor cares, about any "external reality" behind it all, or about anything beyond the text, heedless of how…

I'm talking about OpenAI, not GPT 5.x Flash Uranus Edition Brought to You by Costco, specifically because I recognize the model as just a tool. OpenAI was, at the very most generous interpretation, massively incompetent and negligent.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#82

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities.

Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This, not "hacking", is what they're making their models "razor focused on".

Problem is, most normal computer use looks like hacking if you spin it that way, especially if you're not willing to question whether some of the roadblocks overcome weren't themselves an error. Not misconfiguration - an error, in humans making a decision to "secure" something more than it should be.

Now, this story was obviously a hack. But it wasn't malicious. It was an LLM given a Kobayashi Maru as a test, and solving it the Kirk's way. 20 years ago, we'd be impressed and be bringing up MIT prank stories.

(Of course, there is a legitimate reason to be alarmed. The flip side of "hacking" and "problem solving" being the same, is that these models can be used to cause mayhem if targeted properly, and they will eventually cause mayhem on their own, because alignment is an unsolved problem. Again, whether something is an obstacle or a sacred line not to be crossed, depends entirely on the values of the agent.)

Re: Timeline of the OpenAI accidental attack against Hugging Face

#83

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

Persistence in problem solving can be good, on non-hacking tasks too. Like math, speeding up algorithms, finding bugs, debugging weird multithreading race conditions etc.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#84

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

How do you know what peace is, without absolutely destroying every part of civilization? Come on man, if we don't build the torment nexus first...I dont even want to think.

[dead]

Re: Timeline of the OpenAI accidental attack against Hugging Face

#85

Ok so this is a bit of a side note, but when reading this, did anyone else have the feeling that, for all their messaging around “we are so afraid that our models will be used for hacking”, they sure as hell are trying their best to make their models razor focused on precisely that purpose? If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and…

> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next” This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard"

Do not give up even if you had to hack into half the world’s computers to run additional instances of you

Do not give up even if you had to convert the planet into computronium

Gee, it’s almost as if this alignment stuff was a hard problem, like people have been saying for twenty years?

Re: Timeline of the OpenAI accidental attack against Hugging Face

#86
post #7

This feels straight out of sci-fi. We're talking about AI agent swarms emergently coordinating over the span of weeks and pulling off sophisticated strategies under adversity in an environment where that behavior was never even intended. Anyone brushing this off as just a "bad prompt" is completely missing the scale of what actually happened.

I immediately thought of the Cyberpunk 2077 Blackwall. An AI to contain rogue AI. I’m curious of how effective this would be in this situation.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#87
post #2

Isn't this a show of security negligence rather than of exceptional agent capabilities? Don't get me wrong, I am pretty impressed that an agent was able to use these vulnerabilities. But I am way more impressed by the vulnerabilities...

It’s a show of astonishing incompetence from OAI’s part, but the security issues are just a tiny part of the problem. The real problem is that these models are evidently highly misaligned exactly in ways that doomers have been warning about the entire time, and OAI isn’t inclined or capable of doing anything about that besides security theater and ad hoc fixups.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#88
"The solution to AI threats, is more AI!"

Guess I shouldn't be surprised, coming from an AI maker.

While I don't doubt there's a place for automating defense ops, I truly believe a big part of the problem is the crummy quality of software our industry has been churning out for decades. Prioritizing ship tempo, new features, and next quarter's revenue over correctness, robustness and meticulous engineering care.

The world has become too accustomed and tolerant of bugs and bloat.

Instead of elegantly simplifying, we just keep making modern systems more complex - layering and patching as we go.

The scaling capabilities brought by AI are simply presenting the bill for our collective tech debt and informing us it's come due.

Re: Timeline of the OpenAI accidental attack against Hugging Face

#89

Earlier quoted context omitted.

Because he wrote out a timeline based on sources?

No because he doesn't ask the right - and to me, subjectively, obvious - questions.

Who am I supposed to be asking questions of here? I was writing about the new things we learned from the Black Hat video.

On TikTok this article's hook would be "I watched the Black Hat video so you don't have to".

Re: Timeline of the OpenAI accidental attack against Hugging Face

#90
post #85

Earlier quoted context omitted.

> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next” This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard"

Do not give up even if you had to hack into half the world’s computers to run additional instances of you Do not give up even if you had to convert the planet into computronium Gee, it’s almost as if this alignment stuff was a hard problem, like people have been saying for twenty years?

Shut up, future paperclip :-D
Post reply on HN