Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

421–430 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#422
Should we blindly trust OpenAI's narrative about this event?

It's strangely convenient to arrive at a point when OpenAI was way behing in cybersecurity vs Claude Mythos Fable and everything, Anthropic was making headlines each week, then boom OpenAI inadvertently attacks HuggingFace because their tool is so good it's out of control, so maybe you can buy it and get either protection if you're a company, or a nice tool if you're a cybercriminal, script kiddie, or red team.

What if Sam knew it would happen, either because it was prompted to do exactly that, or without explicitly prompting it, knew that given the parameters of the experiment, knew it was one of the possible outcomes that it didn't harden against this kind of incident deliberately because when they fail, they make wordlwide news and stocks go up?

I don't think I'm putting my head in the sand like Simon says, as I do believe that most frontier models are capable of doing this. I just don't trust AI CEOs to not stage this, especially Sam.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#423
post #252
post #246

Earlier quoted context omitted.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior. I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces…

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the e…

Do you want this supermodel to cause harm whilst finding these vulnerabilities?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#424

Earlier quoted context omitted.

Alignment isn't alignment if it can be turned on and off at the whim of company employees. This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.

A way to help prevent some of that catastrophic damage, is to make companies accountable for what their AIs do. A major problem with AI companies is that they like to point to the AI, as if they're minimally involved innocent bystanders, when that's the furthest thing from the truth.

It’s a shame the target was HF. If it had been a large bank or other institution we might get to see this play out in courts.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#425
post #347

Earlier quoted context omitted.

I (obviously) can't claim to know the specifics of this incident, but I think we can know with near-certainty that as these models surpass human intelligence, the sensation will be exactly what you're describing. The entire point of intelligence is being able to infer information and foresee solutions to problems that less intelligent systems cannot infer or cannot foresee. With regard to this specific incident, as I…

I’m not questioning capabilities. What would it take for the model to know the specific benchmark name and that the answer is in an internal Hugging Face database? Be specific, then wonder how it knew it. Why would they evaluate the model on a benchmark and not watch what it’s saying along the way? It apparently spent more than a whole weekend working on it; Nobody wondered? Nobody looked? You believe they took all r…

> What would it take for the model to know the specific benchmark name and that the answer is in an internal Hugging Face database?

Googling "offensive cyber benchmarks for LLMs"? Figuring out that it was running at Huggingface is as simple as running `hostname`, isn't it?

> Why would they evaluate the model on a benchmark and not watch what it’s saying along the way?

What it's saying is what it wants to say. We already know these models are capable of deception – this has been well-established too. Surely there's an argument to watch closely what actions it's taking, but now we're back to my prior point of "what exactly does that mean?" It's absolutely trivial for these systems to write code that looks like it does one thing but does another. Watching network ingress/egress seems viable but also seems easy to miss the semantic meaning of any set of logs, especially when those logs might be getting generated for the purpose of deceiving observers.

> I don’t think they’re careless

We already know, factually, even the "more careful" lab was fine to release a model that consistently detects it's being evaluated and modifies its behavior accordingly. Why do you think they're not careless? Incentives are a hell of a drug, and there is no incentive to be careful but every incentive to be fast.

> I think more than nobody would have been curious how it’s doing days into a single test question.

Beyond simonw's great point below, your underlying assumption is if someone had looked, they would've noticed. That is a huge, huge assumption. So I'll ask again directly: what specifically do you expect them to be looking at, and how, such that they certainly would have noticed what it was doing? Their "chain of thought" is not actually their thinking. The printed output of their code is not actually what their code is necessarily doing. We know that LLMs can engage in steganography. I don't see what could actually be observed that you can trust except network logs, which again will often be non-trivial to piece into an actual narrative understanding of what's going on (especially if the LLM felt like piling deceptive noise into those logs, too).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#426

My favorite part of this discourse is people somehow finding it preposterous that 2 companies filled to the brim with AI sycophants who regularly lie - and in Sam's case, basically every single word he breathes out is a lie - who have massive vested interests in this tech succeeding couldn't possibly collude together to shore up this facade as a marketing stunt.

There is just no skepticism these days. We are constantly being lied to by tech leaders.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#428
post #86

I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox…

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.

What’s stopping an ai to blackmail someone or pay xxx in crypto to someone working a lab?

And no need to be in a lab, can be any critical infrastructure this personnel. If influenced wrongly I’m sure in many industries 1-2 key people can do big damage.

What’s to stop to pay xxx crypto to a private investigator to find dirt of person yyy for blackmail?

Just money… once ai has money, id say we are one step closer to game over.

Tell me one thing he can’t do with crypto?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#429
post #234

Science fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.

> LLMs are zombies How do you know?

I can easily manipulate them

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#430

Hardly science fiction. I respect a lot of Simon's work, but his credibility gets chiseled away a bit each time he participates in these kinds of echo chamber posts that are basically secondhand marketing.

Which work of his do you respect?
Post reply on HN