Live data from Hacker News

OpenAI’s accidental attack against Hugging Face is science fiction that happened

simonwillison.net

241–250 of 475 posts

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#241
post #236
post #232

Earlier quoted context omitted.

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

Alignment isn't alignment if it can be turned on and off at the whim of company employees.

This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#242
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

Is there a source? Or is that xenophobia?

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#243

Earlier quoted context omitted.

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

Are you serious? Before LLM’s you needed serious skills and experience to pull this off. Now it’s a prompt away on some terminal done by any random dud. And I dont mention the velocity of iteration or that they will be even better in 1 year.

Its like you don't even read the post or the content.

Hint: agentic loops.

No its not a prompt away.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#244
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I wonder how long it takes before someone instructs an LLM to design and launch a "Morris Worm 2.0" and cripple the Internet for a good while. Might even wind up happening by accident (again).

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#245
post #129

The technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of ." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the def…

I believe the Russians and Chinese recognized this years ago, which is why they are using their propaganda machines to make Americans hate datacenters.

American Capitalists don't really need any help there; they're making Americans hate Data Centers all on their own.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#246
post #236
post #232

Earlier quoted context omitted.

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior.

I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces, the fact that guardrails can just be turned off, and these kind of incidents, I am not reassured. We may well get another "oopsie" moment with much more catastrophic consequences even from otherwise well intentioned actors.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#247

Earlier quoted context omitted.

Stuxnet misconfigured industrial equipment that was already set up to run, and all it did was break that equipment. This scenario sets a much, much higher bar.

I guarantee there's misconfigured chemical analysis equipment out there exposed to the internet. I don't think the grandparent was implying the AI would be controlling robot arms to mix things directly (or at least I didn't interpret as such), but it could very well sit in the network until it notices two dangerous compounds in the same machine, and trigger a breakage that causes a harmful mixture. Break a vial conta…

I have no doubt that somewhere there's chemical equipment an Internet attacker could break, or, maybe, if I have to stipulate, misconfigure enough to, like, poison someone. But my question is about the notion that you could synthesize a specific scary substance.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#248

Earlier quoted context omitted.

>The technology held by private AI companies is warfare-capable technology. This is the precisely the response OpenAI is hoping for to raise its valuation, and you fell for it. Look at it this way - whats the difference between tasking AI to break into something, versus taking a whole bunch of smart humans to do the same? The only difference is that AI is slightly easier to orchestrate. Prior to AI, there were alread…

It lowers the economic cost of a given attack, but also lowers the economic cost of protection. Not sure if it’ll be a perfect balance, but right now there’s a manufactured IMbalance due to embargos and winner picking.

It lowers the economic cost of performing an attack in the same way a gun lowers the economic cost of killing a person. It does nothing for consequences of that attack.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#249
post #112
post #82

Replace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)? That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.

I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime. OpenAI gained access to HuggingFaces production database ffs.

> I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime.

I agree, lets use the favorite analogy. OpenAI encouraged a smart and eager junior engineer to find any way whatsoever to get a higher score on the benchmark. Then, the junior breaks into HuggingFace to get a higher score. That would be a big deal involving the FBI, not press releases and blog posts.

Re: OpenAI’s accidental attack against Hugging Face is science fiction that happened

#250
post #236
post #232

Earlier quoted context omitted.

Echoing other responses to you but this isn't a capability problem. You're totally right that this is so last year in terms of LLMs being capable in infosec. The issue here is an alignment one, i.e. the model seemingly isn't "aware" (especially with its guardrails turned off it would seem) that it is doing something immoral/illegal by hacking HF for the answer to its (vague) query of "solve this problem". Or if it is…

I don't think this exposes an alignment failure, because the test here was run with the alignment features deliberately turned off. OpenAI said: > We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. It was a test of raw capabilities of the underlying model.

This is assuming a situation where, A) models "unaligned" by default and B) alignment can be added (though prompts and related things).

The point, which isn't very surprising but still notable, is that the models are "amoral" out of the box. And moreover, we know that there is almost always a means to "jailbreak" them into that out-of-the-box capability (or that sometimes just randomly "jailbreak" in various ways).

Also, saying the models are amoral doesn't mean they don't know good and evil - once they do acts defined as evil, they know "themselves" through their and so self-define themselves as evils (or objectively predict what a secretly/open evil actor would do based on their data). Which is to say I once laughed at the mis-alignment doomers but I can't see strong barriers against the doom scenario now.

Post reply on HN