Live data from Hacker News

OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

bbc.com

51–60 of 110 posts

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#51

i wonder if they even bothered to roleplay this incident or the PR team just made it up

HuggingFace posted an incident report a week ago, which makes it much more likely that this happened. I understand people are suspicious of OpenAI, but I don't think there's any reason to believe this is a made-up event.

https://huggingface.co/blog/security-incident-july-2026

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#52
> was being tested in a controlled environment, but found vulnerabilities and managed to escape.

Uh-huh.

It's worth keeping in mind that we are here talking about people who simply will not be satisfied until they have built Skynet[0]. They will then no doubt experience some very brief satisfaction before we are all annihilated.

I can't help thinking the juice isn't worth the squeeze.

I'm not sure what the alternative is or how to change course unless and until it becomes unambiguously apparent that such a thing is not possible. Currently it seems like we still think it might be possible, so that's where we're heading because if "we" don't do it, somebody else will.

[0] How feasible this really is, or on what timescale it might be possible, I don't think anyone can really say. But, at least for now, this is clearly the aim and trajectory we are on.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#53

Earlier quoted context omitted.

Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.

Is this a better analogy? The top student giving paid lectures about his classes, and another student skipping class and instead studying those lectures to end up with the second highest grade? Maybe it could be improved with the other student not even going to the same school?

In your example the problem is what exactly?

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#54
post #33
post #16

What I wonder is what will Anthropic come up with on the PR front next.

"Our model is so powerful it seduced all our wives. Now every woman in Silicon Valley is pregnant with AI babies and the machines are taking over" New opportunities for AI in the porn industry. I'm only half joking, it used to be a meme on the internet that all new technologies online were driven by the porn industry.

I thought it was military and porn? We're past half way there.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#55

Earlier quoted context omitted.

That’s the play. That these frontier models are so powerful that they must be behind sovereign firewalls and gateways.

Didn't HF use an open weights model running on their own hardware to solve the issue though? Sort of defeats that narrative and plays into one in which frontier == bad_guys and open == good_guys

Public doesnt know about Huggingface. ChatGPT (OpenAI) says it‘s dangerous. They must know.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#56
post #12

This kind of narrative is going to bite them just like the "AI will take your job" narrative has. It feels like the frontier labs are taking a massive gamble with public perception here. I assume the goal is to paint the technology as so powerful and dangerous that only a handful of blessed US companies should be trusted to run it, in an attempt to suppress the rise of the Chinese models that are rapidly catching the…

That’s the play. That these frontier models are so powerful that they must be behind sovereign firewalls and gateways.

>That these frontier models are so powerful

Maybe powerful might NOT be the right word to describe them, they are just non-deterministic, there for we going to see this kind thing more and more.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#57
post #46

Earlier quoted context omitted.

Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.

they're all introducing themselves as claude for one, there are more quantitive and qualitative arguments elsewhere

All chinese models are introducing themselves as claude? Why make claims trivially disproven?

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#58
post #40

Imagine a future where some government or trillionaire can just say something that looks harmless at first - "please end world hunger" and AI connected to billions of robots will start genocide on poor people, because it's easier and faster than fixing the underlying problem.

The logical choice would be to kill the the (relatively) rich people who eat and waste far more food per capita.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#59
post #51

i wonder if they even bothered to roleplay this incident or the PR team just made it up

HuggingFace posted an incident report a week ago, which makes it much more likely that this happened. I understand people are suspicious of OpenAI, but I don't think there's any reason to believe this is a made-up event. https://huggingface.co/blog/security-incident-july-2026

GPT-3 (I think? I forget which one) supposedly tried to deceive researchers and escape the lab. Or at least that was how it was reported. If you actually clicked through several links, it was a "what would you do if" roleplay.

Re: OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

#60
We need to clear up the responsibility of these "AI went rogue" situations, ASAP.

How serious does it have to get before "oops AI did that not me" stops being a valid defense and we start looking into it?

Because I'd bet my life that if I asked ChatGPT to fix a bug that one of my clients reported, and the model fixed it by outright k*lling the client IRL, I'd be held liable instead of anyone at OpenAI. Just a hunch.

Post reply on HN