Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

111–120 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#111

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

Oh come now. It is very much in OpenAI's best interest to make it sound to naive people that that "AI" decided to go rogue on it's own.

Re: Be skeptical of OpenAI's rogue hacker agent story

#112

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

[deleted]

Re: Be skeptical of OpenAI's rogue hacker agent story

#113

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed. Why the OpenAI escape is the most worrying AI mishap yet https://www.economist.com/science-and-technology/2026/07/22/... https://news.ycombinator.com/item?id=49016378

The reason for this is simple. There aren't any (or few) people who understand how AI/LLMs actually work employed by these organisations.

Having said that, if knowledgeable people were to write these articles, you'd end up with boring, dry, truthful content.

Re: Be skeptical of OpenAI's rogue hacker agent story

#114
post #17
post #14

Earlier quoted context omitted.

Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model.

In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own.

Re: Be skeptical of OpenAI's rogue hacker agent story

#115

Earlier quoted context omitted.

I think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily on…

> the extent to which these things go That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.

That's exactly what it feels like they are telling it within self-improving loops or something, when they should be prioritizing how to get the best effective output alongside humans and how our training process effects ground truth. They are just making it sound all-knowing by whatever means necessary and them marketing it as god for the most part.

Re: Be skeptical of OpenAI's rogue hacker agent story

#116
post #11

I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.

What? Both sides are cool with it, why would anyone be arrested and etc.?

Because if a human, say a Aaron Swartz type, were to have done it, they'd destroy him.

Re: Be skeptical of OpenAI's rogue hacker agent story

#117

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

By now I'm pretty confident that some people will keep screeching , "it's not a marketing stunt, these companies really are justified in their trillion dollar valuations, please please please believe them." Even when the economy is crashing around their feet.

"It's not a marketing stunt" is just delusion trying to appear measured and safe.

Re: Be skeptical of OpenAI's rogue hacker agent story

#119

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

Oh come now. It is very much in OpenAI's best interest to make it sound to naive people that that "AI" decided to go rogue on it's own.

[deleted]

Re: Be skeptical of OpenAI's rogue hacker agent story

#120

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

Nobody is saying the models are incapable of what was claimed.
Post reply on HN