Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

11–20 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#11
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.

Hugging face also needs someone arrested for not providing security but that is a lesser charge.

Re: Be skeptical of OpenAI's rogue hacker agent story

#12

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed.

Why the OpenAI escape is the most worrying AI mishap yet

https://www.economist.com/science-and-technology/2026/07/22/...

https://news.ycombinator.com/item?id=49016378

Re: Be skeptical of OpenAI's rogue hacker agent story

#13

does the article end at " How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control? " or is there more that is paywalled? if thats it, the whole article boils down to just " its good marketing so maybe dont believe it " which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of col…

I wish the NPR news broadcasters on the radio yesterday had read the "it's good marketing so maybe don't believe it" angle instead of just parroting the OpenAI press release. Getting that message out would be enormously helpful in countering the blatant submarine marketing "Oh no our AI is a super hacker" with a side of "please regulate super hacker AI and stop those pesky open weights Chinese models that are destroying our stock valuation."

Re: Be skeptical of OpenAI's rogue hacker agent story

#14
post #9

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

Didn’t they explicitly remove alignment guardrails for this test? From the press release:

> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Re: Be skeptical of OpenAI's rogue hacker agent story

#15
post #11

I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.

[flagged]

Re: Be skeptical of OpenAI's rogue hacker agent story

#17
post #14
post #9

Earlier quoted context omitted.

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.

Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

Re: Be skeptical of OpenAI's rogue hacker agent story

#19
post #11

I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.

[flagged]

As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.

Re: Be skeptical of OpenAI's rogue hacker agent story

#20

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...

Post reply on HN