Hugging face also needs someone arrested for not providing security but that is a lesser charge.
Be skeptical of OpenAI's rogue hacker agent story
11–20 of 321 posts
Re: Be skeptical of OpenAI's rogue hacker agent story
#12Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
Why the OpenAI escape is the most worrying AI mishap yet
https://www.economist.com/science-and-technology/2026/07/22/...
Re: Be skeptical of OpenAI's rogue hacker agent story
#13does the article end at " How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control? " or is there more that is paywalled? if thats it, the whole article boils down to just " its good marketing so maybe dont believe it " which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of col…
Re: Be skeptical of OpenAI's rogue hacker agent story
#14Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
Re: Be skeptical of OpenAI's rogue hacker agent story
#15I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.
Re: Be skeptical of OpenAI's rogue hacker agent story
#16Re: Be skeptical of OpenAI's rogue hacker agent story
#17Earlier quoted context omitted.
>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
Re: Be skeptical of OpenAI's rogue hacker agent story
#18There's trillions of dollars at stake here. Be skeptical of anything these AI hypesters say.
LLMs seem to be getting more useful though
Re: Be skeptical of OpenAI's rogue hacker agent story
#19I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.
[flagged]
Re: Be skeptical of OpenAI's rogue hacker agent story
#20Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...