Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

1–10 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#2
Finally mainstream news understands. The unfiltered version:

1) The AI failed to solve ExploitGym problems.

2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

3) Huggingface has no security and the AI broke in using standard script kiddie methods.

OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.

Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.

Re: Be skeptical of OpenAI's rogue hacker agent story

#3
does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?

if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.

Re: Be skeptical of OpenAI's rogue hacker agent story

#4
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.

It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.

It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.

Re: Be skeptical of OpenAI's rogue hacker agent story

#5
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?

Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.

Re: Be skeptical of OpenAI's rogue hacker agent story

#6

does the article end at " How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control? " or is there more that is paywalled? if thats it, the whole article boils down to just " its good marketing so maybe dont believe it " which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of col…

[deleted]

Re: Be skeptical of OpenAI's rogue hacker agent story

#9

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.

>3) Huggingface has no security and the AI broke in using standard script kiddie methods.

Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

Post reply on HN