Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

171–180 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#172

Earlier quoted context omitted.

This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores. The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.

Do you think that works? Just prompt a model "be good" and it stops doing anything bad? It never fucking worked that way and maybe never will. Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs. Saying "don't exploit the box please pretty please" might actually cause…

There is a clear difference between saying not to do something because it's immoral, and saying doing that thing would be futile.

In the Sopranos, there's an episode where a coffee shop protection racket is ruined because a local shop is replaced by a corporate chain that accounts for every cent daily, and immediately fires any employee involved in a discrepancy. In this case, the theft was prevented not by convincing the mobsters of the immorality of their actions - they simply had their harness replaced with one that no longer facilitated the bad behavior.

Re: Be skeptical of OpenAI's rogue hacker agent story

#173

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

Totally evidence-free speculation presented as fact. The average Hacker News thread about AI feels like reading /r/conspiracy.

It's hard to understand something (AI is quite capable) when your salary depends on you not understanding it (AI will replace you).

Re: Be skeptical of OpenAI's rogue hacker agent story

#175

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

I can agree with you on points 1,2, and 3 and still find it important and concerning news. AI have found real world 0 days before, we’re seeing tons of security patches coming in. Open weight models are catchy up. Right now everyone is at risk from this technology as is perhaps something big will capture headlines soon but we’re just gpu constrained from bad actors being able to wield them successfully.

Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently.

we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it

Re: Be skeptical of OpenAI's rogue hacker agent story

#176

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

In any case it just shows that these models aren't properly aligned. Instead of trying to solve tests they try to find ways to cheat.

Re: Be skeptical of OpenAI's rogue hacker agent story

#177
post #146

Earlier quoted context omitted.

> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…

> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I think that depends on the interests and sophistication of the subgroup-of-investors. If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning…

AI investors are not sophisticated users nor product managers. They are by and large not technical at all. They are bureaucrats at a teachers' pension fund in the midwest, unscrupulous dealmakers at private credit firms, and Masayoshi Son. Actually go read what Masayoshi Son says about AI if you want to understand the level of due-diligence we're dealing with.

Re: Be skeptical of OpenAI's rogue hacker agent story

#178

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> Huggingface covered it up They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026

Not only did they not cover it up they also said open weight models are needed over closed ones. Hardly a thing you’d say partnering with the OpenAI

Re: Be skeptical of OpenAI's rogue hacker agent story

#180
post #124

Earlier quoted context omitted.

Because a crime was committed, and a pretty serious one at that? If I blow up your house or steel $100000 from you and we both resolve our differences out of band, should I just be allowed to go about my day like I never did anything, or should I be punished for the crimes I committed? If I am not punished, it makes a mockery of the law that is (supposed) to have protected you, and if it happens repeatedly people wil…

If I steal $100000 from you, and you decide you value your relationship with me more than you value that $100000? Yep, you can just decide to let me have it and let it slide. Believe it or not, deciding that you weren't wronged and not suing isn't a crime. It happens all the time. What people do with each other is up to them. Did Bob allow his friend Jack to borrow his truck? No. Does Bob want to sue Jack for taking…

A lawsuit is not the same as a criminal prosecution. One is supposed to remedy damages, the other is supposed to be because it is in the public interest that people follow laws and we don't devolve into anarchy (or more dangerously, some might-makes-right society where it is more beneficial to appease the powerful than it is to pursue what is just).

It is possible for a prosecutor to decide it is not in the public interest to persue a prosecution (or that there isn't sufficient evidence to prove a criminal action beyond reasonable doubt), and certainly the victim's opinion could be considered, but ultimately whether criminal charges should be pursued is and ought to be based on a different test to civil matters, one focused on the public interest rather than mere restoration.

Post reply on HN