Be skeptical of OpenAI's rogue hacker agent story
171–180 of 321 posts
Re: Be skeptical of OpenAI's rogue hacker agent story
#172Earlier quoted context omitted.
This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores. The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.
Do you think that works? Just prompt a model "be good" and it stops doing anything bad? It never fucking worked that way and maybe never will. Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs. Saying "don't exploit the box please pretty please" might actually cause…
In the Sopranos, there's an episode where a coffee shop protection racket is ruined because a local shop is replaced by a corporate chain that accounts for every cent daily, and immediately fires any employee involved in a discrepancy. In this case, the theft was prevented not by convincing the mobsters of the immorality of their actions - they simply had their harness replaced with one that no longer facilitated the bad behavior.
Re: Be skeptical of OpenAI's rogue hacker agent story
#173Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
Totally evidence-free speculation presented as fact. The average Hacker News thread about AI feels like reading /r/conspiracy.
Re: Be skeptical of OpenAI's rogue hacker agent story
#174Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value
Re: Be skeptical of OpenAI's rogue hacker agent story
#175Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently.
we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it
Re: Be skeptical of OpenAI's rogue hacker agent story
#176There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…
Re: Be skeptical of OpenAI's rogue hacker agent story
#177Earlier quoted context omitted.
> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…
> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I think that depends on the interests and sophistication of the subgroup-of-investors. If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning…
Re: Be skeptical of OpenAI's rogue hacker agent story
#178Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…
> Huggingface covered it up They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026
Re: Be skeptical of OpenAI's rogue hacker agent story
#179Re: Be skeptical of OpenAI's rogue hacker agent story
#180Earlier quoted context omitted.
Because a crime was committed, and a pretty serious one at that? If I blow up your house or steel $100000 from you and we both resolve our differences out of band, should I just be allowed to go about my day like I never did anything, or should I be punished for the crimes I committed? If I am not punished, it makes a mockery of the law that is (supposed) to have protected you, and if it happens repeatedly people wil…
If I steal $100000 from you, and you decide you value your relationship with me more than you value that $100000? Yep, you can just decide to let me have it and let it slide. Believe it or not, deciding that you weren't wronged and not suing isn't a crime. It happens all the time. What people do with each other is up to them. Did Bob allow his friend Jack to borrow his truck? No. Does Bob want to sue Jack for taking…
It is possible for a prosecutor to decide it is not in the public interest to persue a prosecution (or that there isn't sufficient evidence to prove a criminal action beyond reasonable doubt), and certainly the victim's opinion could be considered, but ultimately whether criminal charges should be pursued is and ought to be based on a different test to civil matters, one focused on the public interest rather than mere restoration.