Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

101–110 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#101

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

My take is why not both? It’s good cover to use something that’s eventually going to happen and their name will be associated with the first.

Re: Be skeptical of OpenAI's rogue hacker agent story

#102

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

Never trust somebody who has so much to gain to be entrusted by you. Best lies are not 100% wrong, they have just the right amount of "right" to fool you.

Re: Be skeptical of OpenAI's rogue hacker agent story

#103
post #38

I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid. AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up. Why is today's case so shocking…

Some people oh so desperately want the AI revolution to be a nothingburger.

AI can't be an actual powerful, dangerous technology! Thus, any indication that an AI may attempt concerning things or may possess dangerous capabilities must be secretly a marketing effort!

Especially if an AI has actually succeeded at pulling off a concerning thing out in the wild. Can't have that happen in real life! Nuh-uh! Must be staged!

Re: Be skeptical of OpenAI's rogue hacker agent story

#105

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

The article claims that AI companies have run stunts like this since day 1.

It does not claim that capabilities are not real.

Re: Be skeptical of OpenAI's rogue hacker agent story

#106

Earlier quoted context omitted.

The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score". It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it. Based on the fac…

I think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily on…

> the extent to which these things go

That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.

Re: Be skeptical of OpenAI's rogue hacker agent story

#107
post #9

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

No, a hacking benchmark was exactly what it was tasked with. It wasn't its way to bake a cake.

Re: Be skeptical of OpenAI's rogue hacker agent story

#108
OpenAI has thousands of smartest developers on earth that somehow dropped the ball on the most basic safety hygiene when it comes to sand-boxing that even a high school student knows how to set up.... If that actually happened we are fucking doomed anyways, but hard to believe and most likely its a marketing scheme... which also honestly doesn't bode well.

Re: Be skeptical of OpenAI's rogue hacker agent story

#109
post #79

Earlier quoted context omitted.

Dude.. what are you smoking..? It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct. Meaning you will have to spend more time and tokens to actually get it to do what you want it to do. What do the bible and morals have anything to do with it? I’m criticizing the behavior I see even in the current models. You ask it to do A and instea…

This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores. The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.

Do you think that works? Just prompt a model "be good" and it stops doing anything bad?

It never fucking worked that way and maybe never will.

Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs.

Saying "don't exploit the box please pretty please" might actually cause an LLM to exploit the box more often, for bizarre "don't think of a pink elephant" reasons. 3% rate of exploiting the box (no prompt) -> 11% rate of exploiting the box (with prompt). Because fuck you, that's why. Increased salience -> increased incidence. Welcome to AI tech - good luck and have fun.

Frankly, I expect weirdness like this to be even worse in internal unreleased models that had their behavior fried with who knows what experimental training techniques.

Re: Be skeptical of OpenAI's rogue hacker agent story

#110
post #9

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

>2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. >3) Huggingface has no security and the AI broke in using standard script kiddie methods. Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?

The home directory rm situation also adds credence to this take. The Claude series is much better aligned in comparison.
Post reply on HN