Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

141–150 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#141
post #17

Earlier quoted context omitted.

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model. In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own.

Where did it say that? To me it did not read like that:

> GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes

> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Re: Be skeptical of OpenAI's rogue hacker agent story

#142
post #124

Earlier quoted context omitted.

What? Both sides are cool with it, why would anyone be arrested and etc.?

Because a crime was committed, and a pretty serious one at that? If I blow up your house or steel $100000 from you and we both resolve our differences out of band, should I just be allowed to go about my day like I never did anything, or should I be punished for the crimes I committed? If I am not punished, it makes a mockery of the law that is (supposed) to have protected you, and if it happens repeatedly people wil…

I agree with your stance mostly. I get the sentiment, but both of your examples seem civil to me. Both of those situations can and ideally should be resolved out of the courts, and neither carry social weight.

If one entity is injured by another, and subsequently made whole, however the two parties define that, then it is none of my business.

Re: Be skeptical of OpenAI's rogue hacker agent story

#143
post #23

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

Why not report it? It's still illegal to open a door barred with a piece of cardboard, or to enter a house with no door.

that's the biggest indication of this just being a marketing move to me. I would expect a third party breaking in to huggingface would at the very very least be banned forever.

Re: Be skeptical of OpenAI's rogue hacker agent story

#144
post #126

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> AI managed to escape using standard and well documented script kiddie methods > AI broke in using standard script kiddie methods. I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where…

From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet.

Re: Be skeptical of OpenAI's rogue hacker agent story

#146

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…

> Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.

I think that depends on the interests and sophistication of the subgroup-of-investors.

If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning worldly attachments.

Re: Be skeptical of OpenAI's rogue hacker agent story

#147

Earlier quoted context omitted.

Firstly: yes, very many people are saying this. Secondly: to the people who aren't saying it....then why are you bringing up marketing at all? If the model is capable of it, then the motivation for why OpenAI is talking about it/reporting on it is completely beside the point. Either the capability matters or it doesn't. If the capability doesn't matter, or doesn't matter in the way that some particular person is clai…

I think you’re missing the fact that no one is saying not to be concerned, quite the contrary. OpenAI is using the threat of how powerful its models are to bolster support for regulation in which it’s one of the only players that’s allowed to use the capability. Conveniently, that would also be a moat that makes them more valuable to investors. The marketing of their models as super dangerous has a direct link to the…

Again: yes, very many people (just look through this very thread) are saying that. Tons and tons of people are saying some combination of "the whole thing is made up/a lie" or "This is entirely down to OpenAI incompetence and doesn't matter", or for some other reason (often not stated) making it very obvious that they do not think that this story matters very much.

And, again: to the people who aren't saying that: whatever argument is being made, it seems to me like it probably doesn't need to rely on claiming anything about the motivation of OpenAI. It sounds like you are against government regulation of AI. That's a position that a totally reasonable person could have. You should be able to argue that this event does not justify some particular kind of government regulation without reference to OpenAIs motivation for reporting the story.

Re: Be skeptical of OpenAI's rogue hacker agent story

#148
post #17

Earlier quoted context omitted.

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.

Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given

Re: Be skeptical of OpenAI's rogue hacker agent story

#149

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

I'm inclined toward your same viewpoint here, except OpenAI's announcement of this came with a blog post that gave off a real bragging/marketing vibe.

To me it's clear that even if it was an accident, they kinda liked it, and don't have this incentive to invest that much in preventing it.

Re: Be skeptical of OpenAI's rogue hacker agent story

#150

Earlier quoted context omitted.

Firstly: yes, very many people are saying this. Secondly: to the people who aren't saying it....then why are you bringing up marketing at all? If the model is capable of it, then the motivation for why OpenAI is talking about it/reporting on it is completely beside the point. Either the capability matters or it doesn't. If the capability doesn't matter, or doesn't matter in the way that some particular person is clai…

I think you’re missing the fact that no one is saying not to be concerned, quite the contrary. OpenAI is using the threat of how powerful its models are to bolster support for regulation in which it’s one of the only players that’s allowed to use the capability. Conveniently, that would also be a moat that makes them more valuable to investors. The marketing of their models as super dangerous has a direct link to the…

The vibe in this thread is literally everyone going "script kiddie, it's all smoke and mirrors, marketing stunt" and acting like well that settles it.

Meanwhile literally no one has actually shown the script kiddie thing - nikcub's apparently looked.

Sure, moat, sure, OpenAI milks this, both can be true, but literally has zero to do with misalignment.

Just frustrating that folks are still having this collective delusion about the capabilities of the model.

respectfully.

Post reply on HN