Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

31–40 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#31
post #17
post #14

Earlier quoted context omitted.

Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.

Re: Be skeptical of OpenAI's rogue hacker agent story

#33
post #22

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.

The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.

They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added."

Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.

Re: Be skeptical of OpenAI's rogue hacker agent story

#34
post #19

Earlier quoted context omitted.

[flagged]

As agents become more and more powerful, it would be good to get clear legislation or precedent in place that makes either model creators (OpenAI) or operators (whoever is running the model) liable for their agents' actions.

If you're smaller than OpenAI, you are liable. If you're larger than OpenAI, they are liable.

Re: Be skeptical of OpenAI's rogue hacker agent story

#35

Earlier quoted context omitted.

You aren't going to locate incontrovertible evidence that what happened as or wasn't engineered. Anything like that is going to be private and that is unlikely to change. And that's not really an interesting question anyway. As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details. We are missing, for instance, prompts that were involv…

right. and all that is fine. im just not exactly sure why the basics of media literacy are worthwhile on the site that “optimizes for curiosity”. i saw the domain and thought it was going to be some cool investigative journalism about the incident rather than “be skeptical. the end.”

I hate to be so obnoxious but I think you might be overestimating the average user of this website with regard to media literacy, and if that's the case, then the basics of it seem very relevant for the homepage

Re: Be skeptical of OpenAI's rogue hacker agent story

#36
post #11

I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.

What? Both sides are cool with it, why would anyone be arrested and etc.?

Because incentives are aligned properly if agents aren't liability-proof for crimes- making someone go to jail for instances like this is how to get the labs to behave themselves, whereas going "ha ha what an oopsie-woopsie" will make the next instance worse.

Note that for criminal cases (which this was), the justice system can choose to prosecute even if the victim doesn't want that. It often doesn't, but this is one case where it should.

There’s also an optics issue for the justice system at play here: there’s immense public distrust of and anger at the labs right now. I would go to jail if I hacked HuggingFace, even if I said “it was during an eval!”; not doing the same for the labs makes it look like they’re above the law, which is going to make this anger get worse.

Re: Be skeptical of OpenAI's rogue hacker agent story

#38
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.

AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.

Why is today's case so shocking?

Re: Be skeptical of OpenAI's rogue hacker agent story

#39

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

I worry that cynicism about this:

> if not all was invented and everything was scripted in the first place in order to get desired regulations

ends up covering up what is more worrying:

> OpenAI sandbox is such a horrible hack

I am more worried that this is sloppiness with potentially harmful resources than I am worried that people are juicing the stock price.

Re: Be skeptical of OpenAI's rogue hacker agent story

#40
post #11

I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised. Hugging face also needs someone arrested for not providing security but that is a lesser charge.

> Agents don't work on their own

This is factually false: they both can and clearly did operate in an autonomous and unsupervised manner: https://openai.com/index/hugging-face-model-evaluation-secur...

This does not require sentience, personhood, a soul, or anything of the sort. It further doesn't mean an erasure of legal responsibility, not in principle, and not in historical practice.

I wish people would finally stop with the spiritualistic reasoning around this.

Post reply on HN