Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

251–260 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#251
post #124

Earlier quoted context omitted.

Because a crime was committed, and a pretty serious one at that? If I blow up your house or steel $100000 from you and we both resolve our differences out of band, should I just be allowed to go about my day like I never did anything, or should I be punished for the crimes I committed? If I am not punished, it makes a mockery of the law that is (supposed) to have protected you, and if it happens repeatedly people wil…

I agree with your stance mostly. I get the sentiment, but both of your examples seem civil to me. Both of those situations can and ideally should be resolved out of the courts, and neither carry social weight. If one entity is injured by another, and subsequently made whole, however the two parties define that, then it is none of my business.

Corporations commonly hide criminal activity in order to save face.

In the oil industry there is a portion called land management where the portions of oil and gas from wells can be split across a large number of entities. This can lead to numerous complexities that open up opportunities for fraud/theft in division of the profits. Quite often it is easier for the corporation to cover up that this occurred and pay off the person never to talk under NDA about it rather than have to have their customers find out and potentially take millions in losses.

Computer related hacks are very similar. Quite often these are covered up and never disclosed unless the information shows up in public at some point.

Re: Be skeptical of OpenAI's rogue hacker agent story

#252

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

Considering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO. In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3? If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile…

Yeah, I'm reminded of container escapes, VM escapes etc. There have been plenty in the past; VirtualBox E1000 (I think?) comes to mind from a few years ago. If we are to believe this model can find 0days, I'm on board with the idea it could do so in sandbox.

That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.

Re: Be skeptical of OpenAI's rogue hacker agent story

#253

Earlier quoted context omitted.

>> Agents don't work on their own > This is factually false From your link: > After investigating, we now know that this particular incident was driven by a combination of OpenAI models...while being internally tested on a benchmark of cyber capabilities. Someone set up that test and started it. Whether they outsourced the majority of the work in "setting up" and "starting it" to an LLM or not, they still set it in m…

Setting up is not operating. There's no indication of there having been a human in the loop during its operation: nobody was approving its tool calls, and nobody instructed it to commit these specific actions during its run (via prompting or steering). There's no indication of any supervision of its operation either: OpenAI's engineers acted with significant delay, long after the agent has already meandered its way t…

Unlawful action is quite a messy definition. Is performing a vulnerability scan illegal when it's done internally? What about when there's a device on the network that port forwards information to another server you weren't aware of?

Re: Be skeptical of OpenAI's rogue hacker agent story

#255
post #245

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

All three could be true. Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this. Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme. Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.

Every single time. The next model is always so infinitely powerful it is going to change everything. Said model comes out. Is marginal improvement. Changes nothing and seemingly cannot do what they purported it could do except under the very specific circumstances of the demo.

How many times are we going to go through this.

Re: Be skeptical of OpenAI's rogue hacker agent story

#257

Earlier quoted context omitted.

Basically, they shot someone in public to promote their cool new gun.

Eh, closer to "they pointed a gun at someone and dropped it so it fired at them to promote their cool new self-aiming gun," from the perspective of the skeptics here.

Technically the weapon picked who it wanted to shoot.

Re: Be skeptical of OpenAI's rogue hacker agent story

#259

Yeah, I mean the last 3 or four big releases of models have come with big scary news. I have to think there is at least some intentional marketing effort behind all this

Or, models are potentially dangerous and it works out in the favor of regulatory capture.

I imagine it like the bumper stickers that say "legalize recreational plutonium"

Re: Be skeptical of OpenAI's rogue hacker agent story

#260

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

> The way OpenAI seems to want

This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.

> The second seems to forget that jailbreaks are available for every model

Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.

This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.

You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:

"2. OpenAI’s harness and network security controls were unintentionally [...] bad"

Post reply on HN