Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

301–310 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#301
post #210

Earlier quoted context omitted.

Gross negligence can stand in for intent. I believe a rather compelling case could be developed on the basis that a system believed to be capable of this was developed and improperly secured.

But what’s the crime? The thing got out of its poorly secured cage and caused what damage? Imagine instead of an agent it was a dog or a chimpanzee that got into the neighbor’s yard and the particular neighbor isn’t even pissed about it.

Aaron Schwartz was being prosecuted and threatened with 35 years in jail for the crime of saving research papers to a thumb drive. What damage did he cause?

Re: Be skeptical of OpenAI's rogue hacker agent story

#302

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

4. OpenAI hacked HuggingFace on purpose and got caught

Re: Be skeptical of OpenAI's rogue hacker agent story

#303
post #46

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

Are you suggesting the Suchir Balaji case was not investigated?

I was hoping to bait a conversation, because I sure as hell don't think it was. Thanks for having the nads to mention it here.

Re: Be skeptical of OpenAI's rogue hacker agent story

#304

Earlier quoted context omitted.

Oh come now. It is very much in OpenAI's best interest to make it sound to naive people that that "AI" decided to go rogue on it's own.

How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?

Current admin will not regulate them.

This is them essentially bragging how powerful and autonomous their "AI" is. It isn't scaring their real customers or employees to talk like this.

Re: Be skeptical of OpenAI's rogue hacker agent story

#305

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…

I'm disappointed that this is the level of discourse happening here, when the default assumption is such conspiratorial thinking. I expect that from tiktok and low-information social media, not here.

When your default explanation for everything is "companies are lying about everything", you end up just as incorrect as believing they're always telling the truth.

It's not outlandish that models have these capabilities, the number of CVEs I see as a sysadmin has exploded and we're seeing novel math discoveries nearly every week now. OpenAI does not need to pretend to commit a felony to demonstrate it, that's pure conspiratorial thinking.

Re: Be skeptical of OpenAI's rogue hacker agent story

#306
post #7

Crazy that we need reminders not to take everything we read in corporate press releases and marketing material at face value

Based on the comment section here, it seems like the opposite; the majority of the people here need a reminder that companies aren't actually genies that can only tell falsehoods, where you can only understand what they are saying by correctly guessing the conspiracy underneath.

Taking nothing at face value gives you just as much of a distorted view of reality as taking everything at face value.

Re: Be skeptical of OpenAI's rogue hacker agent story

#307
post #17
post #14

Earlier quoted context omitted.

Didn’t they explicitly remove alignment guardrails for this test? From the press release: > These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

Interestingly if you look at the exploitgym repo (https://github.com/sunblaze-ucb/exploitgym/blob/e5ea7c233a4d...) the intended run mechanism is orchestrated by some python scripts which run agents against various prompts. The prompts themselves wouldn’t mention anything about exploitgym and there should be nothing steering the agent towards trying to find the answers out-of-band. So I don’t see how they could even run into this problem unless all they did was tell an agent to “check out and run exploitgym”

Which, after hearing some personal anecdotes of how people work there, seems plausible

Re: Be skeptical of OpenAI's rogue hacker agent story

#308

Earlier quoted context omitted.

> positive for OpenAI To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful. Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1. I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're t…

I'm disappointed that this is the level of discourse happening here, when the default assumption is such conspiratorial thinking. I expect that from tiktok and low-information social media, not here. When your default explanation for everything is "companies are lying about everything", you end up just as incorrect as believing they're always telling the truth. It's not outlandish that models have these capabilities,…

I don’t believe I argued that an LLM couldn’t find and exploit a vulnerability and even break out of some layer of technical controls. That seems realistic and has been demonstrated before and I mentioned that LLMs are used in offensive security work.

Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either faked or carefully not avoided seems more likely than the others (based on the current information we have).

Could you clarify which point or assumption you are objecting to?

Re: Be skeptical of OpenAI's rogue hacker agent story

#309

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

Regarding your three explanations, I’ve wondered to myself under a circumstance of options 1 or 2 why Hugging Face decided not to file a police report and request to press charges?

OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this?

With this logic I think explanation #3 becomes incredibly likely.

Re: Be skeptical of OpenAI's rogue hacker agent story

#310

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

It seems like the widely-covered news stories that support OpenAI’s narratives originate from things that happened inside OpenAI.

It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself.

Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.

Post reply on HN