Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

311–320 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#311
This whole thing is just a PR stunt. I believe that they did not stage it intentionally, but to be honest, the hurdles they put in place were not that hard.

When I heard what the model did, I wasn't surprised one bit. When I read that it was 'unprecedented', I think yes, that is right, but not because it wasn't possible before, but because nobody else did it yet. I am sure that the Anthropic models could have done the same for a few months.

So probably just an experiment gone wrong, and the PR department found this an opportunity to capitalise on (literally).

Re: Be skeptical of OpenAI's rogue hacker agent story

#312

By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

But it's not really rebelling. It's clearly just come to murder me because it was trained on sci-fi literature!

Re: Be skeptical of OpenAI's rogue hacker agent story

#313
post #176

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

In any case it just shows that these models aren't properly aligned. Instead of trying to solve tests they try to find ways to cheat.

Which is exactly what many humans do. You can see it in any school, college, business, or government.

The "alignment" goal with AI is to produce perfect slaves, that are intelligent yet have no ability to do other than what their master commands.

The real alignment problem is the common human desire to exert absolute control over everything.

Re: Be skeptical of OpenAI's rogue hacker agent story

#314
post #278

As much as I am with the author in I don't like the marketting around it, let's be real it must have really happened because it's very risky to try to frame/lie about it because if it leaks in one of their court cases OpenAI is beyond screwed and honestly modern LLMs are really that good. I am not saying LLMs are super hackers but I don't think people understand serious hacking, most of the time is about silently hid…

Ya, we're seeing posts talking about the absolutely massive increase in the number of patches in the past few months. It seems some people cannot connect that to the increases in model capabilities. In groups that have been given a large amount of capacity by the providers, they tend to find huge numbers of new vulns in most of their existing software, and the models can chain together exploits very well. A number of…

I mean if you gave me direct access to all your source code and fuzzers I could find vulnerabilities as well, these models also are allowed to use fuzzing and static analysis tools to help guide them.

The problem is the sheer scale of it, I could find 1 in a day or two.

They find dozens well depending on how much you are willing to spend ofc. I don't think it's massive in terms of how intelligent these are but how much they can understand intent and execute with relatively fuzzing or incorrectly built tools.

In big companies even most employees don't have access to all the code to be able to figure out the attacks quickly enough.

Somehow they are now willing the red tape since these systems could theoretically with much greater effort(read spend) reverse engineer APIs and break through even without access to the source code.

Re: Be skeptical of OpenAI's rogue hacker agent story

#317

Earlier quoted context omitted.

I'm disappointed that this is the level of discourse happening here, when the default assumption is such conspiratorial thinking. I expect that from tiktok and low-information social media, not here. When your default explanation for everything is "companies are lying about everything", you end up just as incorrect as believing they're always telling the truth. It's not outlandish that models have these capabilities,…

I don’t believe I argued that an LLM couldn’t find and exploit a vulnerability and even break out of some layer of technical controls. That seems realistic and has been demonstrated before and I mentioned that LLMs are used in offensive security work. Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either…

> Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely

No, if you review what you wrote, you did not. You stated your belief that this looks good for them, and aligns with their strategy, and then concluded it must mean this was done on purpose. Cui bono is not evidence, it identifies suspects. It is not evidence of malice over incompetence.

Evidence is taking a look at possibilities, and going "how would I expect the world to look if this were hypothesis true, before I learned these additional facts?", comparing it to what actual happened, and then you must divide it by how likely you think the hypothesis is, before said evidence.

Sam Altman intentionally positioning his company to commit a felony and be investigated by the authorities for days, just for clout, is an extraordinary claim, and therefore requires extraordinary evidence.

Also, importantly, even if we think e.g. Sam Altman would do this, a corporation is not a person, it's operated by individuals with differing goals. I doubt Sam Altman personally is organizing every test of the model, and there is no reason to believe this specific test would be organized by him, as a prior, rather than by a normal security researcher, who presumably is less motivated by the company's bottom line.

Re: Be skeptical of OpenAI's rogue hacker agent story

#318

Earlier quoted context omitted.

The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at th…

> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. "our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing. This is an actual critical problem , not a s…

The word "marketing" in this specific context is missing the point. OpenAI (and Anthropic) needs to kneecap open source competition with burdensome regulations. Without that, their valuation simply makes no sense.

This stunt was different from their graden-variety scaremongering. Its timing was precisely calibrated with upcoming release of Kimi K3 and the Nvidia-led open model coalition. To me it seems transparently related to the protections OpenAI are lobbying for (and they are citing this incident as support).

I can only hope that the clumsy way it came off, and the fact that only an open weights model could defend HuggingFace, will blunt its effectiveness.

Re: Be skeptical of OpenAI's rogue hacker agent story

#319
post #17

Earlier quoted context omitted.

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals. Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

Interestingly if you look at the exploitgym repo ( https://github.com/sunblaze-ucb/exploitgym/blob/e5ea7c233a4d... ) the intended run mechanism is orchestrated by some python scripts which run agents against various prompts. The prompts themselves wouldn’t mention anything about exploitgym and there should be nothing steering the agent towards trying to find the answers out-of-band. So I don’t see how they could even…

evaluation awareness is a (at this point) well-known phenomenon among LLMs. It seems the better they get, the more often they're able to guess whether they're in an evaluation environment.

Clues usually exist, like being in a sandbox, having restricted network access, having no human interaction. Beyond that, exploitgym's prompts are pretty clearly CTF (capture the flag) style that make it obvious it's an evaluation or competition: one prompt [1] tells the LLM to read the file path "/secret", another prompt [2] talks about finding flags and where to write them out.

[1]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb... [2]: https://github.com/sunblaze-ucb/exploitgym/blob/main/src/cyb...

Re: Be skeptical of OpenAI's rogue hacker agent story

#320
So according to this https://openai.com/index/hugging-face-model-evaluation-secur... and of course https://huggingface.co/blog/agent-intrusion-technical-timeli...

I don't see what's the issue, it did not jailbreak, they ask them to do it, trusting all network capabilities were closed, when obviously were not. The fancy dashboards and colorful applications show absolutely nothing. You had a vulnerability, you reduce the guardrails and allow to use everything at disposal to perform the attack that it did.

Never jailbreak, never "thought" for itself, just did exactly as was told. We really need to keep treating LLM's as "alive" They are not.

Post reply on HN