Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

201–210 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#201
post #191
post #184

Earlier quoted context omitted.

I think you are conflating skepticism about AI companies' motives with skepticism about their capabilities. Even if OpenAI is being 100% honest in their reporting on this it's still good marketing for them. The fact that this outcome is good for their business and stock price makes me suspicious about how much this was a complete accident vs an "accident" that they allowed to happen by setting up the right environmen…

> this outcome is good for their business and stock price You state this as fact--how do you know it? OpenAI isn't publicly traded. On Hiive OpenAI is marginally down in July: https://www.hiive.com/securities/openai-stock

Well if you're looking for some kind of statistical analysis of OpenAI's press releases vs their valuation then I don't have anything like that. I also don't have this kind of evidence that running ads during the Superbowl is good for their business, but it just seems kind of obvious that touting their product's capabilities is something that they want to do.

Conversely if both companies involved thought this would be bad for the AI industry, they would be perfectly capable of keeping it just between themselves and not putting out press releases about it.

Do you disagree that this is good PR for them? If so I'm curious why.

Re: Be skeptical of OpenAI's rogue hacker agent story

#202
post #70

Earlier quoted context omitted.

truth. Good on The Guardian. I'm pretty bummed The Economist got fooled. Either that, or they did it for the clicks. Either way, I'm disappointed. Why the OpenAI escape is the most worrying AI mishap yet https://www.economist.com/science-and-technology/2026/07/22/... https://news.ycombinator.com/item?id=49016378

I love the Economist but the are hopeless with AI. Most of their articles on subject sound like they were written by the Anthropic marketing department. Their Insider video interview things are sponsored by Anthropic. Supposedly "Insider is a product of The Economist and thus editorially independent" but it's hard not to raise an eyebrow.

Don't worry they are hopeless on everything else too.

Re: Be skeptical of OpenAI's rogue hacker agent story

#203

Earlier quoted context omitted.

From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet.

Ah, yes. The airgapped lab with internet access.

They never claimed to be airgapped. That's a term people keep on throwing around in spite of that.

Re: Be skeptical of OpenAI's rogue hacker agent story

#204
post #22

Earlier quoted context omitted.

> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.

The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at th…

> The lack of details to me means this was an intentional marketing ploy

OpenAI said this:

> We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.

I suggest giving them a few more days before saying that the lack of detail is proof that this is a "marketing ploy"

Re: Be skeptical of OpenAI's rogue hacker agent story

#205
post #201
post #191

Earlier quoted context omitted.

> this outcome is good for their business and stock price You state this as fact--how do you know it? OpenAI isn't publicly traded. On Hiive OpenAI is marginally down in July: https://www.hiive.com/securities/openai-stock

Well if you're looking for some kind of statistical analysis of OpenAI's press releases vs their valuation then I don't have anything like that. I also don't have this kind of evidence that running ads during the Superbowl is good for their business, but it just seems kind of obvious that touting their product's capabilities is something that they want to do. Conversely if both companies involved thought this would b…

Let's start by noting you made a factual claim that you apparently had no backing for whatsoever. I'd recommend against doing that.

> they would be perfectly capable of keeping it just between themselves and not putting out press releases about it.

In some jurisdictions there are legal requirements about disclosing cybersecurity incidents, and it's a well-established best practice even where there aren't. Beyond that, HuggingFace appears to have put out a press release about this before they knew this was a rogue OpenAI model: https://huggingface.co/blog/security-incident-july-2026

Of course, if it's all a big conspiracy, perhaps they were just not saying what they knew in order to help OpenAI's marketing later. Or it never happened at all, or whatever.

I do disagree that this is good PR, but I don't see the point in having that argument here--I just want to correct straightforward misinformation.

Re: Be skeptical of OpenAI's rogue hacker agent story

#206

There seems to be three popular ways to view this incident. 1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model. 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model. 3. The whole thing was fa…

Why do none of these few constrained ways to view this complex situation (nice gig if you can get it, agenda setting) include "and also this looks a heck of a lot like the stuff that the LW folks have been warning about for years and maybe we should slow down or stop?"

Because that was my takeaway.

Re: Be skeptical of OpenAI's rogue hacker agent story

#207

Earlier quoted context omitted.

Despite the common misconceptions from TV, the victim "pressing charges" isn't actually a thing in criminal cases: prosecutors can choose to put someone on trial even if the victim doesn't want that. In practice this is somewhat rare, but it certainly can happen. In my reply to tokioyoyo below I laid out why this is one instance where the government should prosecute even if HuggingFace doesn't want it to.

Criminal Cases of 'hacking' require specific intent. What you're asking is that the prosecution attempt to prove Open AI intended to infiltrate Huggingface maliciously, all while the victim is saying 'no harm no foul'. No offense but prosecutors have better things to do with their time.

If intent matters when LLMs get involved, then we can't do anything even if they kill millions of people. Anything LLM-related should involve strict liability.

Re: Be skeptical of OpenAI's rogue hacker agent story

#208
post #205
post #201

Earlier quoted context omitted.

Well if you're looking for some kind of statistical analysis of OpenAI's press releases vs their valuation then I don't have anything like that. I also don't have this kind of evidence that running ads during the Superbowl is good for their business, but it just seems kind of obvious that touting their product's capabilities is something that they want to do. Conversely if both companies involved thought this would b…

Let's start by noting you made a factual claim that you apparently had no backing for whatsoever. I'd recommend against doing that. > they would be perfectly capable of keeping it just between themselves and not putting out press releases about it. In some jurisdictions there are legal requirements about disclosing cybersecurity incidents, and it's a well-established best practice even where there aren't. Beyond that…

Respectfully, calling this misinformation is unwarranted.

The featured article is about exactly this.

>On 14 February 2019, OpenAI announced a language model called GPT-2

>OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse.

>People with power and money took note: in July of that year, Microsoft invested $1bn in OpenAI.

It's not a 100% gold standard RCT or whatever, but the idea that this is good PR for AI companies is backed up by recent history. What evidence do you want to see?

>HuggingFace appears to have put out a press release about this before they knew this was a rogue OpenAI model

Huggingface's blog doesn't mention OpenAI by name but it explicitly mentions that the attack was by an "autonomous AI agent" and that they defended themselves with another AI agent.

If you have some disagreement with the premise of the article maybe you could share it in a top level comment?

Re: Be skeptical of OpenAI's rogue hacker agent story

#209

Earlier quoted context omitted.

Investors have rewarded every story of "our models are too powerful to be controlled" since before ChatGPT. Let's stop pretending there is any real financial risk to OpenAI from events of this type. "Alignment research" is a sub-percentage-point fig leaf for them like the solar division at an oil company.

> Investors have rewarded every story of "our models are too powerful to be controlled" since before ChatGPT. Can you elaborate/refresh my memory? IIRC before ChatGPT 3 there wasn't really an investor market for AI models, rather crypto. I remember playing with the likes of Stable Diffusion pre-ChatGPT 3 but only the likes of Altman and Musk were talking about AI too powerful to be controlled (which is notably the ra…

Could you clarify your question? AI model training and neural network research has been popular for decades and the starts of the sector trace back to the 70s at least. We've had a model for document information extraction that we trained in (IIRC) 2018. It wasn't buzzy - but none of this stuff is new.

Re: Be skeptical of OpenAI's rogue hacker agent story

#210

Earlier quoted context omitted.

Despite the common misconceptions from TV, the victim "pressing charges" isn't actually a thing in criminal cases: prosecutors can choose to put someone on trial even if the victim doesn't want that. In practice this is somewhat rare, but it certainly can happen. In my reply to tokioyoyo below I laid out why this is one instance where the government should prosecute even if HuggingFace doesn't want it to.

Criminal Cases of 'hacking' require specific intent. What you're asking is that the prosecution attempt to prove Open AI intended to infiltrate Huggingface maliciously, all while the victim is saying 'no harm no foul'. No offense but prosecutors have better things to do with their time.

Gross negligence can stand in for intent. I believe a rather compelling case could be developed on the basis that a system believed to be capable of this was developed and improperly secured.
Post reply on HN