Live data from Hacker News

Be skeptical of OpenAI's rogue hacker agent story

theguardian.com

131–140 of 321 posts

Re: Be skeptical of OpenAI's rogue hacker agent story

#131

Earlier quoted context omitted.

You don't really need anyone to be lying here. It is likely that the broad strokes of the narrative are true and that no collusion or conspiracy took place here. The issue is that a lot of important details in that narrative are missing, and the devil is really in the details here. I suspect that those details would make the result seem less exciting and that this event would move the needle far less for them if they…

Where do you see the claim that "long-horizon goals in real world settings are now effectively settled"? The argument you put in their mouth would be a bad one, but I don't see anyone making it.

https://openai.com/index/hugging-face-model-evaluation-secur...

> UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.

I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess.

We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting.

Re: Be skeptical of OpenAI's rogue hacker agent story

#132
post #120

Earlier quoted context omitted.

Nobody is saying the models are incapable of what was claimed.

Firstly: yes, very many people are saying this. Secondly: to the people who aren't saying it....then why are you bringing up marketing at all? If the model is capable of it, then the motivation for why OpenAI is talking about it/reporting on it is completely beside the point. Either the capability matters or it doesn't. If the capability doesn't matter, or doesn't matter in the way that some particular person is clai…

I think you’re missing the fact that no one is saying not to be concerned, quite the contrary. OpenAI is using the threat of how powerful its models are to bolster support for regulation in which it’s one of the only players that’s allowed to use the capability. Conveniently, that would also be a moat that makes them more valuable to investors.

The marketing of their models as super dangerous has a direct link to the regulatory moat they’re pursuing.

Re: Be skeptical of OpenAI's rogue hacker agent story

#133
post #42

Earlier quoted context omitted.

Being a poorly equipped victim still isn’t a crime thankfully. It may or may not be a crime and typically the damaged party is pressing the charges. One would argue there is no actual damage here.

Despite the common misconceptions from TV, the victim "pressing charges" isn't actually a thing in criminal cases: prosecutors can choose to put someone on trial even if the victim doesn't want that. In practice this is somewhat rare, but it certainly can happen. In my reply to tokioyoyo below I laid out why this is one instance where the government should prosecute even if HuggingFace doesn't want it to.

But what do you honestly expect would happen? It’s “an accident” with no actual damages.

Re: Be skeptical of OpenAI's rogue hacker agent story

#134

Earlier quoted context omitted.

The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at th…

> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind. "our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing. This is an actual critical problem , not a s…

It's a critical problem like when a drug dealers supply kills someone and they get a bump in business because they're selling "the real deal"

Re: Be skeptical of OpenAI's rogue hacker agent story

#135
post #22

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> AI managed to escape using standard and well documented script kiddie methods. I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.

I don't understand why "OpenAI says" should be considered any more meaningful than "someone on HN says" when they provide equal amounts of evidence. Sure, OpenAI would plausibly have more pertinent info, but given that they actively are choosing not to share it and have way more incentive to lie than a random HN stranger, the case they're making literally couldn't be any weaker.

Re: Be skeptical of OpenAI's rogue hacker agent story

#137

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

> AI managed to escape using standard and well documented script kiddie methods

> AI broke in using standard script kiddie methods.

Go ahead and show us how easy it is to break into HuggingFace (and OpenAI) networks.

Re: Be skeptical of OpenAI's rogue hacker agent story

#139

Earlier quoted context omitted.

Where do you see the claim that "long-horizon goals in real world settings are now effectively settled"? The argument you put in their mouth would be a bad one, but I don't see anyone making it.

https://openai.com/index/hugging-face-model-evaluation-secur... > UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings. I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a…

[deleted]

Re: Be skeptical of OpenAI's rogue hacker agent story

#140

Finally mainstream news understands. The unfiltered version: 1) The AI failed to solve ExploitGym problems. 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods. 3) Huggingface has no security and the AI broke in using standard script kiddie methods. OpenAI and Huggingface covered it up and used it for public relations. That is, if not al…

Are we finally now in 2026 coming around to the idea that sometimes entities may find themselves incentivized to conspire with each other? Is theorizing about such no longer off-limits due to a thought-terminating cliche?
Post reply on HN