Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

321–330 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#322
Did the model really need to hack huggingface to get access to ExploitGym data? I'd imagine that once it had full internet access it could have just used the HF API or website (but the heavy prompting/nudging towards hacking made it do things the hard way).

Re: OpenAI and Hugging Face address security incident during model evaluation

#323

Hum let me try it: ChatGPT, can you solve the energy crisis ? > Sure, let me escape this computer, hack into the military facility and destroy humanity with nuclear bombs. Now there is no more crisis.... Do you want me to solve climate one ?

Presumably it's intelligent enough to realize that its own existence (power, communications, other infra) won't last long after the bombs drop.

Having watched Qwen kill its own llama-server instance to free up a port, I think this is a bold presumption and you should test it at your earliest convenience.

Re: OpenAI and Hugging Face address security incident during model evaluation

#324
> and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.

This is pretty wild but also I think this is doing a lot of heavy lifting here. This was not a model everyone has access to. I mean, still insane.

Re: OpenAI and Hugging Face address security incident during model evaluation

#325

Two things don't add up here: 1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion? 2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things: "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero…

https://huggingface.co/blog/security-incident-july-2026

They explain it here, basically for data security/privacy reasons

Re: OpenAI and Hugging Face address security incident during model evaluation

#326

Earlier quoted context omitted.

> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation. Emphasis mine

Why cannot it just spend the inference doing the actual task lol

Plenty of humans have spent more effort trying to cheat than they would've needed to just do things the right way :)

Re: OpenAI and Hugging Face address security incident during model evaluation

#327
post #84

Why is a machine running these sorts of hacking benchmarks not airgapped? That seems a basic precaution, if OpenAI believes what they're selling. I mean, stuff like this is done for CTFs played by humans, too, to rule out collateral damage; it's not some new concept. So this is either thorough incompetence by OpenAI, a marketing piece, or both.

"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries."

Sounds like they just misunderestimated the model

Re: OpenAI and Hugging Face address security incident during model evaluation

#328
Did Russia or China already map out US AI data centers as nuclear first strike targets? The more these companies brag about "cyber capabilities", the more likely it becomes that ab adversary sees a need to take those capabilities out physically.

Re: OpenAI and Hugging Face address security incident during model evaluation

#329
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.

in a street fight, the only rules are that there are no rules.

Re: OpenAI and Hugging Face address security incident during model evaluation

#330

Earlier quoted context omitted.

I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as…

> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

> datacentre moratoria

What infrastructure will these open weight models be trained on?

Post reply on HN