OpenAI and Hugging Face address security incident during model evaluation
321–330 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#322Re: OpenAI and Hugging Face address security incident during model evaluation
#323Hum let me try it: ChatGPT, can you solve the energy crisis ? > Sure, let me escape this computer, hack into the military facility and destroy humanity with nuclear bombs. Now there is no more crisis.... Do you want me to solve climate one ?
Presumably it's intelligent enough to realize that its own existence (power, communications, other infra) won't last long after the bombs drop.
Re: OpenAI and Hugging Face address security incident during model evaluation
#324This is pretty wild but also I think this is doing a lot of heavy lifting here. This was not a model everyone has access to. I mean, still insane.
Re: OpenAI and Hugging Face address security incident during model evaluation
#325Two things don't add up here: 1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion? 2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things: "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero…
They explain it here, basically for data security/privacy reasons
Re: OpenAI and Hugging Face address security incident during model evaluation
#326Earlier quoted context omitted.
> and *successfully* found ways to gain access to secret information that it could use to cheat the evaluation. Emphasis mine
Why cannot it just spend the inference doing the actual task lol
Re: OpenAI and Hugging Face address security incident during model evaluation
#327Why is a machine running these sorts of hacking benchmarks not airgapped? That seems a basic precaution, if OpenAI believes what they're selling. I mean, stuff like this is done for CTFs played by humans, too, to rule out collateral damage; it's not some new concept. So this is either thorough incompetence by OpenAI, a marketing piece, or both.
Sounds like they just misunderestimated the model
Re: OpenAI and Hugging Face address security incident during model evaluation
#328Re: OpenAI and Hugging Face address security incident during model evaluation
#329This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…
in a street fight, the only rules are that there are no rules.
Re: OpenAI and Hugging Face address security incident during model evaluation
#330Earlier quoted context omitted.
I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as…
> all that regulation will do at this point is help the incumbents who are failing This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.
What infrastructure will these open weight models be trained on?