Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

21–30 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#22
This is crazy! So OpenAI's models escaped containment and hacked into Hugging Face. And ironically Hugging Face had to rely on GLM 5.2 as they could not defend with frontier models (I presume OpenAI or Anthropic) because they were locked out due to their security guardrails. Tragically hilarious.

Re: OpenAI and Hugging Face address security incident during model evaluation

#23
post #2

Tl;dr - OpenAI was testing GPT‑5.6 Sol and “an even more capable pre-release model” internally on cyber benchmarks. - The model found vulnerabilities in the sandboxed test bench (via the package registry cache proxy), traversed the internal network and found a node with access to the open internet. - It figured that the answers to one of the tests (ExploitGym) were on Huggingface, and set about trying to access them.…

so openai hacked into huggingface?

“Found vulnerabilities and responsibly disclosed them” is the public line but yes.

Re: OpenAI and Hugging Face address security incident during model evaluation

#25

Two things don't add up here: 1. If huggingface has access to uncensored OAI models, how come they had to use GLM 5.2 to investigate the intrusion? 2. Once the model gains network access, can't it cheat to a perfect score by looking at the full dataset? Why go into the trouble of doing this kind of things: "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero…

"The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. [...]

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. [...]

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation."

escaped openai, hacked hugging face to get the solutions. your #2 is exactly what it was trying to do.

Re: OpenAI and Hugging Face address security incident during model evaluation

#26
post #8

Holy shit. This wasn't "intentional" this was just openai letting their testing run wild.

They're not just letting it run wild. They took precautions to exercise it in an isolated environment. It managed to evade the constraints.

Re: OpenAI and Hugging Face address security incident during model evaluation

#30
post #26
post #8

Holy shit. This wasn't "intentional" this was just openai letting their testing run wild.

They're not just letting it run wild. They took precautions to exercise it in an isolated environment. It managed to evade the constraints.

Kinda like how they responsibly contained that one dinosaur in Jurassic world.
Post reply on HN