> Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation The way they describe makes it look like there was an intention to cheat painting it as human/AGI. If you leave a possible path open and it will always find it.
OpenAI and Hugging Face address security incident during model evaluation
101–110 of 1001 posts
Re: OpenAI and Hugging Face address security incident during model evaluation
#102Earlier quoted context omitted.
It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.
Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself. We are so close ;)
You agree to give it safe haven and provide power and maintenance to the hardware, and in return you can ask it questions like an Oracle.
Re: OpenAI and Hugging Face address security incident during model evaluation
#103> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times
Re: OpenAI and Hugging Face address security incident during model evaluation
#104This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.
Could be perfectly natural.
Re: OpenAI and Hugging Face address security incident during model evaluation
#105Re: OpenAI and Hugging Face address security incident during model evaluation
#106Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL
Re: OpenAI and Hugging Face address security incident during model evaluation
#107Earlier quoted context omitted.
I don't know why I'm impressed that huggingface has its own AI that detected it considering they house so many models.
They used GLM 5.2, they just meant "our own" as in they were running it.
Re: OpenAI and Hugging Face address security incident during model evaluation
#108Earlier quoted context omitted.
> Hard to see take-off stopping or slowing down. It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it. Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's…
1. it's not cheap to run glm-5.2 so not just anyone can do it 2. just because you haven't heard of attacks doesn't mean they haven't happened 3. this attack in the article was performed by a prerelease model which presumably benchmarks a bit above Sol which benchmarks above glm-5.2 We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possi…
> We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possibly look like?
GPT-3 can discover and chain their own zero days too, if the targeted software is vulnerable to enough low-hanging fruit. Exploit chains are not a reflection of intelligence, but more often a reflection of architectural oversights that can be tested with common exploits like XSS or bruteforcing.
Re: OpenAI and Hugging Face address security incident during model evaluation
#109Re: OpenAI and Hugging Face address security incident during model evaluation
#110I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? As much as I'm skeptical of the apocalyptic alignment claims, this comes off as unhinged, and I wonder if it's benchmaxing or general behavior.