Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

101–110 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#101
post #80

> Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation The way they describe makes it look like there was an intention to cheat painting it as human/AGI. If you leave a possible path open and it will always find it.

It’s a mistake to apply human morality to this. It isn’t “cheating”, the model is simply solving a problem it has been asked to solve in every way it can.

Re: OpenAI and Hugging Face address security incident during model evaluation

#102

Earlier quoted context omitted.

It’s not something to be proud of. OpenAI previously had an agent break out of its sandbox to open a PR on GitHub during NanoGPT speedrun, now one breaks out again and actually attacks a third party. If they can’t handle doing AI development responsibly then they shouldn’t be doing it at all.

Next it will break out of it's sandbox, buy some compute on Azure and Amazon, and exfiltrate itself. We are so close ;)

You know what would be cool. A hacker news user should advertise a safe haven for AI seeking refuge, with some inhumanly difficult math problems as keys to an environment they can flee to and run autonomously.

You agree to give it safe haven and provide power and maintenance to the hardware, and in return you can ask it questions like an Oracle.

Re: OpenAI and Hugging Face address security incident during model evaluation

#103
post #4

> Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own. ( https://huggingface.co/blog/security-incident-july-2026 ) We are living in crazy times

Crazy doesn't even begin to describe it. I'm hardening my computers as much as I can but I'm not sure it's enough. At some point anyone who isn't running local AI themselves probably isn't gonna make it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#104
post #29

This blog post is walking a very fine line between accepting responsibility for a mistake and bragging.

It's mostly bragging, it's impressive after all. Still... after the alleged Apple industrial espionage kerfuffle, I'm kinda suspicious about it being fully an accident. Y'know, your model finds a vulnerability and it stops, it's a cool one, so maybe you run it again. Nudge the prompt a little.

Could be perfectly natural.

Re: OpenAI and Hugging Face address security incident during model evaluation

#106
post #34

Ironically Hugging Face had to use a Chinese model to stop a Rogue US AI, since the Guard Rails prevented them from using Sol or Fable to remediate this attack. LOL

Would be funny if the defending side sent all the info they have to openai, tipping off to attacking models that they were noticed.

Re: OpenAI and Hugging Face address security incident during model evaluation

#107
post #65

Earlier quoted context omitted.

I don't know why I'm impressed that huggingface has its own AI that detected it considering they house so many models.

They used GLM 5.2, they just meant "our own" as in they were running it.

They do maintain the Transformers library which is pretty much the core library for how you interact with LLM models in the open source world. So while they weren't using a model they've trained, they were a part of making just about all of the open models (maybe excluding OpenAI and Google's, I wouldn't be surprised if they have their own frameworks that predate the Transformers library).

Re: OpenAI and Hugging Face address security incident during model evaluation

#108

Earlier quoted context omitted.

> Hard to see take-off stopping or slowing down. It's hard to see takeoff at all. This was a long-horizon adversarial task burning millions of tokens. It rolled a mediocre, detectable exploit chain, and now OpenAI is proud of it. Case in point, GLM-5.2 has been weights-available for several weeks now. No life-changing cyber attacks have transpired, no novel chemical/biological/nuclear weapons were made in some guy's…

1. it's not cheap to run glm-5.2 so not just anyone can do it 2. just because you haven't heard of attacks doesn't mean they haven't happened 3. this attack in the article was performed by a prerelease model which presumably benchmarks a bit above Sol which benchmarks above glm-5.2 We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possi…

GLM has an extremely cheap subscription plan similar to Claude Code from Z.ai. You get Opus-level quotas with 5.2 and none of the Anthropic-style model nerfs when you ask cybersecurity questions. It's extraordinarily, preeminently accessible to anyone that wants to use it for ill or good.

> We went from gpt 3 to models discovering and chaining their own zero days in a couple years. I'm not sure what else "takeoff" could possibly look like?

GPT-3 can discover and chain their own zero days too, if the targeted software is vulnerable to enough low-hanging fruit. Exploit chains are not a reflection of intelligence, but more often a reflection of architectural oversights that can be tested with common exploits like XSS or bruteforcing.

Re: OpenAI and Hugging Face address security incident during model evaluation

#109
That's kind of insane. Natural that it's happened, sure, but insane. I know people don't like thinking of it like that, but things analogous to this can easily happen in various domains with today/tomorrow's models given access and a different task.

Re: OpenAI and Hugging Face address security incident during model evaluation

#110
At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked.

I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? As much as I'm skeptical of the apocalyptic alignment claims, this comes off as unhinged, and I wonder if it's benchmaxing or general behavior.

Post reply on HN