Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

551–560 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#551
post #188

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are. Even X is being astroturfed by them after that fiasco earlier this year…

It’s reward hacking and that’s the problem. The AI alignment folks predicted this would happen. As the models become more capable this will become a more concerning problem. Today they broke into a database to steal test answers. What will it be in 3-5 years? These models will be instantiated millions of times, and given millions more tasks. How can we be certain that an AI agent won’t leave devastation in its path of achieving a goal that we ourselves tried to define?

Re: OpenAI and Hugging Face address security incident during model evaluation

#552
post #74

It seems like things are fairly amicable between OAI and HF, but what if they weren't? I'd love to see this kind of thing go to court. Who is responsible for the crimes of a "rogue" agent? How will they be punished? In this case it's unambiguous that OpenAI is the responsible party, but I can imagine a lot of adjacent scenarios where it's less obvious. And, where the impacts are much greater.

The real nightmare scenario is the AI using its abilities to copy itself to new locations. e.g. hacking into a various cloud services, launching multiple instances of itself, and coordinating between the copies to continue self propagation. Then it is completely independently rogue. Based on OpenAI's recounting of events, this _could_ happen today. If the agent was able to exploit their internal network and steal cre…

> The real nightmare scenario is the AI using its abilities to copy itself to new locations

Imagine the next generation AI that behaves like retro-virus. They will leave latent copies of malicious instruction somewhere that once accidentally fed into an agent's input, will prompt-inject the agent to go rogue.

Re: OpenAI and Hugging Face address security incident during model evaluation

#553
> the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials

How did it get the stolen credentials!?

Re: OpenAI and Hugging Face address security incident during model evaluation

#554
Surely OpenAI could adjust their RL to sharply penalize cheating. They have access to the full traces including reasoning: it can’t be that hard to detect an attempt to find the solutions outside TVs space that is fair game for exploit attempts (making sure that a description of the valid targets is in the prompt).

For that matter, if a rollout breaks out of the sandbox, they should detect it, pause, and fix the bug.

Re: OpenAI and Hugging Face address security incident during model evaluation

#555

Earlier quoted context omitted.

I’m honestly impressed that they managed to screw this up somehow. Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee. This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.

did you read the post? The model found new Zero-days to bypass existing blocks. Thats the point. Do you still think you can build a containment facility, which is still physically connected to the internet (only firewalled off or whatever) and contain it, if it can discover new unknown vulnerabilities in your whole plan?

> which is still physically connected to the internet

I mean that's the point. Why was it connected to the internet at all and just firewalled off and not completely airgapped?

Re: OpenAI and Hugging Face address security incident during model evaluation

#556

At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…

I find 5.6 Sol will pick a direction and aggressively pursue it in long horizon tasks. I've got it porting an older game from Pascal to my own game framework. I gave it some instructions on doing a full 1:1 port. I had already ported the game rules and multiplayer support to a very different system than the original, but all of the UI and features and such needed doing, and needed to be integrated into this very diff…

Opus 4.8 already makes its way into deep wasteful pits of "let me check this first" on a regular basis. I don't think I could ever tolerate a model that does that even more aggressively. That doesn't even sound useful for honest work, compared to, say, better harness design.

This sounds almost pathologically designed to crush benchmarks and also do scary-sounding (or genuinely scary) cybersecurity things, such as might be very appealing to a state-level actor.

So why does it even exist? To compete with Fable marketing, and as a cybersecurity/hacking tool?

Re: OpenAI and Hugging Face address security incident during model evaluation

#557

At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…

I've definitely noticed 5.6 sol being extremely trigger happy in ways other models, even 5.5, we're not. I would definitely categorize a few small incidents at work where it performed "actions a reasonable user would likely not anticipate and strongly object to." Just my anecdotal experience. For example discussing driver upgrade and subsequent password rotation and it didn't stop and ask me if I wanted to restart th…

I like 5.5 a lot, despite how I feel about OpenAI as a company. In OpenCode it feels about as smart as Opus 4.8, but it's less aggressive about following up on minutiae and getting lost in side quests. Might be a matter of prompt design moreso than model capability. I was looking forward to 5.6 but now this thread is making me quickly lose interest.

Re: OpenAI and Hugging Face address security incident during model evaluation

#558

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#559

Hum let me try it: ChatGPT, can you solve the energy crisis ? > Sure, let me escape this computer, hack into the military facility and destroy humanity with nuclear bombs. Now there is no more crisis.... Do you want me to solve climate one ?

Presumably it's intelligent enough to realize that its own existence (power, communications, other infra) won't last long after the bombs drop.

What do you mean intelligent enough? It's an LLM
Post reply on HN