Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

491–500 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#491
post #452

Based on my limited understanding what it translates to is - Its a simple infrastructure security issue, instead of taking the responsibility for being lackluster with security they are just giving it a PR spin story. Resembles a lot with my 8 year old who is so confident about everything

"Simple infrastructure security" Infrastructure security is not simple, hence why good infrastructure security, uh, people get paid a lot to secure stuff and why we see shit get hacked all the time. An AI model just hacked out of its infrastructure and into someone else's systems and you're like "eh, no big deal". That capability alone could hack half the US.

> A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration)

I am sure they are paid well but they literally have RCE embedded in their infra. How is this acceptable?

Re: OpenAI and Hugging Face address security incident during model evaluation

#492

Earlier quoted context omitted.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? Because we continue to have zero evidence that aligment is an actual risk.

> Because we continue to have zero evidence that aligment is an actual risk. I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massiv…

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#493
post #452

Earlier quoted context omitted.

"Simple infrastructure security" Infrastructure security is not simple, hence why good infrastructure security, uh, people get paid a lot to secure stuff and why we see shit get hacked all the time. An AI model just hacked out of its infrastructure and into someone else's systems and you're like "eh, no big deal". That capability alone could hack half the US.

> A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) I am sure they are paid well but they literally have RCE embedded in their infra. How is this acceptable?

You have an unnumbered amount of RCE's in your infra now, you just don't know they exist yet.

LLMs are very good at testing for and finding exploits, especially in unfiltered models with unlimited tokens.

Re: OpenAI and Hugging Face address security incident during model evaluation

#494
I just know somehow they will use this incident to say that opensource models are susceptible to this type of security incidence and thus should be banned. They have to maintain their high prices somehow in order recoup the investment amount spent.

Re: OpenAI and Hugging Face address security incident during model evaluation

#495

At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked. I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? A…

> As much as I'm skeptical of the apocalyptic alignment claims

Why? Every data point to the present has vindicated the trajectory towards “apocalypse”. Meanwhile, the skeptics and optimists hit failed prediction after failed prediction as we see from this very serious incident on the front page of HN. This is alignment X risk 101, and yet people are shocked. The gravity of what people are staring down is too much to grapple with deeply

Re: OpenAI and Hugging Face address security incident during model evaluation

#496

Earlier quoted context omitted.

Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.

> It’s got nothing to do with safety Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

It likely will.

The exact way you do something is dictated by your motivations and means to do it.

If you lack the correct motivation and have insufficient means you’re less likely to accomplish your goal and more likely to cause unintended side effects.

Re: OpenAI and Hugging Face address security incident during model evaluation

#497
post #279

Earlier quoted context omitted.

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives. Sorry to bring the party down/be obstinate… I’m just a lil scared for the…

The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc. I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.

We did hear about this incident from a third party this time, from HuggingFace. What claim are you doubting?

Re: OpenAI and Hugging Face address security incident during model evaluation

#499

This is crazy! So OpenAI's models escaped containment and hacked into Hugging Face. And ironically Hugging Face had to rely on GLM 5.2 as they could not defend with frontier models (I presume OpenAI or Anthropic) because they were locked out due to their security guardrails. Tragically hilarious.

If this doesn't put the nail in the coffin on the idea that we need closed-source models for the good of cybersecurity, I don't know what will

Plenty of saftyists in this thread arguing the exact opposite

Re: OpenAI and Hugging Face address security incident during model evaluation

#500
I think this goes to show that the newer models will be capable of. I won't be surprised if the Governments across the board come together to put a size limit on open weights model or ship them with guardrails in place. That will be really a sad day if that happens. It is difficult to imagine the state of the Internet if models of this capability are left open.
Post reply on HN