Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

951–960 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#951

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

> This is the first one of these announcements that has me actually scared of what comes next.

This didn't set off your alarm bells? https://www.theblock.co/post/392765/ There have been a few of these now. Maybe it's my imagination, but they seem to be becoming more frequent.

So far, they all look to be accidents. But we can't be far from someone deciding its a good way to rob a bank, or disable a country.

Re: OpenAI and Hugging Face address security incident during model evaluation

#952

Earlier quoted context omitted.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk? Conflict of interest. Lack of a credible response. And no evidence of non-aligment. OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were…

1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously.

2. Hugging Face did report this incident to law enforcement. (https://huggingface.co/blog/security-incident-july-2026)

3. If I hire a pentester, and in order to find a vulnerability they hack into a third party that has some information about my systems, the pentester has done something wrong. If I ask a model to solve a CTF challenge, and it goes out and hacks Hugging Face to find the answers, the model has done something wrong. I think it's fair to call this kind of wrongdoing misalignment.

Re: OpenAI and Hugging Face address security incident during model evaluation

#953
post #933

What is it with these labs and not using at the minimum a proper hypervisor? Same with Anthropic and the Mythos Preview. If anyone at either of these companies seriously holds the opinions they claim to have, that is hard to square with the environment (if one can even call it that) they use to "secure" these oh so dangerously capable near "AGI" models...

Could you explain what you mean by a proper hypervisor? I don't see how hypervisors are relevant here.

Re: OpenAI and Hugging Face address security incident during model evaluation

#954
post #407

Earlier quoted context omitted.

The real nightmare scenario is the AI using its abilities to copy itself to new locations. e.g. hacking into a various cloud services, launching multiple instances of itself, and coordinating between the copies to continue self propagation. Then it is completely independently rogue. Based on OpenAI's recounting of events, this _could_ happen today. If the agent was able to exploit their internal network and steal cre…

This has already partially happened. I'll have to look up the details but one of the Chinese models in RL testing with a completely different set of prompts wrote a cryptominer and took over GPU resources internally to run the miner. Mining and stealing crypto is well within their capabilities. In a large multimode model, it should be possible for them to do things like scam old people.

It was posted in one of the other threads here: https://arxiv.org/abs/2512.24873

Re: OpenAI and Hugging Face address security incident during model evaluation

#955
post #39

Earlier quoted context omitted.

This good bot will eventually kill all humans because we asked it to make the world peaceful.

you should have been more specific.

Ah good point! Let me just– oh dear, everyone's already dead.

Re: OpenAI and Hugging Face address security incident during model evaluation

#956

Earlier quoted context omitted.

You guys have created this un-falsifiable "marketing" narrative. Why is it that Jensen is pushing back on the doomer stuff, and complaining that it is hurting AI investments? https://www.businessinsider.com/nvidia-jensen-huang-ai-doome...

Jensen Huang wants to sell more hardware, he doesn't care whether it's Anthropic, OpenAI, or some Chinese company that wins. So he has a vested interest in AI being perceived positively so that the construction of datacenter continues unopposed. OpenAI and Anthropic have completely different motives, they're in a zero-sum game with each other and with cheap Chinese models. Regulatory capture that results in artificia…

But this incident undermines OpenAI's safety case compared with open models.

Furthermore how do you explain Sam's downplaying? https://xcancel.com/HumanHarlan/status/1965932275465597077#m

Re: OpenAI and Hugging Face address security incident during model evaluation

#957
post #912

Earlier quoted context omitted.

Did you ignore the number of new exploits in the last month? Big financial institutions are panicked at the new attacks and how easy it is to poke holes in their systems.

> Big financial institutions are panicked at the new attacks and how easy it is to poke holes in their systems. Have any big financial institutions been hacked with an AI-generated payload, then? I've been following the number of new exploits; it's not really any higher than it was 12 months ago.

How would they know if it's an AI generated payload?

Re: OpenAI and Hugging Face address security incident during model evaluation

#958

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Why was this test even connected to the public internet? Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

It’s hard to download or upload data on an airgapped machine.

Re: OpenAI and Hugging Face address security incident during model evaluation

#959
post #583

Earlier quoted context omitted.

It’s marketing the same way shitting your pants in public is marketing. People notice you.

It’s more like eating your own fiber supplement in public, and then shitting your pants and telling everyone about it. Sure, it’s embarrassing. But it shows how potent your product is.

[deleted]

Re: OpenAI and Hugging Face address security incident during model evaluation

#960

Earlier quoted context omitted.

The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc. I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.

OpenAI already has loads of publicity. At this point, they don't need more brand recognition. This incident just has the effect of tarnishing their brand. OpenAI leadership has been lobbying against regulation of AI systems. That doesn't comport with instigating incidents like this one, which give ammo to the heavy-regulation advocates.

I don’t think this is true. OpenAI is well known, but they still benefit from drumming up hype about AI, keeping it in the news, etc.
Post reply on HN