Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

831–840 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#832

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Remember when the pre-GPT3 days when the main argument against AI alignment concerns was that "we simply won't let it out of the box"? So quaint in hindsight.

Re: OpenAI and Hugging Face address security incident during model evaluation

#833
post #590

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. Researcher: hack me Model: understood Researcher: oh my god

Researcher: hack me

Model: I committed a crime

Researcher: oh my god

Re: OpenAI and Hugging Face address security incident during model evaluation

#834

From https://huggingface.co/blog/security-incident-july-2026 , this is frickin' hilarious: > When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident respon…

so an on-premise and open-weight model was more useful than a commercial frontier model?

Re: OpenAI and Hugging Face address security incident during model evaluation

#835

Earlier quoted context omitted.

What disturbs me is that there likely won’t be a big enough reaction to this policy wise. There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong han…

I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as…

[dead]

Re: OpenAI and Hugging Face address security incident during model evaluation

#836

I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this: Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart m…

Anthropic in general seems to have better security...but they also had reported an internal AI gained access to outside email services to contact an Anthropic developer

Re: OpenAI and Hugging Face address security incident during model evaluation

#838

Earlier quoted context omitted.

You know that people can plan the whole theatre?

I'm not quite sure what you mean by that, but it sounds like you're suggesting the lying about it.

The events happened, it doesn't mean those were accidental.

Re: OpenAI and Hugging Face address security incident during model evaluation

#839
post #823

Earlier quoted context omitted.

We are not going to know we have crossed the line until it's been crossed and we can look back and say "Oops. We should have done something back then."

People have been sounding that alarm for well over a year now, just not anyone who has the power to do anything about it.

Lesswrong / “doomers” have been sounding the alarm on this for over a decade. Sam Altman himself was familiar with the milieu and agreed “Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity.”

It’s baffling these concerns have been drowned out for so long until now that we’re finally at a point they are unequivocally undeniable, we get people saying the alarm has just been sounded. Ya because “doomers” were repeatedly dismissed as their predictions became true year after year

Re: OpenAI and Hugging Face address security incident during model evaluation

#840
post #162

As grounded as this article comes across I can’t help but find this whole situation reckless and worrying. There is essentially nothing us private citizens can do while these companies develop super machine capabilities that if they were to slip into the wrong hands could cause massive real world problems. They’re moving fast and breaking things and the only defense we have is paying them money in the hopes that the…

It's computer Gain of Function research.

Gain of function research is not anywhere near as dangerous as the public believes. In the US, it was very convenient to blame it for the pandemic, even though SARS-CoV-2 is of natural origin and almost certainly spilled over at the Huanan wet market in Wuhan. The threat of viruses comes almost exclusively from nature, which is constantly cooking up new viruses all by itself and exposing people all over the world to them. A few people doing tightly controlled research under high-biocontainment are a drop in the ocean. But their research is the main way we can prepare to deal with future pandemics, not to mention understanding the usual viruses that already afflict humanity.
Post reply on HN