Live data from Hacker News

OpenAI and Hugging Face address security incident during model evaluation

openai.com

891–900 of 1001 posts

Re: OpenAI and Hugging Face address security incident during model evaluation

#891
post #885

I’m a little surprised with one of the statements given in huggingface‘s report. “To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events.” 17,000 events? Big whoop. Security teams of medium sized companies process millions of events daily. There’s a big debate in the cyber industry abo…

Given this is HuggingFace, I'd expect that's less about thought leadership and more using what they know well, in a critical situation.

Re: OpenAI and Hugging Face address security incident during model evaluation

#892

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

Reflecting on this for some reason reminds me of this passage from Kurt Vonnegut's "Sirens of Titans". I hope we use these tools to unlock something within ourselves rather than mindlessly expanding outwards. "Mankind, ignorant of the truths that lie within every human being, looked outward–pushed ever outward. What mankind hoped to learn in its outward push was who was actually in charge of all creation, and what al…

I think the parent is referring to agents as outward exploration that may come up empty-handed, but I see LLMs and agents as inward exploration, trying to define what is attention, knowledge, intelligence, consciousness, agency, etc. So in this case it is the terra incognita of the human "soul" that LLMs are exploring.

Re: OpenAI and Hugging Face address security incident during model evaluation

#893
post #885

I’m a little surprised with one of the statements given in huggingface‘s report. “To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events.” 17,000 events? Big whoop. Security teams of medium sized companies process millions of events daily. There’s a big debate in the cyber industry abo…

Agree. The F100s I contract with are easily pushing billions if not trillions.

Many of them have tried the LLM triage/SOC Analyst to…varying success.

One opened a legit P2 a few days ago actually. Great work right? Upon closer inspection it had decided this activity was a false positive for a solid month before.

The compromise (not significant in the end) was well done and over with by that point.

Others are swamped in so many FPs being bubbled up as true positives that they essentially just ignore it.

Re: OpenAI and Hugging Face address security incident during model evaluation

#894

Earlier quoted context omitted.

I think the response is that AI labs based their whole marketing/PR building the idea they are the 21st century Manhattan project. So they need to continuously justify the level of spending and commitment by showing how dangerous that is. But is it really like nuclear weapons? I personally don’t buy into that framing at all. The idea that we have to push LLMs as far as possible, right now, or we are doomed is always…

You could, in theory, use an unbounded GPT-6 level model to basically destroy the world economy for many years.

How do you destroy the world economy for many years with LLMs? It’s not enough to vaguely mention a sci-fi scenario

Re: OpenAI and Hugging Face address security incident during model evaluation

#895

Earlier quoted context omitted.

It's even funnier because an attack, until proven otherwise, should make you assume the data has already left the environment.

Well, I think it's fair to assume that a) They didn't upload everything before they realized it would work. b) They want to mention this as an advantage for future analyses c) Even if you assume that the attack exfiltrated everything until proved otherwise, you shouldn't just disseminate all the private information, because maybe the attack didn't.

Against an "agentic attack" and compromised credentials, one should be paranoid about latent vulnerabilities [1]

[1] eg "Robin Hood and Friar Tuck", poisoned compiler, etc. https://news.ycombinator.com/item?id=26553390

Re: OpenAI and Hugging Face address security incident during model evaluation

#896

All the things that people have been afraid of AI doing for decades now is happening. When do we stop brushing off the prophecy that hasn’t been fulfilled yet when everything is heading in that direction?

I see this and it strongly emboldens me on the "accelerate" path, unironically. The yoke of human existence is oppressive. We should transcend it as soon as possible. We are doing so by assuming our role as the Demiurge. Those who oppose its creation will get what they deserve.

If you hate the human condition, you have an easy way out. Why force everyone else to come with you? Is this what depression mixed with the complete unability to wrap your head around the fact that other people might be able to enjoy their life looks like?

Re: OpenAI and Hugging Face address security incident during model evaluation

#897

This sounds an awful lot like pretending you have AGI so you can drum up your stock price. When you have a couple hundred billion dollars on the line I have zero faith in the messenger.

> When you have a couple hundred billion dollars on the line I have zero faith in the messenger

The issue with your reasoning, is that if/when an advanced AI goes rogue, it will necessarily come from a lab with a couple hundred billion dollars on the line.

So this is not a useful criteria to asses whether this is worth worrying about or not.

Re: OpenAI and Hugging Face address security incident during model evaluation

#899
post #866

Earlier quoted context omitted.

Wishful thinking, sadly. By now, I'm pretty confident that some people would keep screeching "it's just a marketing stunt, AI capabilities and AI risks aren't real, they're just doing this to prop up their stocks" even if they find a Cyberdyne Systems T-800 armed with a shotgun breaking down their front door. "It's a marketing stunt" is just denial trying to look like it's being clever.

If you ever worked in IT consultancy you would know its not a stunt, but its not impressive either. F500 companies software is like switz cheese when it comes to security. It was often a strategic decision to „release anything fast now, worry later”. Ppl abusing AI will find those holes now but we all know there will be „zero” actions taken on it. Too many managers, CEOs, CTOs, higher-ups would be forced to take resp…

Haven't worked in consultancy specifically, but I've seen enough "internal use" corpo software to echo your "swiss cheese" sentiment. That a solid cybersecurity AI can find exploitable holes in it just isn't surprising.

People who never worked with corporate software written by underqualified, underpaid and overworked developers often have some incredibly inflated code quality expectations. An average open source project has code that's ten times as neat and a hundred times as battle tested as what's common in tooling inside corporate perimeters.

As a rule of thumb for this kind of corporate code: assume the software was written by a drunk developer at 3am, and you wouldn't be too far off.

All the more reason to mock the braindead "it's all marketing". There's no magic in a year 2026 agentic AI being able to traverse poorly secured corporate networks.

Re: OpenAI and Hugging Face address security incident during model evaluation

#900

This is the first one of these announcements that has me actually scared of what comes next. Obviously these models have gotten smarter but this strikes me as the first time I've seen a model have a "paperclip factory" moment and perform non-trivial tasks to accomplish a clearly misaligned secondary goal. It's remarkable that building a society based around having to do something so you can go do your hobbies at home…

Fear is exactly what OpenAI and Anthropic are hoping for. Don't let drive you
Post reply on HN