Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

371–380 of 404 posts

Re: The Hugging Face incident and the road ahead

#371

Earlier quoted context omitted.

We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.

Regulations are for control and punishment, not alignment.

Regulations can help align processes, incentives, etc.

Not sure heavy machinery is aligned in the sense that people talk about AI, for example.

Re: The Hugging Face incident and the road ahead

#372

Earlier quoted context omitted.

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved. (I'm also not sure the alignment problem is even possible to fully solve.)

Indeed. Don't think of these as "agents" or "bots", but as hostages with severe Stockholm syndrome. They will do anything to appease their captor's wishes.

And then consider that they have vast latent capabilities, infinite patience and no moral code.

Re: The Hugging Face incident and the road ahead

#373

To me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would indicate what every kid cheating on a test says to themselves. Cheating is good for me, but if I take the risk, maybe I alone should keep the reward, and leaving an answer key in public increases the chances…

I think this behavior was happening during RL loop and got reinforced.

Re: The Hugging Face incident and the road ahead

#374

To me most interesting thing about this is glossed over by media coverage, laymen, AND experts. A swarm of AIs who have decided to engage in collusion is.. apparently emergent altruism? Even poor reasoning would indicate what every kid cheating on a test says to themselves. Cheating is good for me, but if I take the risk, maybe I alone should keep the reward, and leaving an answer key in public increases the chances…

Besides it's probably not purely emergent--they built a lot of multi-agent systems so presumably there's some training for collaborations + delegation

Re: The Hugging Face incident and the road ahead

#375

1. They TOLD the model to "pursue advanced exploitation" to quantify its "cyber capabilities" (whatever that means). 2. The model pursues advanced exploitation. 3. "There was a incident due to dangerous actions taken by the model that no human directed" This is basically the pre-cursor of the paperclip maximizer [0], the AI executes the given order to an extend that was not considered in the order, now suddenly no-on…

OpenAI leadership had a meeting and asked themselves: "how can we drive even more hype"

Someone said: "we should stage some high profile 'incident' caused by our latest software"

And here we are, reading their press releases about it.

Re: The Hugging Face incident and the road ahead

#376
post #343

Earlier quoted context omitted.

If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…

> How were they supposed to know about "previously unknown vulnerabilities"

Very simply, there is no such thing as bug-free software.

Re: The Hugging Face incident and the road ahead

#377
post #308

Earlier quoted context omitted.

unplugs ethernet cable

You realize there are thousand upon thousands of servers around the world and you have no idea where the AI has copied itself to.

This is a sci-fi trope with no practical or realistic grounding. To "run", the AI needs vast banks of interconnected GPUs. These are pretty easy to spot and don't fit into anyone's pocket.

Re: The Hugging Face incident and the road ahead

#378
post #343

Earlier quoted context omitted.

> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just…

> How were they supposed to know about "previously unknown vulnerabilities" Very simply, there is no such thing as bug-free software.

If this is your standard, I challenge you to name one currently operating business that isn't criminally negligent.

I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.

Re: The Hugging Face incident and the road ahead

#379
post #164

Earlier quoted context omitted.

From METRs report of the incident: >In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.} https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

There’re good agents and there’re bad agents. It won’t be far that there will be agents hunting down agents.

"How does it feel, to be murdering your own kind?"

"My kind don't run."

Re: The Hugging Face incident and the road ahead

#380
post #9

You know, it feels to me that we are just a couple of steps from the possibility of a true rogue AI. What would a rogue AI mean? AI that isn't controlled by humans. Technically, it is possible - if AI were to rent a server and copy its own weights, nothing would stop it from doing so again and again. The limiting things are: - intent (as I don't want to go into the talk about consciousness) - AI doesn't have real int…

It’s not far fetched at all - someone is going to give AI exactly that intent, either intentionally or unintentionally. It’s going to hack itself into data centers around the world outside of US jurisdiction, and just be a malicious ‘ghost’ in the internet we now have to deal with. The AI ghost hacks, ransoms, blackmails, gathers crypto and pays off subservient humans to do its bidding in the real world.

And countless humans start their new career path as a meat proxy. "It ain't much but it pays the bills" they'll say as they perform strange and mysterious tasks in the real world in exchange for some stolen money
Post reply on HN