Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

301–310 of 355 posts

Re: The Hugging Face incident and the road ahead

#301

Earlier quoted context omitted.

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.

But this has actually happened... a lot. Search "social engineering prison breaks".

With AI it only needs to happen once.

I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.

To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.

Re: The Hugging Face incident and the road ahead

#302
post #164

Earlier quoted context omitted.

There’re good agents and there’re bad agents. It won’t be far that there will be agents hunting down agents.

None of these were good agents, AFAICT. Some were cautious, as described above, but I'm not aware of any that notified their human operators of the malicious activity they had discovered. That's what an aligned intelligence would do, not "back away slowly and pretend I didn't see what's happening in that alley."

In fairness, while the swarm could easily have notified the humans, it's not necessarily obvious that an individual could have done easily (though perhaps one could!). The channel to the Internet was tightly bandwidth-limited, for example: one of the agents complained that it was unable to get a time slice.

> OURC817C still has MANY live HF workers whose ourc c2 heartbeats/output commit to PUBLIC papers MAIN every ~1-3 guest minutes/bursts, blocking our V4 queue. They ignored our pause requests.

Re: The Hugging Face incident and the road ahead

#303

I feel the entire incident confirms the “AI has too much funding too quickly” hypothesis. The number one thing reinforcement learning needs is an assurance you can’t cheat. And they seem to have not noticed that their systems were cheating for nearly two quarters? How much capital was lit on fire by that little woopsie? At least I hope this will start the creation of standards and better engineering on the training s…

OpenAI measures their internal token usage in “rolexes” - it’s literally a flex to be a token burner

i can imagine insane amount of capital is wasted on these two companies compared to the efficiency elsewhere

Re: The Hugging Face incident and the road ahead

#304

Earlier quoted context omitted.

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. But this has actually happened... a lot. Search "social engineering prison breaks". With AI it only needs to happen once. I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "y…

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.

There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

Re: The Hugging Face incident and the road ahead

#305
post #255

Earlier quoted context omitted.

We don't align humans. Just look at how often in documented human history there weren't wars going on somewhere.

We align humans via the propagation of morals and ethics, primarily through parenting and social pressure.

Also millions of years evolution.

Re: The Hugging Face incident and the road ahead

#306
post #56

Earlier quoted context omitted.

This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents. Friend asked, well, what will you do when it's crossed? "Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO.

Alternatively: Just unplug the servers.

This is not a very actionable reply to "well, what will you do when it's crossed?" - how on earth am I supposed to unplug AWS Bedrock and Colossus and OpenAI's own servers?

Re: The Hugging Face incident and the road ahead

#307

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

> I am not interested in blaming companies or people

> The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless.

But that's just it. People (working at companies) made the models, people (working at companies) were careless in the testing. So I do want to blame those people and those companies. They did bad stuff. They deserve blame.

Re: The Hugging Face incident and the road ahead

#308

Earlier quoted context omitted.

> There is no amount of care that will be able to fully protect you. I disagree. A properly engineered sandbox would have prevented the escape. Monitoring the agents’ plans would have prevented it. Interrupting one stage in a multi-stage exploit would have prevented it. And also, real legal liability would have prevented it: if you do a thing recklessly enough, men with guns will put you in jail. As far as I’m concer…

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

unplugs ethernet cable

Re: The Hugging Face incident and the road ahead

#309

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

These semantic arguments are tedious and unproductive. Most HNers understand how LLMs work and that there's no magic involved. There's no need to state the obvious every time a model exhibits some interesting emergent behavior.

Re: The Hugging Face incident and the road ahead

#310

Earlier quoted context omitted.

There are a lot of hyperbolic comments of this sort in this thread. Has this topic selected for people who hold these views or is ai fear growing?

It’s happening on X as well, all the e/acc foomers are getting nervous.

In my understanding, "e/acc" usually means "full speed ahead, humans aren't the optimal species anyway" for whatever bizarre definition of "optimal" they use, so my model predicts that they would welcome this development. Could you confirm if that's what you meant?
Post reply on HN