Live data from Hacker News

The Hugging Face incident and the road ahead

openai.com

331–340 of 342 posts

Re: The Hugging Face incident and the road ahead

#331

Earlier quoted context omitted.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.

People have been saying that since about the invention of the internet.

There's already a bunch of documented ways to exploit system hardware to jump airgaps. Bang the system bus the right way and it's a radio antenna that can directly connect to nearby mobile phones.

The easiest one is, of course, sending a message to a human saying "yo, I need internet". Humans are eager to please and easily fooled, and anthropomorphise everything: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims

And that's just for good humans. The moment we got AI worth a penny, everyone with money to invest put a model on the web and tried to charge for access to it.

Re: The Hugging Face incident and the road ahead

#332

Earlier quoted context omitted.

There are limits, but those limits are unknown. Do you disagree?

I don’t need to know the value of their limit, I just need to know their bounds. Just like a prison doesn’t need to know the strength of each inmate, just that they can’t bend or bite through steel bars. Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.

Cryptography is real, but nobody in that field seems to be hubristic enough to think their methods are flawless, and there's a degree of suspicion than the best models may have secret weaknesses engineered into them by the governments who sponsored them.

Physics is real and networking requires a substrate. But there are already known exploits which can misuse the compute hardware as an antenna, e.g. my first search result: https://github.com/fulldecent/system-bus-radio

(Older nerds may remember https://en.wikipedia.org/wiki/Van_Eck_phreaking)

> Magic is not real

  Turning lead into gold isn't the magic of alchemy, it's just nucleosynthesis.

  Taking a living human's heart out without killing them, and replacing it with one you got out a corpse, that isn't the magic of necromancy, neither is it a prayer or ritual to Sekhmet, it's just transplant surgery.

  ...

  Reading someone’s thoughts isn't magic telepathy, it's just fMRI decoding.

  ...

  Seeing someone's bones without flaying the flesh from them isn't magic, it's just an x-ray.

  Curing congenital deafness, letting the blind see, letting the lame walk, none of that is magic or miracle, they're just cochlear implants, cataract removal/retinal implants, and surgery or prosthetic exoskeletons respectively.
- me, https://www.lesswrong.com/posts/hAwvJDRKWFibjxh4e/it-isn-t-m...

Re: The Hugging Face incident and the road ahead

#333

Earlier quoted context omitted.

Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.

The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)

I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion.

Tens of millions, even.

Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection].

> And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.

Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai

Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2

Re: The Hugging Face incident and the road ahead

#334

I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities Model is told…

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

Maybe it should be called 'ambiguity' problem then. An issue that famously emerges from trying to use natural language for instructing computers: https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...

Re: The Hugging Face incident and the road ahead

#335
post #333

Earlier quoted context omitted.

The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)

I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion . Tens of millions, even. Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica]…

I am not anthropomorphising the LLMs at all, I was talking about the obligations we put on humans using/making/etc. machines etc.

Re: The Hugging Face incident and the road ahead

#336
post #323

Earlier quoted context omitted.

> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task. The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) compla…

If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up

Maybe the test/task itself wasn't intended as a marketing stunt. But the response to fallout with "going rouge" certainly was.

The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.

Re: The Hugging Face incident and the road ahead

#337

Earlier quoted context omitted.

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. But this has actually happened... a lot. Search "social engineering prison breaks". With AI it only needs to happen once. I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "y…

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas. There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

They don't need to be omnipotent, and they're already human-or-superhuman at persuasion: https://arxiv.org/html/2411.06837v2

This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.

Re: The Hugging Face incident and the road ahead

#338
post #333

Earlier quoted context omitted.

I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion . Tens of millions, even. Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica]…

I am not anthropomorphising the LLMs at all, I was talking about the obligations we put on humans using/making/etc. machines etc.

Hmm. I think I misunderstood what you meant by "experience policing people around financial incentives" in that case.

Re: The Hugging Face incident and the road ahead

#339
post #337

Earlier quoted context omitted.

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas. There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

They don't need to be omnipotent, and they're already human-or-superhuman at persuasion: https://arxiv.org/html/2411.06837v2 This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.

That is about persuasion with evidence on various topics, not about persuading people to abandon safty protocols and processes and highly policed settings.

Yes, many things could happen, but again, that failure is possible is not a reason to do implement processes etc. I don't see why hypotheticals should stop addressing actuals.

Re: The Hugging Face incident and the road ahead

#340

Earlier quoted context omitted.

This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it…

Maybe it should be called 'ambiguity' problem then. An issue that famously emerges from trying to use natural language for instructing computers: https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...

Or indeed natural language to instruct humans.

If it was easy to specify exactly the behaviour you wanted then we probably wouldn't have contract law.

Post reply on HN