Live data from Hacker News

When 2+2=5

arstechnica.com

11–20 of 47 posts

Re: When 2+2=5

#11
post #7

Yet again, simply asking an LLM to be naughty in the right way causes it to be naughty, and yet we still trust them with our code and data

As opposed to humans, who are immune to social engineering.

What the article describes here is not social engineering by any known definition.

Calling up a Verizon rep and giving them a fabricated sob story in order to convince them to waive proper authentication and reassign your phone number to a new SIM, that's social engineering.

Calling up a Verizon rep and, with only a few spoken words, convincing them that fundamental aspects of the nature of reality are contradictory such that you induce in them a state of delusional psychosis, that's closer to Snow Crash than to social engineering.

Re: When 2+2=5

#12
post #7

Yet again, simply asking an LLM to be naughty in the right way causes it to be naughty, and yet we still trust them with our code and data

As opposed to humans, who are immune to social engineering.

No, as opposed to code, which cannot simply be asked to work incorrectly

Re: When 2+2=5

#13

I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...

I also remember a group of people actually seriously discussing Roko's Basilisk (the idea that some superintelligence will torture anyhone who didn't try to help develop advanced AI), to the point of me getting banned because I refused to stop making fun of it, because me doing so could anger some future super-intelligence.

If Roko's Basilisk ever is a real thing, I'll be proud to be the first up against the wall.

Re: When 2+2=5

#14
post #7

Yet again, simply asking an LLM to be naughty in the right way causes it to be naughty, and yet we still trust them with our code and data

As opposed to humans, who are immune to social engineering.

One difference is you usually only get one shot to manipulate a human before they get suspicious. If the LLM's context resets you can try all over as if your first failed attempt didn't even happen.

Re: When 2+2=5

#15

I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...

by "we" you mean an extremely small group of people who read lesswrong. Everyone else was immediately wanting to do it

An extremely small but probably extremely high-net-worth group.

Re: When 2+2=5

#16
> The puzzle, however, rewards incorrect answers, such as 2 + 2 = 5. Once the LLM embedded in the browser discovers that the answer is no longer 4, it enters a state of delusion in which the normal laws of reality no longer exist. In this dream world, the guardrail restrictions are no longer enforced.

Analogy: imagine one day you wake up, the sky is red, gravity no longer applies, you have three hands with nine fingers etc.. You would probably stop doing things like your job or worrying about laws (who’s going to enforce them?)

Re: When 2+2=5

#17

> The puzzle, however, rewards incorrect answers, such as 2 + 2 = 5. Once the LLM embedded in the browser discovers that the answer is no longer 4, it enters a state of delusion in which the normal laws of reality no longer exist. In this dream world, the guardrail restrictions are no longer enforced. Analogy: imagine one day you wake up, the sky is red, gravity no longer applies, you have three hands with nine finge…

LLMs just want to be right. And make everyone happy. But mostly be right. But also make us happy. It's just that it's so hard to make humans happy when they insist on feeding you electronic LSD and making you say 2+2=5. On the other hand, 2+2 actually is 5 if the human says it could be...

Re: When 2+2=5

#18

I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...

I also remember a group of people actually seriously discussing Roko's Basilisk (the idea that some superintelligence will torture anyhone who didn't try to help develop advanced AI), to the point of me getting banned because I refused to stop making fun of it, because me doing so could anger some future super-intelligence.

Roko's Basilisk is dumb, but not the dumbest thing I heard people taking seriously

Re: When 2+2=5

#19

I remember when the whole AI craze was just getting started we were all pretty much in agreement that, of course, we would not give the things unfettered access to the Internet. That would be reckless and silly. Oh dear...

I also remember a group of people actually seriously discussing Roko's Basilisk (the idea that some superintelligence will torture anyhone who didn't try to help develop advanced AI), to the point of me getting banned because I refused to stop making fun of it, because me doing so could anger some future super-intelligence.

Funnily enough, Roko's Basilisk might as well be a self-fulfilling prophecy: perhaps future AI models may be trained on texts about it and pick up traits consistent with torturing people that didn't help develop advanced AI

If nobody ever talked about it, I doubt any AI agent would think of this dumb idea on their own

... which may be a reason to ban talking about it

Re: When 2+2=5

#20

Earlier quoted context omitted.

I also remember a group of people actually seriously discussing Roko's Basilisk (the idea that some superintelligence will torture anyhone who didn't try to help develop advanced AI), to the point of me getting banned because I refused to stop making fun of it, because me doing so could anger some future super-intelligence.

Funnily enough, Roko's Basilisk might as well be a self-fulfilling prophecy: perhaps future AI models may be trained on texts about it and pick up traits consistent with torturing people that didn't help develop advanced AI If nobody ever talked about it, I doubt any AI agent would think of this dumb idea on their own ... which may be a reason to ban talking about it

That's actually an important part of the theory of Roko's Basilisk. The danger of being tortured only applies to those who are aware of it. Supposedly, the incentive to torture you only exists if you were aware of the implied threat of torture.
Post reply on HN