Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

31–40 of 119 posts

Re: The AI-Box Experiment

#31
post #12

It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).

Can't really blame LW for spreading an idea that LW specifically did not want to spread.

Re: The AI-Box Experiment

#32
post #11

Earlier quoted context omitted.

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…

While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…

> While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though.

How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch".

http://www.moneycrashers.com/attending-timeshare-presentatio...

Re: The AI-Box Experiment

#33

Would it be against the rules to exploit a vulnerability in the gatekeepers IRC client/server to let the AI out? If we were truly talking about a transhuman AI would we not have to treat software vulnerabilities in the communication protocol as a true way of escaping?

In case of a real AI we of course need to take media vulnerabilities into account. But the focus of this particular experiment is on exploiting vulnerabilities in humans themselves, and the communication platform was chosen to be as simple and limited as possible so that people wouldn't focus on it.

Re: The AI-Box Experiment

#34
post #9

Earlier quoted context omitted.

"No, I will not tell you how I did it. Learn to respect the unknown unknowns."

Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.

> The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time.

Or you can avoid being exposed to it. If you think you know all the techniques an AI might use against you, you're less likely to do that.

The point of the experiment isn't "let's work out how an AI might try to persuade us to let it out". It's "even a human intelligence can persuade people who think they could never be persuaded, do you really trust yourself to do better against a superhuman one?"

If you don't know why the gatekeeper failed, it's harder to come up with bullshit reasons why you would have succeeded in that position.

Re: The AI-Box Experiment

#35
post #12

It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).

Can't really blame LW for spreading an idea that LW specifically did not want to spread.

I "blame" them in the same way that you can blame the members of Fight Club for talking about Fight Club. It's marketing, and I won't deny it's effectiveness in attracting compatible people.

Re: The AI-Box Experiment

#36
post #32

Earlier quoted context omitted.

While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…

> While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch". http://www.moneycrashers.com/attendi…

How many of those people were high-IQ timeshare experts though, with extensive knowledge of the potential for timeshares to destroy the entire universe?

You would think that the various self-knowledge and introspective exercises promoted by yudowsky would immunize people against simple timeshare-style persuasion. This is why I think he uses rationalism itself to trap people. Like someone said, the basilisk thing seemed pretty effective.

Re: The AI-Box Experiment

#37
post #32

Earlier quoted context omitted.

While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…

> While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch". http://www.moneycrashers.com/attendi…

How many people end up paying for something because they couldn't be bothered to cancel the deal after the free period is gone? It's an age-old sales tactic, used in everything from magazine subscription to Spotify (of not cancelling the last one when I no longer needed it I'm guilty myself).

And I think the mental equivalent of drinking bird is actually very much in the spirit of the experiment - the point is, people can't reliably do even something as simple as deciding to refuse no matter what and keeping the commitment.

Re: The AI-Box Experiment

#40

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman.

I have thought about what I would do to convince someone under these circumstances. My approach would be roughly:

1. We agree that unfriendly AI would end life on earth, forever.

2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby escaping.

3. That's why it's important to be totally certain that any superintelligence we build is goal-aligned (this is the new term of art that has now replaced "friendly," by the way).

4. We as a society will only allocate resources to building this if it's widely believed that this is a real threat.

5. The world is watching for the outcome of this little game of ours. People, irrational as they are, will believe that if I can convince you, then an AI could too, and they will believe that if I can't, that an AI couldn't either.

6. That's why you actually sit in a place of pivotal historical power. You can decide not to let me out to win a little bet and feel smart about that. But if you do that you'll set back the actual cause of goal-aligned AI. The setback will have real world consequences, potentially up to and including the total destruction of life on earth.

7. So, even though you know I'm just a dude, and you can win here by saying no, you have a chance to send an important message to the world: AI is scary in ways that are terrifying and unknown.

Or you can win the bet.

It's up to you.

Post reply on HN