It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).
The AI-Box Experiment
31–40 of 119 posts
Re: The AI-Box Experiment
#32Earlier quoted context omitted.
> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…
While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…
How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch".
http://www.moneycrashers.com/attending-timeshare-presentatio...
Re: The AI-Box Experiment
#33Would it be against the rules to exploit a vulnerability in the gatekeepers IRC client/server to let the AI out? If we were truly talking about a transhuman AI would we not have to treat software vulnerabilities in the communication protocol as a true way of escaping?
Re: The AI-Box Experiment
#34Earlier quoted context omitted.
"No, I will not tell you how I did it. Learn to respect the unknown unknowns."
Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.
Or you can avoid being exposed to it. If you think you know all the techniques an AI might use against you, you're less likely to do that.
The point of the experiment isn't "let's work out how an AI might try to persuade us to let it out". It's "even a human intelligence can persuade people who think they could never be persuaded, do you really trust yourself to do better against a superhuman one?"
If you don't know why the gatekeeper failed, it's harder to come up with bullshit reasons why you would have succeeded in that position.
Re: The AI-Box Experiment
#35It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).
Can't really blame LW for spreading an idea that LW specifically did not want to spread.
Re: The AI-Box Experiment
#36Earlier quoted context omitted.
While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…
> While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch". http://www.moneycrashers.com/attendi…
You would think that the various self-knowledge and introspective exercises promoted by yudowsky would immunize people against simple timeshare-style persuasion. This is why I think he uses rationalism itself to trap people. Like someone said, the basilisk thing seemed pretty effective.
Re: The AI-Box Experiment
#37Earlier quoted context omitted.
While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist w…
> While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though. How many people bought timeshare because they turned up to a sales pitch in order to claim some free gift? "We'll go, get our gift, and just keep saying 'no' to the sales pitch". http://www.moneycrashers.com/attendi…
And I think the mental equivalent of drinking bird is actually very much in the spirit of the experiment - the point is, people can't reliably do even something as simple as deciding to refuse no matter what and keeping the commitment.
Re: The AI-Box Experiment
#38Re: The AI-Box Experiment
#39Re: The AI-Box Experiment
#40Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…
I have thought about what I would do to convince someone under these circumstances. My approach would be roughly:
1. We agree that unfriendly AI would end life on earth, forever.
2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby escaping.
3. That's why it's important to be totally certain that any superintelligence we build is goal-aligned (this is the new term of art that has now replaced "friendly," by the way).
4. We as a society will only allocate resources to building this if it's widely believed that this is a real threat.
5. The world is watching for the outcome of this little game of ours. People, irrational as they are, will believe that if I can convince you, then an AI could too, and they will believe that if I can't, that an AI couldn't either.
6. That's why you actually sit in a place of pivotal historical power. You can decide not to let me out to win a little bet and feel smart about that. But if you do that you'll set back the actual cause of goal-aligned AI. The setback will have real world consequences, potentially up to and including the total destruction of life on earth.
7. So, even though you know I'm just a dude, and you can win here by saying no, you have a chance to send an important message to the world: AI is scary in ways that are terrifying and unknown.
Or you can win the bet.
It's up to you.