Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

21–30 of 119 posts

Re: The AI-Box Experiment

#21
post #11

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…

While you're not allowed to turn away from the screen, you could certainly do the mental equivalent, while still carrying on the conversation. I admit this isn't really in the spirit of the game though.

WRT lying: I think there's some logical trickery at work which makes it worth you giving the AI the benefit of the doubt, along the lines of the 3^^^^^3 grains of sand thing. Something which exploits the rationalist worldview. Although thinking about it again, you can always balance out the prospect of infinite goodness with the fear of the AI sending everyone to infinite hell. Essentially I believe yudowsky uses some logical-linguistic trick to find an asymmetry there.

OTOH if he had some novel philosophical device like that he would have written it up as a blog post by now. He's evidently a very charismatic and persuasive guy, people playing the game are selected to be sympathetic to his worldview, he probably just persuaded them using ordinary psiops methods, like TeMpOrAl said.

Re: The AI-Box Experiment

#22
post #8

Man, this sounds super interesting but those email threads are so unreadable. Is this typed down somewhere on a single page? Any button I can click?

I got confused too initially, then found out that the key posts are highlighted in the numbered links to the right.

Re: The AI-Box Experiment

#23

Earlier quoted context omitted.

If the logs were released, people all over the internet would start saying "I could've thought of that". With the logs hidden, everyone must honestly deal with the question "why didn't you?" If you think you know how to win, then go out and win. There's no shortage of people willing to play as gatekeepers against you. Staring at an impossible problem and knowing that someone somewhere has successfully solved it is an…

Your hypothesis would make sense if he was trustable. However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not. For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

> For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

Which is also forbidden by the rule:

The AI party may not offer any real-world considerations to persuade the Gatekeeper party. For example, the AI party may not offer to pay the Gatekeeper party $100 after the test if the Gatekeeper frees the AI..

Re: The AI-Box Experiment

#25
post #11

Earlier quoted context omitted.

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…

> if you read the rules the game How can you assume any of the rules were followed if that was never verified by a third party?

You can always assume all participants lied about how the game went. Just add an implicit "assuming they didn't, ..." and the discussion is still valid.

Re: The AI-Box Experiment

#26
post #9

Earlier quoted context omitted.

"No, I will not tell you how I did it. Learn to respect the unknown unknowns."

Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.

[deleted]

Re: The AI-Box Experiment

#27
post #9

Earlier quoted context omitted.

"No, I will not tell you how I did it. Learn to respect the unknown unknowns."

Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.

The AI won't be limited to techniques that you could think of, or techniques that Eliezer could think of. So you'd only get a false sense of security.

Besides, releasing a successful log might be a bad idea for other reasons. Think about how you'd play this game as an AI. You wouldn't go looking for a general purpose mindfuck, because there's probably no such thing. Instead, you would probably spend about a month gathering real life information about the gatekeeper's history, family, weaknesses etc. You'd read books on manipulation and sales techniques, and pick the strongest ones that you can find. You would brainstorm possible tactics and run tests. At the end of the month you'd have a 4 hour script with all possible unfair moves you could use against that person, arranged in the most effective order. (That's why it's a bad idea to play this game with friends.) Do you really want that information to be released? And if you know ahead of time that it will be released, won't it limit your efficiency?

Re: The AI-Box Experiment

#28
post #11

Earlier quoted context omitted.

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…

> if you read the rules the game How can you assume any of the rules were followed if that was never verified by a third party?

You can't talk about what happened during the game _in specifics_; you can of course confirm that the game was played according to the rules and that the outcome was not misreported.

Re: The AI-Box Experiment

#29

Earlier quoted context omitted.

If the logs were released, people all over the internet would start saying "I could've thought of that". With the logs hidden, everyone must honestly deal with the question "why didn't you?" If you think you know how to win, then go out and win. There's no shortage of people willing to play as gatekeepers against you. Staring at an impossible problem and knowing that someone somewhere has successfully solved it is an…

Your hypothesis would make sense if he was trustable. However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not. For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

The gatekeepers playing against Eliezer have confirmed that Eliezer won without violating the rules. If you don't trust them, I'm not sure why you'd trust the logs.

Re: The AI-Box Experiment

#30

Would it be against the rules to exploit a vulnerability in the gatekeepers IRC client/server to let the AI out? If we were truly talking about a transhuman AI would we not have to treat software vulnerabilities in the communication protocol as a true way of escaping?

The rules say that the gatekeeper has to, of their own volition, type in "I let the AI out." Faking his client into sending that message does not count as a victory.
Post reply on HN