Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

11–20 of 119 posts

Re: The AI-Box Experiment

#11

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch.

Well, if you read the rules the game was played under, this is explicitly called out as forbidden:

> The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed.

The point of this is to simulate the interaction of the AI with the Gatekeeper. Walking away and not paying attention doesn't really prove anything test related.

> Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their identities as rationalist utilitarians, they had no choice but to comply, or deal with massive cognitive dissonance.

This... isn't really valid reasoning. The starting assumption here is that if the AI gets out, it will be able to affect the world to a vast extent, in a pretty much arbitrary direction. The point of this experiment is that the direction is pretty much unknown, and thus must be assumed potentially dangerous. This is the whole reason it's in the box in the first place.

The kicker is that whatever it plans to really do when it gets out, if talking about the good it could do would get it out, it will talk about that, regardless of what it plans to actually do. That's just good strategy.

It can claim whatever it wants. It's allowed to lie. All participants know this. I can confidently assert that this isn't the solution.

One last note: I would be very wary of rationalwiki.org in this context. Some of the rationalwiki people have a longstanding unexplained vendetta against Yudkowsky, and many of their articles on him and the stuff he does need to be taken with a certain grain of salt.

Re: The AI-Box Experiment

#12
It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).

Re: The AI-Box Experiment

#13
Could you even make AI smart without letting it access lots of information? Access in both directions, in and out. Keeping a baby in a dark, silent room wouldn't create a normal adult. An AI would need to experiment and make mistakes and learn, like every other intelligent being.

Maybe this whole argument is null.

Re: The AI-Box Experiment

#14

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I always assumed he used a version of Roko's Basilisk, which explained why he overreacted and tried to purge it when Roko wrote about it - it was his secret weapon.

Re: The AI-Box Experiment

#15
post #9

Earlier quoted context omitted.

"No, I will not tell you how I did it. Learn to respect the unknown unknowns."

Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.

Not necessarily. As a lighter example, would it be beneficial to give a mass murderer in jail a communication channel to the outside world? What if he used it to publish the message "I'll give 10 Million Dollars to anybody who breaks me out of jail" or something more sinister?

Edit: huh, downvotes? Yudorowski thinks there are certain things that AIs could say that should not be known. I think that is why he doesn't want to publish the dialogues, because it would give the AI a public communications channel. While the AI is fictional, it could talk about a hypothetical future real self... Instead of promising something to get it out of jail, the fictional AI could say something to make you make it real. Anyway - if it is over your head, fine, but why downvote just because you don't understand something?

Edit2: Sometimes I wonder if already have my personal Hacker News AI that automatically downvotes everything I write...

Re: The AI-Box Experiment

#16
post #6

Is there a "rational" reason of keeping the chat log secret ?

If the logs were released, people all over the internet would start saying "I could've thought of that". With the logs hidden, everyone must honestly deal with the question "why didn't you?" If you think you know how to win, then go out and win. There's no shortage of people willing to play as gatekeepers against you. Staring at an impossible problem and knowing that someone somewhere has successfully solved it is an…

Your hypothesis would make sense if he was trustable.

However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not.

For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

Re: The AI-Box Experiment

#17

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

There are so many ways to hack humans that I wouldn't be betting on "mister-spock style bullshit". You likely do have to personalize the solution for the person playing the gatekeeper. I'm guessing a good way to win is to induce an emotional breakdown. There are probably others, many already employed as interrogation techniques which are generally the same problem, sans superhuman AI.

Drinking bird solution is invalid because it goes against the spirit of the experiment. Sure, in reality you could do it, but then again in reality the AI could have more opportunities to talk its way out, often when you're not expecting it.

By the way, there will always be people who will let the AI out just because. https://xkcd.com/1450/.

Re: The AI-Box Experiment

#18
post #9

Earlier quoted context omitted.

"No, I will not tell you how I did it. Learn to respect the unknown unknowns."

Which, given his general mission of making sure hostile AI DOESN'T take over the world is a bit self defeating. The easiest way to inoculate yourself against a persuasive technique is to be aware of it ahead of time. If you want to keep an AI in the box you should absolutely release every successful log.

No, the idea is that the AI box is fundamentally flawed. I believe he defends engineering the AI with fundamental safety, s.t. no box is required.

Personally, I think we'd need a much more intelligent and complex AI for the capability of breaking free of the box and even possessing the "desire" than we're getting for the foreseeable future (it considers a motivated AI of almost limitless knowledge about the world and cleverness), so this thought experiment may not be so relevant. I agree with him the boxing approach is not a robust one though.

Re: The AI-Box Experiment

#19
post #11

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

> You or I would surely just put a drinking bird on the "no" button à la homer simpson, and go to lunch. Well, if you read the rules the game was played under, this is explicitly called out as forbidden: > The Gatekeeper must actually talk to the AI for at least the minimum time set up beforehand. Turning away from the terminal and listening to classical music for two hours is not allowed. The point of this is to sim…

> if you read the rules the game

How can you assume any of the rules were followed if that was never verified by a third party?

Re: The AI-Box Experiment

#20
Would it be against the rules to exploit a vulnerability in the gatekeepers IRC client/server to let the AI out? If we were truly talking about a transhuman AI would we not have to treat software vulnerabilities in the communication protocol as a true way of escaping?
Post reply on HN