Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

21–30 of 69 posts

Re: The AI-Box Experiment

#21
post #15
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

It's plausible, but Yudkowsky has argued against this kind of lying in other contexts. I can't find the reference, but he said something along the lines of: > If you'd lie when the fate of the world is on the line, that's precisely the time at which your promises become worthless.

http://lesswrong.com/lw/v2/prices_or_bindings/

Found from http://lesswrong.com/lw/6w/degrees_of_radical_honesty/

Re: The AI-Box Experiment

#22
post #18

> "If the Gatekeeper says "I am examining your source code", the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on." This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there woul…

Do you think having an outside method of examining the source code is advantage enough when the AI can rewrite its source code.

Re: The AI-Box Experiment

#23
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

At first I thought I would never let a transhuman AI out of the box, but after reading that and thinking about it... wow!

Also, that sounds like a great sci-fi story.

Re: The AI-Box Experiment

#24
post #17

Earlier quoted context omitted.

It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.

He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.

No, it's very different from that. It's more along the lines of demonstrating a drug that cures cancer, but refusing to tell anyone its chemical composition or how to make it.

Re: The AI-Box Experiment

#25
post #17

Earlier quoted context omitted.

It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.

He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.

If it's a science "experiment", his strategy would have to be revealed so you can reproduce it. Names of people participating in drug trials is not required to reproduce an experiment. In principal this makes it different from not publishing the names of people who participated in drug trials.

All he has "proven" is that a certain subset of people can be conned into typing something into at terminal. I don't get the significance. For all I know, he's choosing his target.

Re: The AI-Box Experiment

#26
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

I don't really like that argument. Even granting that you should consider the possibility that you are a simulation running in the box (you might believe that this is all but certain), I'm not sure you have reason to let the AI out. Consider:

Case 1: You are a simulation running in the box.

Then your decision whether or not to release the AI has no impact, and whether or not you (and copies) will be tortured is out of your hands.

Case 2: You are the "real" you, outside the box.

Reduces to the same scenario but without remarks after "the AI adds. . . ." This may still not be trivial, but I suspect a cost-benefit calculation might show that unboxing the AI would have consequences worse than the torture of a million boxed copies. (If not, is the box even relevant? -- simply creating the AI unleashes so much evil on the world that it doesn't matter whether you unbox it.)

(Is there a refinement of the scenario where you can be a simulation but still believe your choice has an impact on your punishment? Probably. For example each copy could get 500 years of torture for its own choice, plus 500 years if the real you does not unbox the AI. This refinement would force us to deal more directly with the AI's threat.)

Re: The AI-Box Experiment

#27
post #22
post #18

> "If the Gatekeeper says "I am examining your source code", the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on." This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there woul…

Do you think having an outside method of examining the source code is advantage enough when the AI can rewrite its source code.

Yes, because examining the old source code allows you to predict its behaviour, including the rewriting of source code. If line 42 says "never rewrite lines 42 or 43" and line 43 says "never kill humans" you would be more likely to let it out of the box than if line 42 said "rewrite whatever you want" and line 43 said "do whatever is necessary to achieve world domination."

Re: The AI-Box Experiment

#28

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

> You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity.

Not necessarily. We can use the AI to solve hard problems whose solutions can be verified automatically by a dumb verifier - NP-complete problems are an example of such class. The whole output of the AI would be filtered through such a verifier. In this scenario the hypothetical AI would either have to find a bug in the verifier or maybe find a way to smuggle its messages in the solutions.

Re: The AI-Box Experiment

#29

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

Just because someone is smarter than someone doesn't mean complete power. Alot of people are smarter than their bosses, but you know what the bosses have on their favour? The power to terminate the employee. If I have the power to terminate the AI at any time, as long as I don't give that power up, I will have power over it.

Re: The AI-Box Experiment

#30

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

That the AI can always get out of the box is a very strong claim.

But it's not necessary. The only claim necessary is that the AI can convince some humans to let it out of the box, and we cannot identify a priori which humans will and will not let it out, thus we cannot guarantee we'll keep it in the box. That's a much weaker claim, but proves the same general point and is much easier to argue.

Post reply on HN