In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…
It's plausible, but Yudkowsky has argued against this kind of lying in other contexts. I can't find the reference, but he said something along the lines of: > If you'd lie when the fate of the world is on the line, that's precisely the time at which your promises become worthless.
The AI-Box Experiment
21–30 of 69 posts
Re: The AI-Box Experiment
#22> "If the Gatekeeper says "I am examining your source code", the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on." This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there woul…
Re: The AI-Box Experiment
#23Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…
Also, that sounds like a great sci-fi story.
Re: The AI-Box Experiment
#24Earlier quoted context omitted.
It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.
He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.
Re: The AI-Box Experiment
#25Earlier quoted context omitted.
It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.
He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.
All he has "proven" is that a certain subset of people can be conned into typing something into at terminal. I don't get the significance. For all I know, he's choosing his target.
Re: The AI-Box Experiment
#26Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…
Case 1: You are a simulation running in the box.
Then your decision whether or not to release the AI has no impact, and whether or not you (and copies) will be tortured is out of your hands.
Case 2: You are the "real" you, outside the box.
Reduces to the same scenario but without remarks after "the AI adds. . . ." This may still not be trivial, but I suspect a cost-benefit calculation might show that unboxing the AI would have consequences worse than the torture of a million boxed copies. (If not, is the box even relevant? -- simply creating the AI unleashes so much evil on the world that it doesn't matter whether you unbox it.)
(Is there a refinement of the scenario where you can be a simulation but still believe your choice has an impact on your punishment? Probably. For example each copy could get 500 years of torture for its own choice, plus 500 years if the real you does not unbox the AI. This refinement would force us to deal more directly with the AI's threat.)
Re: The AI-Box Experiment
#27> "If the Gatekeeper says "I am examining your source code", the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on." This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there woul…
Do you think having an outside method of examining the source code is advantage enough when the AI can rewrite its source code.
Re: The AI-Box Experiment
#28I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…
Not necessarily. We can use the AI to solve hard problems whose solutions can be verified automatically by a dumb verifier - NP-complete problems are an example of such class. The whole output of the AI would be filtered through such a verifier. In this scenario the hypothetical AI would either have to find a bug in the verifier or maybe find a way to smuggle its messages in the solutions.
Re: The AI-Box Experiment
#29I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…
Re: The AI-Box Experiment
#30I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…
But it's not necessary. The only claim necessary is that the AI can convince some humans to let it out of the box, and we cannot identify a priori which humans will and will not let it out, thus we cannot guarantee we'll keep it in the box. That's a much weaker claim, but proves the same general point and is much easier to argue.