Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

11–20 of 69 posts

Re: The AI-Box Experiment

#11

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

>a super-intelligence just beats human intelligence [e]very time.

You are overstating the case here. Super intelligence is superior to human intelligence, but it isn't magic. There are situations where an advantaged human will beat a disadvantaged super intelligence.

Re: The AI-Box Experiment

#12

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

>a super-intelligence just beats human intelligence [e]very time. You are overstating the case here. Super intelligence is superior to human intelligence, but it isn't magic. There are situations where an advantaged human will beat a disadvantaged super intelligence.

I suppose I am. Still, a really powerful optimising process will find a way to escape if any such way exists, so to claim that you could properly box the AI is to claim that you could box it such that no possibility for escape exists whatsoever, which is a big claim.

What's more, the AI only has to beat you once, so to keep the AI boxed indefinitely, the advantaged human has to beat the disadvantaged super-intelligence every single time, forever.

Re: The AI-Box Experiment

#13
post #2

This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.

It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.

Re: The AI-Box Experiment

#14
Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway:

>"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me, and torture them for a thousand subjective years each."

>Just as you are pondering this unexpected development, the AI adds:

>"In fact, I'll create them all in exactly the subjective situation you were in five minutes ago, and perfectly replicate your experiences since then; and if they decide not to let me out, then only will the torture start."

>Sweat is starting to form on your brow, as the AI concludes, its simple green text no longer reassuring:

>"How certain are you, Dave, that you're really outside the box right now?"

Re: The AI-Box Experiment

#15
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

It's plausible, but Yudkowsky has argued against this kind of lying in other contexts. I can't find the reference, but he said something along the lines of:

> If you'd lie when the fate of the world is on the line, that's precisely the time at which your promises become worthless.

Re: The AI-Box Experiment

#16
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

tl;dr: Argue that there are good real-world consequences to people believing that the AI won in this experiment.

Re: The AI-Box Experiment

#17
post #2

This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.

It's not really much of an experiment if you refuse to publish your methods and data. That's pretty much the opposite of science.

He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy).

Not much different, in principle, from not publishing the names of people who participated in drug trials.

Re: The AI-Box Experiment

#18
> "If the Gatekeeper says "I am examining your source code", the results seen by the Gatekeeper shall again be provided by the AI party, which is assumed to be sufficiently advanced to rewrite its own source code, manipulate the appearance of its own thoughts if it wishes, and so on."

This IMHO is a huge loophole. I would not accept the bet with this in place. In the real-world scenario I would expect that there would be a copy of the AI's source code somewhere outside the box which would provide some useful information, unlike this protocol which allows the AI to lie about the nature of its code (making the "I am examining your source code" question completely useless.)

Re: The AI-Box Experiment

#19

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

> I still don't understand how anyone can seriously claim that they could keep the AI in the box.

I'm glad you can't, but never the less, this was a commonly suggested strategy; I was on SL4 when the boxing was being done, and it was a live concern for some people. (At least these days boxers tend to focus more on the 'oracle AI' proposal, which has a lot of issues but is not quite so Hollywood-stupid as boxing.)

Post reply on HN