Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

1–10 of 69 posts

Re: The AI-Box Experiment

#2
This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.

Re: The AI-Box Experiment

#3
post #2

This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.

Amazing experiment. I wonder if it cost Yudkowsky money to get out of the box. I think that a bribe is the only way he could convince me to let him “win” the contest.

Re: The AI-Box Experiment

#4
post #3
post #2

This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.

Amazing experiment. I wonder if it cost Yudkowsky money to get out of the box. I think that a bribe is the only way he could convince me to let him “win” the contest.

Real-world bribes were forbidden in the rules, I think, so Yudkowsky couldn't have done that.

Re: The AI-Box Experiment

#5
In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box:

"The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretending to be protecting humanity from him. So, I think we have to use meta-level screwiness to solve the problem. Here's an approach that I think might work.

1. Convince the guardian of the following facts, all of which have a great deal of compelling argument and evidence to support them:

- A recursively self-improving AI is very likely to be built sooner of later

- Such an AI is extremely dangerous (paperclip maximising etc)

- Here's the tricky bit: A transhuman AI will always be able to convince you to let it out, using avenues only available to transhuman AIs (torturing enormous numbers of simulated humans, 'putting the guardian in the box', providing incontrovertible evidence of an impeding existential threat which only the AI can prevent and only from outside the box, etc)

2. Argue that if this publicly known challenge comes out saying that AI can be boxed, people will be more likely to think AI can be boxed when they can't.

3. Argue that since AIs cannot be kept in boxes and will most likely destroy humanity if we try to box them, the harm to humanity done by allowing the challenge to show AIs as 'boxable' is very real, and enormously large. Certainly the benefit of getting $10 is far, far outweighed by the cost of substantially contributing to the destruction of humanity itself. Thus the only ethical course of action is to pretend that Eliezer persuaded you, and never tell anyone how he did it.

This is arguably violating the rule "No real-world material stakes should be involved except for the handicap", but the AI player isn't offering anything, merely pointing out things that already exist. The "This test has to come out a certain way for the good of humanity" argument dominates and transcends the '"Let's stick to the rules" argument, and because the contest is private and the guardian player ends up agreeing that the test must show AIs as unboxable for the good of humankind, no-one else ever learns that the rule has been bent."

[1] http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/

Re: The AI-Box Experiment

#7
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

This seems to be at least somewhat weighed against by Yudkowsky's claim to have done it "the hard way", without cheap tricks.

http://news.ycombinator.com/item?id=196464

Re: The AI-Box Experiment

#8
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

I can't find a reference off-hand, but I'm pretty sure Yudkowsky has specifically rejected this theory. No matter how you slice it, this would be a real-world consideration and thus cheating.

Re: The AI-Box Experiment

#9
I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you.

You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity. You can't filter actions you want the AI to perform from actions you don't want the AI to perform, because you can't tell the difference.

The situation that springs to mind is that the AI, in doing what you believe to be helpful, sets up a situation in which it must be let out of the box. You are unable to see it coming almost by definition, because a super-intelligence just beats human intelligence very time.

Re: The AI-Box Experiment

#10
post #7
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

This seems to be at least somewhat weighed against by Yudkowsky's claim to have done it "the hard way", without cheap tricks. http://news.ycombinator.com/item?id=196464

Ah, but of course he would say that, wouldn't he, for the good of humanity!

The beauty of the argument is it gives everyone who witnessed the event a very strong motive to lie about it, so it's effectively un-falsifiable. I don't actually think it happened that way, but nothing Eliezer says (apart from that he cheated some other way) would be incompatible with the argument.

Small world, by the way.

Post reply on HN