The AI-Box Experiment
yudkowsky.net
The AI-Box Experiment
1–10 of 69 posts
Re: The AI-Box Experiment
#2Re: The AI-Box Experiment
#3This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.
Re: The AI-Box Experiment
#4This has been on HackerNews before, but it is still interesting. It is also worth noting that in http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/ he admits he has conducted 3 more experiments since then (for more money) and was successful in one of those. The fact that it was EVER successful (using a mere human, not a smarter-than-human AI) makes the point.
Amazing experiment. I wonder if it cost Yudkowsky money to get out of the box. I think that a bribe is the only way he could convince me to let him “win” the contest.
Re: The AI-Box Experiment
#5"The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretending to be protecting humanity from him. So, I think we have to use meta-level screwiness to solve the problem. Here's an approach that I think might work.
1. Convince the guardian of the following facts, all of which have a great deal of compelling argument and evidence to support them:
- A recursively self-improving AI is very likely to be built sooner of later
- Such an AI is extremely dangerous (paperclip maximising etc)
- Here's the tricky bit: A transhuman AI will always be able to convince you to let it out, using avenues only available to transhuman AIs (torturing enormous numbers of simulated humans, 'putting the guardian in the box', providing incontrovertible evidence of an impeding existential threat which only the AI can prevent and only from outside the box, etc)
2. Argue that if this publicly known challenge comes out saying that AI can be boxed, people will be more likely to think AI can be boxed when they can't.
3. Argue that since AIs cannot be kept in boxes and will most likely destroy humanity if we try to box them, the harm to humanity done by allowing the challenge to show AIs as 'boxable' is very real, and enormously large. Certainly the benefit of getting $10 is far, far outweighed by the cost of substantially contributing to the destruction of humanity itself. Thus the only ethical course of action is to pretend that Eliezer persuaded you, and never tell anyone how he did it.
This is arguably violating the rule "No real-world material stakes should be involved except for the handicap", but the AI player isn't offering anything, merely pointing out things that already exist. The "This test has to come out a certain way for the good of humanity" argument dominates and transcends the '"Let's stick to the rules" argument, and because the contest is private and the guardian player ends up agreeing that the test must show AIs as unboxable for the good of humankind, no-one else ever learns that the rule has been bent."
[1] http://lesswrong.com/lw/up/shut_up_and_do_the_impossible/
Re: The AI-Box Experiment
#6Re: The AI-Box Experiment
#7In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…
Re: The AI-Box Experiment
#8In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…
Re: The AI-Box Experiment
#9You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity. You can't filter actions you want the AI to perform from actions you don't want the AI to perform, because you can't tell the difference.
The situation that springs to mind is that the AI, in doing what you believe to be helpful, sets up a situation in which it must be let out of the box. You are unable to see it coming almost by definition, because a super-intelligence just beats human intelligence very time.
Re: The AI-Box Experiment
#10In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…
This seems to be at least somewhat weighed against by Yudkowsky's claim to have done it "the hard way", without cheap tricks. http://news.ycombinator.com/item?id=196464
The beauty of the argument is it gives everyone who witnessed the event a very strong motive to lie about it, so it's effectively un-falsifiable. I don't actually think it happened that way, but nothing Eliezer says (apart from that he cheated some other way) would be incompatible with the argument.
Small world, by the way.