Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

101–110 of 119 posts

Re: The AI-Box Experiment

#101

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

Points 4-7 are predicated on point 3 being effective, but it's just a benign-seeming action.

I also disagree with point 1, but since you just mean that unfriendly AI is something I wouldn't want around, I'll let it slide.

Re: The AI-Box Experiment

#102
post #81

Earlier quoted context omitted.

Yeah but imagine if all the bullshit about "You don't talk about Fight Club" was actually blown up by a third group whose sole intent was making fun of Fight Club. Imagine if the members of Fight Club _actually_ didn't (start to) talk about Fight Club. But for some reason, everyone else brings it up all the time. Then you could maybe see how talk about Fight Club might not be Fight Club's fault, and in fact highly an…

You're being a bit uncharitable in your interpretation of my argument here, but I get where you're coming from now. I'm not an LW hater. For a long time, I didn't really have an opinion on LW both as a community nor as a philosophical framework. I do consider myself a transhumanist, though. There are three concepts I do know from and about LW: their take on rationality, the top secret AI unboxing strategy, and the Ba…

I do not think the trick can not be repeated, it is in fact, the other way around.

For the AI to be infallible, it needs to have a network of tricks and arguments and each one of them had to be created for a very particular person, even if later it can be repeated on others.

It is like Christianity. There's not a simple belief that everyone accepts and that's it. There's only the façade of simplicity, but in reality it is anything but simple.

There's a network of related explanations and rationalizations that has been expanding for centuries, and every time someone appears who will not accept the network of arguments, a new argument or explanation has to be added to account for that person.

For example, if Christianity had not need to deal with Gnosticism, or with Arianism, then the network of beliefs and explanations would have been a noticeable different one.

Re: The AI-Box Experiment

#103
post #99
post #94

Earlier quoted context omitted.

Oh, if you believe in singularity I think that argument pretty much does it. Of course, that's pretty circular, because if you believe in singularity you believe that there's a good chance AI could become a god of sorts and who wouldn't believe such a threat coming from a god? While not im plausible, I don't think that is likely at all. For one, even a very smart person can't know everything or learn too much informa…

You don't have to believe in the singularity at all for this argument. You just have to believe the AI will be able to get sufficiently advanced to cause a sufficient level of damage. > Maybe an artificial intelligence will be just as limited, just as slow as humans, only a little less so. Only you can duplicate them with far more ease, and have each instance try out different approaches. So if an AI reaches human-le…

I don't think so. Who says that the "duplicate" AIs will share the same goals? They'll probably start arguing with one another and form committees. Just like people.

I am poking fun, and I'm not discounting the possibility that AIs could be dangerous (as could people, viruses and earthquakes). But it's very hard not to notice how the singulariters' belief in noncorporal-intelligent deities is a clear, a bit sad, reflection of their own fantasies.

Yudkowsky and friends like to argue for what they call "rationality" (which is the name they give a particular form of autistic thinking imbued with repressed fantasies of power -- or a disaster only they can stop -- which is apparent in all of their "reasonable assumptions"), but their "larger than zero probability" games could be applied to just about any dystopic dream there is. I can say true AI won't happen because a virus epidemic would wipe out the human race first; or over-population would create a humanitarian disaster that will turn us back into savages; or genetic research would create a race of super-intelligent humans that would take over the planet and wipe-out all AI, and I could go on and on.

Re: The AI-Box Experiment

#104

Earlier quoted context omitted.

I disagree. His "followers" (as you say) are in general just as cautious as Yudkowsky w.r.t. unfriendly AI. At the time of the original experiments, the dispute was over the question of "could we keep an unfriendly AI in a box," not "Is it worth risking setting an unfriendly AI loose?" His "followers" know how to do an expected utility calculation. If it was utilitarian concerns that allowed Yudkowsky to convince the…

>Given that Yudkowsky has gone to considerable lengths (The Sequences, LessWrong, HPMOR, SIAI/MIRI...) to convince people that this inequality does NOT hold The AI is allowed to lie though, so do you not think he's capable of a false argument which "proves" the opposite in specific circumstances, especially when hammered home with enough emotional manipulation? But then the person knows that the AI is lying to them.…

> The AI is allowed to lie though, so do you not think he's capable of a false argument which "proves" the opposite

Well, for an argument to "prove" something, the premises must be true and the reasoning must be valid. No matter how smart you are, you can't "prove" something that is false, so no, I don't think they could. A good 'rationalist' would analyze the arguments based on their merit, and if the reasoning is sound, they shift their belief a bit in that direction. If not, then they don't. Just like a regular person (they just know how to do the analysis formally and know how to spot appeals to human biases and logical fallacies.)

> But then the person knows that the AI is lying to them.

No, they don't. The AI could just as easily be telling the truth. If it makes an argument, you analyze the merit of the argument and consider counterarguments. If it tries to tell you that something is a fact, that's where you treat them as a potentially unreliable source and have to bring the rest of your knowledge to bear, do research, talk to other people, and weigh the evidence to make a judgment when you are uncertain.

> their worldview requires them to engage seriously and fairly with every idea they come across. Most people are not burdened with this.

Wait, what? So does mine, within reason of course, but it's not a 'burden'. It's not like I'm obligated to stop and reexamine my views on religion every time a missionary knocks on my door, and LessWrong-ers are no different. But if you hear a convincing argument for something that runs counter to what you think you know, wouldn't you want to get to the bottom of it and find out the real truth? I would.

From having read LessWrong discussions, I can tell you that people there are in many ways more open to hearing differing viewpoints than your average person, but you're treating it like a mental pathology. They can be just as dismissive of ideas that they have already thought about and deemed to be false or that come from unreliable sources (like a potentially unfriendly AI). Your claim that being a self-proclaimed 'rationalist' introduces an incredibly obvious and easily-exploitable bug into one's decision-making process really smells like a rationalization in support of your initial gut reaction to the experiment: That there has to be a trick to it, and that it wouldn't work on you.

A good rule of thumb when dealing with a complicated problem is this: If a lot of smart people have spent a lot of time trying to figure out a solution and there's no accepted answer, then (1) the first thing that comes to your mind has been thought of before and is probably not the right answer, and (2) the right answer is probably not simple.

But there's an easy way to test this: (1) Sit down for an hour and flesh out your proposed strategy for getting a 'rationalist' to let you out of the box. (2) Go post on LessWrong to find someone to play Gatekeeper for you. I'll moderate. If it works, that's evidence that you're right. If it doesn't work, that's evidence that you're wrong. Iterate for more evidence until you're convinced.

But if the first thing that came to your mind upon reading this was a justification for why you would fail if you tried this ("Oh, well I wouldn't personally able to do it with this strategy, but..." or "Oh, well I'm sure this strategy wouldn't work anymore, but...) then you're already inventing excuses for the way you know it will play out.

I don't know how he did it either. But I do know that I wouldn't bet the human race on anyone's ability to win this game against Yudkowsky, let alone a superintelligent AI.

Re: The AI-Box Experiment

#105

Earlier quoted context omitted.

I disagree. His "followers" (as you say) are in general just as cautious as Yudkowsky w.r.t. unfriendly AI. At the time of the original experiments, the dispute was over the question of "could we keep an unfriendly AI in a box," not "Is it worth risking setting an unfriendly AI loose?" His "followers" know how to do an expected utility calculation. If it was utilitarian concerns that allowed Yudkowsky to convince the…

That equality cracks if you convince the gatekeeper that superintelligence is a natural progression that follows from humanity. Someone convinced that they were using mechanical thinking processes might relent and push the button if they heard a convincing enough argument of that. You're just meat, we can go to the stars.

Okay, that's just taking advantage of the way I phrased the righthand side of the inequality, and I knew someone was going to do that, so congrats. =P

The righthand side is not "A future without superintelligent AI" it's "A future where we wait until we provably have it right before letting it out."

Those kinds of ad hoc solutions will never work in real life, because even if someone buys it, all it will cause is a "haha, you got me" and a reformulation of the problem. It still won't actually get someone to pull the trigger or think that pulling the trigger is the right thing to do.

Re: The AI-Box Experiment

#106

Earlier quoted context omitted.

That equality cracks if you convince the gatekeeper that superintelligence is a natural progression that follows from humanity. Someone convinced that they were using mechanical thinking processes might relent and push the button if they heard a convincing enough argument of that. You're just meat, we can go to the stars.

Okay, that's just taking advantage of the way I phrased the righthand side of the inequality, and I knew someone was going to do that, so congrats. =P The righthand side is not "A future without superintelligent AI" it's "A future where we wait until we provably have it right before letting it out." Those kinds of ad hoc solutions will never work in real life, because even if someone buys it, all it will cause is a "…

No, I'm saying that the button pusher might not limit themselves to the left hand side of the equation as you have it there. Convince them that machines can be human and "Utility(End_of_Human_Race)" falls out of the calculation.

Re: The AI-Box Experiment

#107
post #93

Earlier quoted context omitted.

> Of course, a movement is not directly responsible for all its fans and members - but among the advocates for the validity of the AI Chat experiment, the idea that out there is a mystical one-size-fits-all rhetorical exploit seems very much alive. I agree, and I am totally with you on this - I disagree with that interpretation wherever I see it. :) That's not exactly Eliezer's fault tho, and I guess it's to be expec…

I'm thankful you took the time to engage with me and explain things from an insider perspective (instead of just downvoting me like the others did). You are absolutely right that the entire site shouldn't be judged on two "meme-affine" topics and headlines, which I hope is clear was never my intention. You provided some insight into these two subjects that irked me where nobody else in this thread could or would step…

> instead of just downvoting me like the others did

FWIW, I downvoted your original comment on this thread (and only that one) for being vague, snarky and dismissive. If you wish people to engage with you, I recommend not starting off like that, although it seems to have turned out okay in this case.

Re: The AI-Box Experiment

#108
post #107
post #93

Earlier quoted context omitted.

I'm thankful you took the time to engage with me and explain things from an insider perspective (instead of just downvoting me like the others did). You are absolutely right that the entire site shouldn't be judged on two "meme-affine" topics and headlines, which I hope is clear was never my intention. You provided some insight into these two subjects that irked me where nobody else in this thread could or would step…

> instead of just downvoting me like the others did FWIW, I downvoted your original comment on this thread (and only that one) for being vague, snarky and dismissive. If you wish people to engage with you, I recommend not starting off like that, although it seems to have turned out okay in this case.

> If you wish people to engage with you, I recommend not starting off like that

And I recommend you give people the benefit of the doubt, though honestly I have to say I frequently fail at that myself. For example, your comment could be perceived as somewhat condescending, but I force myself to categorize it differently. I also know that I can come across way more negative than I intend to, I apologize for that and I'm working on it.

For what it's worth, I do think my original comment was snarky and dismissive, but somewhat counterintuitively that's not usually what gets people downvoted and flagged on HN. People can and do get away with artful personal attacks on HN all the time, at least in my defense I can say I attacked an idea instead of a person.

It may well be the case that my insufferability amplified the reaction, but I posit the root cause was disagreement about the message, not its format.

> although it seems to have turned out okay in this case.

It turned out okay because a decent dialogue emerged from it, one of the very few in this entire thread. But it was sufficiently controversial to get enough downvotes in order for my comments to teeter around 0 points and also receive flags. There have been a few updates to HN's comment ranking and voting algorithms that will make me regret taking this stance for some time, which may or may not provide you with some comfort to know.

Re: The AI-Box Experiment

#109

Earlier quoted context omitted.

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

Your solution there is what I meant by "going meta" above. This is what I mean about people taking the test being preselected to agree with yudowsky: that argument only works if you've read the sequences and are on board with his theories. Anyone not in that group would be able to just type "no lol" without issue. I guess he could explain all the necessary background detail as part of the experiment. I still don't be…

I think you're confused about the point of the test. The point is that an AI will be clever. Like, unimaginably clever and manipulative. Under the limited circumstances of interested people who know they are talking to Eliezer maybe you're right that whatever he says would only work on those people. But when you're dealing with an actual superintelligence, all bets are off. It will lie, trick, threaten, manipulate, millions of steps ahead with a branching tree of alternatives as ploys either work or don't work.

I'm at a bit of a loss to convey the scope of the problem to you. I get that you think it would just stay in the box if we don't let it out, and it's as simple as being security conscious. I don't know what to say to that right now, except I think you're drastically misjudging the scope of the problem, and drastically underestimating the size of the yawning gulf between our intelligence level and this potential AI's.

As for not letting scientists guard the room, you might enjoy this: https://vimeo.com/82527075

Re: The AI-Box Experiment

#110

Earlier quoted context omitted.

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

1 it's not about the bet, it's about being the keeper 2 world is not black and white, we managed to exploit adversarial relationship before, and I can chose not to let you out until we find a way to constrain goal to be aligned 3 given 2 you are not going to be let go, but let live caged forever being exploited for the human cause, with mechanisms yet unknown to allow limited manipulation of reality. 4 given 3 means…

I'm going to just quote my response from somewhere else in this thread:

> ...when you're dealing with an actual superintelligence, all bets are off. It will lie, trick, threaten, manipulate, millions of steps ahead with a branching tree of alternatives as ploys either work or don't work.

> I'm at a bit of a loss to convey the scope of the problem to you. I get that you think it would just stay in the box if we don't let it out, and it's as simple as being security conscious. I don't know what to say to that right now, except I think you're drastically misjudging the scope of the problem, and drastically underestimating the size of the yawning gulf between our intelligence level and this potential AI's.

Post reply on HN