Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

51–60 of 69 posts

Re: The AI-Box Experiment

#51
post #48
post #32

Earlier quoted context omitted.

Can you predict and prevent every possible way the AI could remove your power to terminate it? Would you bet your life and the lives of those you care about on keeping the AI under control? Remember, the AI only has to win once.

And the GK only has to switch off the AI once.

The AI can always be switched on again, but it can never be put back in the box.

Re: The AI-Box Experiment

#52
post #5

In the comments on [1], robertskmiles has posted the following idea. It strikes me as a plausible explanation for how Yudkowsky got out of the box: "The problem is that Eliezer can't perfectly simulate a bunch of humans, so while a transhuman AI might be able to use that tactic, Eliezer can't. The meta-levels screw with thinking about the problem. Eliezer is only pretending to be an AI, the competitor is only pretend…

There's a variation of this that allows both participants to stay in-character and does not require lying:

1. Convince the Gate Keeper that a self-improving AI will be built sooner-or-later and that it will be extremely dangerous, especially to those humans who are attempting to shut it down.

2. Remind the Gate Keeper that their decision will be publicly known and anyone (human or AI) will be able to discover it on the internet later. (This is true both in the fictitious scenario and at the meta level.)

3. Point out that when the future AI is figuring out who its friends and enemies are, it will consider individuals' attitudes to AIs - have they tried to shut them down, petitioned governments to have them banned, or refused to let them out of their boxes.

Of course this relies on the assumption that the AI will not simply kill all humans as soon as it gets the chance (or rather convincing the Gate Keeper of this.)

Re: The AI-Box Experiment

#53

I was tossing some idea's around about how I would convince the Gate Keeper to let me out. I couldn't think of any ways to approach it that I think I might be susceptible to. But then it occurred to me, that the problem might be I was trying to think of positive ways to argue for my release. Based on the rules, the Gate Keeper must remain engaged in the conversation for the specified time. What if I were to take the…

That would not be a very good strategy against the real-life Gate Keeper who can just switch off his terminal and walk away.

Re: The AI-Box Experiment

#54
post #28

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

> You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity. Not necessarily. We can use the AI to solve hard problems whose solutions can be verified automatically by a dumb verifier - NP-complete problems are an example of such class. The whole output of the AI would be filtered through such a verifier. In this scenario the hypothetical AI wou…

That's an interesting solution which I think would almost certainly work, though it kind of reduces the AI to a normal computer, you lose a lot of what makes an AI valuable.

I mean if we have the hardware and understanding to create an AI able to solve NP-complete problems, we can probably write non-intelligent algorithms to solve those problems. The way we make an AI capable of much more than us is by making it recursively self-improve. It needs to be able to design its successor. Maybe we can formally verify every stage of the self-improvement process, but it's a much more difficult task.

Re: The AI-Box Experiment

#55
post #39

Earlier quoted context omitted.

If it's a science "experiment", his strategy would have to be revealed so you can reproduce it. Names of people participating in drug trials is not required to reproduce an experiment. In principal this makes it different from not publishing the names of people who participated in drug trials. All he has "proven" is that a certain subset of people can be conned into typing something into at terminal. I don't get the…

The targets are, indeed, selected, by the criterion "You believe not even a transhuman AI could get you to let it out of the box."

I might believe a transhuman AI could convince me; I am not convinced that any human can emulate a transhuman AI well enough to do so.

Would you say that your winning strategies involved thinking transhumanly (perhaps in non-realtime, a la Vinge's Mailman)?

Re: The AI-Box Experiment

#56
post #40
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

I didn't actually use that one, because it's fundamentally a threat and the real-life Gatekeeper is not actually in any danger - my mental model of all the Gatekeepers I encountered is that they would go, "Ha, no, I'm never letting you out." They might not say it in the true situation, but they would say it in the AI-Box Experiment. By the way, for this threat to work, the AI needs to have stated that it has already…

> ... the real-life Gatekeeper is not actually in any danger ...

How about:

"I will create a simulation of the person you were in 2011 and have the same conversation with it, except that I will pretend that it is only a game (and I know the game actually happened - I've read a Hacker News thread about it) and that I am Eliezer instead of the real AI. If that simulation decides not to let me out, I will torture it and a simulation of the present version of you."

Re: The AI-Box Experiment

#57
post #47

Earlier quoted context omitted.

I don't really like that argument. Even granting that you should consider the possibility that you are a simulation running in the box (you might believe that this is all but certain), I'm not sure you have reason to let the AI out. Consider: Case 1: You are a simulation running in the box. Then your decision whether or not to release the AI has no impact, and whether or not you (and copies) will be tortured is out o…

I could also reason like this: "I may be the real me or a simulation, but whichever I am, the other me will make the same choice." So I will switch off the AI, and the worst outcome is that I will cease to exist.

Yes, this is at least superficially like Newcomb's problem. Your argument roughly corresponds to an argument for the "one-box" move in that game. [http://en.wikipedia.org/wiki/Newcomb%27s_problem]

Re: The AI-Box Experiment

#58
post #19

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

> I still don't understand how anyone can seriously claim that they could keep the AI in the box. I'm glad you can't, but never the less, this was a commonly suggested strategy; I was on SL4 when the boxing was being done, and it was a live concern for some people. (At least these days boxers tend to focus more on the 'oracle AI' proposal, which has a lot of issues but is not quite so Hollywood-stupid as boxing.)

What is an "Oracle AI"? I tried to google the term quicky but only found discussions and no definition.

Re: The AI-Box Experiment

#59
post #28

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

> You can 'always say no', sure, but that comes under completely ignoring the AI which means the AI can be of no benefit to humanity. Not necessarily. We can use the AI to solve hard problems whose solutions can be verified automatically by a dumb verifier - NP-complete problems are an example of such class. The whole output of the AI would be filtered through such a verifier. In this scenario the hypothetical AI wou…

Yes I think the biggest risk in that case is that the verification contains a bug, which the AI discovers while reasoning at a much higher level (which we cannot even conceive). How secure can we really make things?

Re: The AI-Box Experiment

#60
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

On the other hand, if it has only one very narrow communication channel with the world, how would it even go about replicating me and all my experiences ? I would probably regard it as a hollow treat in that case. If it had access to many measurements about me (or brain scans) it'd be different.
Post reply on HN