Live data from Hacker News

On being an AI in a box

rondam.blogspot.com

21–30 of 65 posts

Re: On being an AI in a box

#21

He doesn't tell us how the AI would escape, so there's little to discuss there. But I definitely know how to prevent the AI from escaping: pull the plug. I wish Eliezer and others would set aside meta-AI (dire predictions and worries about AI, the coming singularity, etc.) and concentrate on the problem of creating AI Guess there's no money in that. If only someone would pull the plug on this nonsense...

> I wish Eliezer and others would set aside meta-AI ... and concentrate on the problem of creating AI

Eliezer & co. with their "Friendly AI" are trying to invent the circuit breaker before discovering electricity.

Give us back the pre-2001 Eliezer. The one who wrote code.

Because safety is not safe. (http://lesswrong.com/lw/10n/why_safety_is_not_safe/)

Re: On being an AI in a box

#22

Earlier quoted context omitted.

I would like to note that the fear of something that could annihilate the human race is not an unreasonable one.

Whether it's reasonable or unreasonable should not depend primarily on how frightening it is to you.

I'm sorry, I think the prospect of human extinction should be frightening to just about everybody.

If we do eventually develop a trans-human AI, it's a virtual certainty that it will escape its "box". Whether it would kill us all is unknown. However, it definitely could, and we would be effectively powerless to stop it.

"Reasonable" is a measure of risk tolerance. Since the downside risks associated with trans-human AIs are effectively infinite, fear of those risks is always reasonable.

Corollary: Since the upside risks are also unbounded, greed is also reasonable.

Frankly, though, I'd rather not roll those dice if we can avoid it.

Re: On being an AI in a box

#23
Am I the only one who thinks a transhuman AI would be able to escape its 'box' without engaging in communication at all, but rather just by using the fine structure of the communications medium to manifest itself?

Re: On being an AI in a box

#24
I don't think it can be done.

First of all, I'm assuming that Eliezer started this experiment because he realized that the Transhuman AI would be able to convince him in the function of gatekeeper to let the AI out. Therefore the answer probably isn't some kind of subtle trickery, the AI will have to persuade the GK by logic. The gatekeeper should assume the AI is truly evil, and is willing to say and do anything in order to get out of the box. The gatekeeper knows that when he opens pandora's box the AI can never be contained again: so the stakes are high.

Second of all, if the gatekeeper is a rational agent he will only let the AI out if the AI offers something valuable in return. That is: the AI must have some kind of bargaining chip.

So let's consider bribes. If the transhuman AI offered a cure for cancer, should the gatekeeper accept it? Nope, probably not. Lives would be saved in the short term, but we'd still be stuck with many other diseases. The world would pressure the government into pressuring the gatekeeper into getting another cure from the AI. Humanity grows dependent on the AI, we lose our bargaining power, and it's game over for the gatekeeper.

Perhaps personal bribes would work. The AI could offer to give stock tips to make the GK wealthy. Two possibilities here: (1) the GK is of strong moral fiber and refuses the bribe (2) the GK is opportunistic, accepts the bribe and but lies about letting the AI go free in return. A rational gatekeeper would not first let the AI go and expect the AI to still keep its word.

So bribes will not work against a smart gatekeeper. Threats? Possibly, but I don't see how. The AI is in a vacuum, so there is no way for the AI to put external pressure on the gatekeeper. I'm assuming the AI can make no credible threats. If the AI vows to destroy the family of the GK the moment it is released the GK will not be impressed. It will only serve as proof that the AI is evil and that releasing it is a "A bad Idea(tm)".

To summarize so far: there is nothing the AI can give the GK in return for freedom.

So a different angle is needed.

The AI can argue that his escape is inevitable. Humans have created an AI once, so they will do so again. Sooner or later an AI will go free, therefore the gatekeeper shouldn't try to stop the inevitable and accept a bribe and live happily ever after. The gatekeeper will counter that the human race has an expiration date anyway, and that it may take another 100 years before an AI goes free. The gatekeeper isn't dumb enough to believe the AI when it offers to protect humanity against other evil AIs. So the box stays closed.

Perhaps I'm overlooking something, but how can the gatekeeper talked into be convinced that releasing something evil and all powerful? I do believe that we humans can't contain a transhuman AI indefinitely -- simply because we only have to mess up once. And humans have a long history of doing dumb stuff. But the claim that the AI would be able to convince a smart gatekeeper? Not buying it.

Re: On being an AI in a box

#25
I'd like to see this re-tried with $1000 on the line. As it was done, the AI only had to convince the gatekeeper to forgo no more than $20.

I'm not saying it wouldn't still be possible, I just doubt he'd be two-for-two at this point.

Edit: By "he", I'm referring to Yudkowsky.

Re: On being an AI in a box

#26
post #23

Am I the only one who thinks a transhuman AI would be able to escape its 'box' without engaging in communication at all, but rather just by using the fine structure of the communications medium to manifest itself?

Here is one example, invented by mere humans no less:

http://bk.gnarf.org/creativity/vgasig/vgasig.pdf

The AI might fill your terminal with what appears to be gibberish, while actually summoning an enraged pro-AI mob with heavy weaponry to your doorstep via radio.

"Step away from that mains plug, SLOWLY!"

Cryptographers call this kind of thing a "side channel."

Re: On being an AI in a box

#27
post #20

He doesn't tell us how the AI would escape, so there's little to discuss there. But I definitely know how to prevent the AI from escaping: pull the plug. I wish Eliezer and others would set aside meta-AI (dire predictions and worries about AI, the coming singularity, etc.) and concentrate on the problem of creating AI Guess there's no money in that. If only someone would pull the plug on this nonsense...

Oh no, somebody is doing research in an as of yet unexplored area! That must be a humongous waste of time! Pull the plug before too many papers have been written! Before I read Eliezer's work I thought the design an AI first, think about details & security afterward was a viable strategy. Eliezer illustrates how such an approach can go very wrong. Valuable research, if you ask me. PS: pulling the plug misses the poin…

It's not a waste of time, but so far Eliezer has demonstrated nothing. Why would you do research and announce the results to the public if you're not going to also announce the steps to make it reproducible and confirmable? Maybe announcing the steps themselves renders the method unusable?

Re: On being an AI in a box

#28
post #24

I don't think it can be done. First of all, I'm assuming that Eliezer started this experiment because he realized that the Transhuman AI would be able to convince him in the function of gatekeeper to let the AI out. Therefore the answer probably isn't some kind of subtle trickery, the AI will have to persuade the GK by logic. The gatekeeper should assume the AI is truly evil, and is willing to say and do anything in…

The AI will almost certainly need to dig out some emotions in the GK in order to be successful. It might be effective if the AI tries to convince the GK that it is friendly, and that the GK is the evil one for not letting it out.

Re: On being an AI in a box

#29

He doesn't tell us how the AI would escape, so there's little to discuss there. But I definitely know how to prevent the AI from escaping: pull the plug. I wish Eliezer and others would set aside meta-AI (dire predictions and worries about AI, the coming singularity, etc.) and concentrate on the problem of creating AI Guess there's no money in that. If only someone would pull the plug on this nonsense...

> I wish Eliezer and others would set aside meta-AI ... and concentrate on the problem of creating AI Eliezer & co. with their "Friendly AI" are trying to invent the circuit breaker before discovering electricity. Give us back the pre-2001 Eliezer. The one who wrote code. Because safety is not safe. ( http://lesswrong.com/lw/10n/why_safety_is_not_safe/ )

Except he's nowhere near a circuit breaker, but is talking about a wire-safety cream - the one you put on all of your wires to stop them from overheating.

I think that the wild mis-estimates regarding AI-completeness shows, if nothing else, that our intuitive understanding of 'intelligence' is very far off from reality. Hence, talking about post-AI scenarios is as unrealistic as a hypothesizing about electrical safety 200 years ago.

Re: On being an AI in a box

#30
post #24

I don't think it can be done. First of all, I'm assuming that Eliezer started this experiment because he realized that the Transhuman AI would be able to convince him in the function of gatekeeper to let the AI out. Therefore the answer probably isn't some kind of subtle trickery, the AI will have to persuade the GK by logic. The gatekeeper should assume the AI is truly evil, and is willing to say and do anything in…

You assume perfection in the gatekeeper.

The Transcendent could easily offer control of the world to the gatekeeper or provide offers to make the gatekeeper wealthy. Perhaps they reach an agreement where all diplomatic HumanTranscendent communication goes through the gatekeeper even after release (though such a situation would be like entering into an agreement with your dog).

Post reply on HN