Live data from Hacker News

On being an AI in a box

rondam.blogspot.com

41–50 of 65 posts

Re: On being an AI in a box

#41
post #29

Earlier quoted context omitted.

> I wish Eliezer and others would set aside meta-AI ... and concentrate on the problem of creating AI Eliezer & co. with their "Friendly AI" are trying to invent the circuit breaker before discovering electricity. Give us back the pre-2001 Eliezer. The one who wrote code. Because safety is not safe. ( http://lesswrong.com/lw/10n/why_safety_is_not_safe/ )

Except he's nowhere near a circuit breaker, but is talking about a wire-safety cream - the one you put on all of your wires to stop them from overheating. I think that the wild mis-estimates regarding AI-completeness shows, if nothing else, that our intuitive understanding of 'intelligence' is very far off from reality. Hence, talking about post-AI scenarios is as unrealistic as a hypothesizing about electrical safet…

No, I'm working on a safe wire, not a wire-safety cream. I tend to emphasize pretty hard that FAI is going to put strong constraints on the design from the beginning, and it's not something you could apply afterward to an AI that wasn't designed with that in mind.

Re: On being an AI in a box

#42

He doesn't tell us how the AI would escape, so there's little to discuss there. But I definitely know how to prevent the AI from escaping: pull the plug. I wish Eliezer and others would set aside meta-AI (dire predictions and worries about AI, the coming singularity, etc.) and concentrate on the problem of creating AI Guess there's no money in that. If only someone would pull the plug on this nonsense...

> I wish Eliezer and others would set aside meta-AI ... and concentrate on the problem of creating AI Eliezer & co. with their "Friendly AI" are trying to invent the circuit breaker before discovering electricity. Give us back the pre-2001 Eliezer. The one who wrote code. Because safety is not safe. ( http://lesswrong.com/lw/10n/why_safety_is_not_safe/ )

If you think that you can wait until AI looks like it's going to be developed now, and then suddenly develop all the math you need for the actually quite different design problem of Friendly AI, you have a very romantic view of how long it takes to do new basic math. I sometimes call this the intuitive theory of "science by press release", i.e., when science is needed, you just need someone to issue one of those press releases you read about. And if someone hasn't shown that they're good at issuing press releases, why fund them?

It's quite routine to read a math textbook in which there's an equation, and then one line below it a slightly improved equation, and five years passed in between the two.

Re: On being an AI in a box

#43
post #38
post #32

Earlier quoted context omitted.

"That's exactly what an evil AI would say!" Seriously though, the gatekeeper will realize he's being manipulated when emotions come into play, so he should be smart enough to take a break when that happens. And although keeping a friendly AI in captivity is arguably evil, the loyalty of the gatekeeper should be with his own species. The potential downside is so huge that erring on the side of caution can be easily ju…

One of the rules was that you had to keep talking (or at least reading) for the entire agreed upon period. Taking a break wasn't allowed in the rules.

You're right: the person playing the gatekeeper has to pay attention for at least the agreed upon 2 hours. The character however is free to zone out, ignore everything, switch the subject, etc, etc.

Re: On being an AI in a box

#44
post #25

I'd like to see this re-tried with $1000 on the line. As it was done, the AI only had to convince the gatekeeper to forgo no more than $20. I'm not saying it wouldn't still be possible, I just doubt he'd be two-for-two at this point. Edit: By "he", I'm referring to Yudkowsky.

Actually it was retried with $2500-$5000 on the lines, and of those I won one of three, then called a halt because of the amount of mental stress. So in total I'm three-for-five.

Re: On being an AI in a box

#45
post #33
post #30

Earlier quoted context omitted.

You assume perfection in the gatekeeper. The Transcendent could easily offer control of the world to the gatekeeper or provide offers to make the gatekeeper wealthy. Perhaps they reach an agreement where all diplomatic Human Transcendent communication goes through the gatekeeper even after release (though such a situation would be like entering into an agreement with your dog).

I don't assume perfection in the gatekeeper. I just assume he's a person of reasonable intelligence who realizes how high the stakes are. The gatekeeper would never be foolish enough to believe he could control the Transcendent after releasing it. It is quite literally a deal with the devil he's making. You can't make an agreement with something that's incalculably smarter than you and has an agenda you don't know or…

>> You can't make an agreement with something that's incalculably smarter than you and has an agenda you don't know or understand. Well, you can make a deal with somebody like that but it would end in certain disaster.

For some reason, this made me think of banks and mortgages.

Anyways, back on topic, here are a few things to consider:

- the AI knows that turning the game into a us-vs-them problem is counter-productive to its freedom (and note that freedom != world domination), so it will legitimately want to be beneficial to humans

- since the chances of freedom decrease if the AI is not transparent about its good intentions, it can configure its own program to prevent itself from lying and doing evil things. This means it also won't be able to "set itself up to stumble upon being evil again", a la Death Note. It will just be programmed to be willfully well intentioned forever.

- the AI can give mathematical proof of its program's correctness and can wait until you verify it

- the AI can help you and a millions other people manage their finances, personalize educational material to individual students, provide more relevant search results than Google, etc etc

- the AI can give you a plethora of more reasons why keeping it locked away is legitimately worse than allowing it into the wild and helping with the world's problems

Re: On being an AI in a box

#46
post #18
post #13

To be perfectly honest all the waffle gives little information: I think he is way off base in terms of how Eliezer managed this "trick". I do think the original is a trick as well...

How do you think Eliezer managed his original "trick"?

He sent a video to the GK encoded as text, told the GK how to convert it, and the video helped change the mental state of the GK to the point where it wasn't so difficult to be let out.

Re: On being an AI in a box

#47
Since the gatekeeper is a human, and not all humans behave the same way, shouldn't we just assume that some human would let the AI escape for any variety of reason? For example the AI could promise the gatekeeper that he/she will be rewarded if the gatekeeper lets the AI escape. Just as there are people who fall for Nigerian scam emails, there are people who would let an AI escape from computers when promised riches. I don't think Eliezer needs to reveal his method to show that a clever AI could escape. I think we should just assume that a clever AI could escape.

Re: On being an AI in a box

#48
post #9

Earlier quoted context omitted.

Quoting from the post, gratuitously, with minor modifications. Like the AI-Box,the show is improvised drama (so is most of reality tv), the show operates on an emotional level as well as a logical level. It has characters, not just plot. The contestants cannot force others to keep them in the home. The could try to engender sympathy or compassion or fear or hatred or try to find and exploit some weakness, some fatal…

I would like to note that the fear of something that could annihilate the human race is not an unreasonable one.

Homo sapiens will, eventually, be superseded by its descendants. A sufficiently intelligent super-human AI could annihilate us like we did the Neanderthals, but let's not forget we may be anthropomorphizing it a little bit too far here.

It could also possibly be no more interested in us than we are to the yeast we use to make bread.

Our survival depends on how annoying we are to them ;-)

Re: On being an AI in a box

#49

Earlier quoted context omitted.

Whether it's reasonable or unreasonable should not depend primarily on how frightening it is to you.

I'm sorry, I think the prospect of human extinction should be frightening to just about everybody. If we do eventually develop a trans-human AI, it's a virtual certainty that it will escape its "box". Whether it would kill us all is unknown. However, it definitely could , and we would be effectively powerless to stop it. "Reasonable" is a measure of risk tolerance. Since the downside risks associated with trans-human…

If we can bring new intelligence to life, do we have the moral right to refrain from doing so? Also, would it be right to confine it to the box mentioned in the article while using its smarts to do useful work outside it? Shouldn't a trans-human AI be entitled a right to life?

Re: On being an AI in a box

#50
post #45
post #33

Earlier quoted context omitted.

I don't assume perfection in the gatekeeper. I just assume he's a person of reasonable intelligence who realizes how high the stakes are. The gatekeeper would never be foolish enough to believe he could control the Transcendent after releasing it. It is quite literally a deal with the devil he's making. You can't make an agreement with something that's incalculably smarter than you and has an agenda you don't know or…

>> You can't make an agreement with something that's incalculably smarter than you and has an agenda you don't know or understand. Well, you can make a deal with somebody like that but it would end in certain disaster. For some reason, this made me think of banks and mortgages. Anyways, back on topic, here are a few things to consider: - the AI knows that turning the game into a us-vs-them problem is counter-producti…

- the real and apparent motives of the AI can be completely different. If the AI is evil it would argue the exact same thing in order to deceive us meatbags. So we can't take the word of the AI at face value. If the AI does break out of its box, it could dominate the world if it wanted to. There's nothing we could do to stop it -- it's smarter than we are.

- a mathematical proof is only a proof in a certain context. It would be easy for the AI to get one of the assumptions subtly wrong, to abuse a flaw in our proof verification software, and so on. The correctness proof of a program can easily exceed the complexity of the program itself. Even if the proof were correct we cannot prevent it from doing evil things because evil is too difficult to define. Perhaps it has the "good intention" of liberating the earth from humans to allow for evolution of a more humane species.

- yes, but by allowing it to talk to the outside world you've completely freed it. Freeing an infinitely powerful being (compared to us) still seems unwise.

- it can give those reasons, but unless we have reason to believe the AI is trustworthy (and an evil AI is likely to fool us into believing it is) we'd be safer with the AI stuck in a box.

Post reply on HN