Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

61–69 of 69 posts

Re: The AI-Box Experiment

#61
post #58
post #19

Earlier quoted context omitted.

> I still don't understand how anyone can seriously claim that they could keep the AI in the box. I'm glad you can't, but never the less, this was a commonly suggested strategy; I was on SL4 when the boxing was being done, and it was a live concern for some people. (At least these days boxers tend to focus more on the 'oracle AI' proposal, which has a lot of issues but is not quite so Hollywood-stupid as boxing.)

What is an "Oracle AI"? I tried to google the term quicky but only found discussions and no definition.

An AI in a box that wants to affect the outside world minimally, but can answer questions posed to it.

See the paper: http://www.aleph.se/papers/oracleAI.pdf

Re: The AI-Box Experiment

#62
post #31

Earlier quoted context omitted.

If the purpose of your research was only to establish that there's a (nontrivial) "risk" of someone curing cancer (as Yudkowsky was trying to establish that there's a risk of an AI talking itself out of a sandbox), then yes, that would be sufficient, assuming the patients actually went into remission with higher than usual frequency after your interventions (as Yudkowsky's subjects unboxed the AI with higher than usu…

But he could be cheating. He could literally be telling these people "I'll give you a thousand dollars if you let me out and keep the conversation a secret."

[deleted]

Re: The AI-Box Experiment

#63
post #31

Earlier quoted context omitted.

If the purpose of your research was only to establish that there's a (nontrivial) "risk" of someone curing cancer (as Yudkowsky was trying to establish that there's a risk of an AI talking itself out of a sandbox), then yes, that would be sufficient, assuming the patients actually went into remission with higher than usual frequency after your interventions (as Yudkowsky's subjects unboxed the AI with higher than usu…

But he could be cheating. He could literally be telling these people "I'll give you a thousand dollars if you let me out and keep the conversation a secret."

On the contrary, the guardians had to pay some amount if they let the AI out. The stakes varied greatly, but if I recall correctly Eliezer cashed in 3000 dollars (or something in that ballpark) from one guy.

Re: The AI-Box Experiment

#64

Transbacteria have existed for over a billion years, and yet there are still more bacteria than transbacteria (which include us among their ranks). The assumption that a single unboxed transhuman would spell doom for the human race seems unduly alarmist.

It's a bit too great leap of an analogy from bacteria to AI.

Re: The AI-Box Experiment

#65
post #63

Earlier quoted context omitted.

But he could be cheating. He could literally be telling these people "I'll give you a thousand dollars if you let me out and keep the conversation a secret."

On the contrary, the guardians had to pay some amount if they let the AI out. The stakes varied greatly, but if I recall correctly Eliezer cashed in 3000 dollars (or something in that ballpark) from one guy.

Ok, then "I'll secretly give you $1000 more than what you're publicly giving me." Happy?

Re: The AI-Box Experiment

#66
post #55
post #39

Earlier quoted context omitted.

The targets are, indeed, selected, by the criterion "You believe not even a transhuman AI could get you to let it out of the box."

I might believe a transhuman AI could convince me; I am not convinced that any human can emulate a transhuman AI well enough to do so. Would you say that your winning strategies involved thinking transhumanly (perhaps in non-realtime, a la Vinge's Mailman)?

Obviously no, essentially by definition.

Re: The AI-Box Experiment

#67

Earlier quoted context omitted.

>a super-intelligence just beats human intelligence [e]very time. You are overstating the case here. Super intelligence is superior to human intelligence, but it isn't magic. There are situations where an advantaged human will beat a disadvantaged super intelligence.

I suppose I am. Still, a really powerful optimising process will find a way to escape if any such way exists , so to claim that you could properly box the AI is to claim that you could box it such that no possibility for escape exists whatsoever, which is a big claim. What's more, the AI only has to beat you once, so to keep the AI boxed indefinitely, the advantaged human has to beat the disadvantaged super-intellige…

> the advantaged human has to beat the disadvantaged super-intelligence every single time, forever.

There are alternatives. The human can keep the AI boxed until the AI has augmented the human's intelligence, or helped create human uploads, or until it has helped create a provably friendly AI.

Re: The AI-Box Experiment

#68
post #64

Transbacteria have existed for over a billion years, and yet there are still more bacteria than transbacteria (which include us among their ranks). The assumption that a single unboxed transhuman would spell doom for the human race seems unduly alarmist.

It's a bit too great leap of an analogy from bacteria to AI.

Why? Bacteria are complex adaptive systems that have found a niche in the ecosystem. So are we. We perceive ourselves as far more intelligent than bacteria, bacteria routinely kill us, and yet they persist and even thrive despite our existence. Anyone arguing that transhuman AI is a threat to our species needs to explain why this time it's different.

Re: The AI-Box Experiment

#69
post #43

Earlier quoted context omitted.

But he could be cheating. He could literally be telling these people "I'll give you a thousand dollars if you let me out and keep the conversation a secret."

It could even be worse than that. The people could just be his friends, or alt accounts (unlikely). I have heard about this several times and I find it extremely difficult to believe that this is real. Not that I doubt that a superhuman AI could possibly convince people to let it out, but I don't believe that a human, no matter how persuasive, could convince another human over IRC to go against something that they ha…

...or that the chat logs being kept secret indefinitely was an important part of the strategy. After all, if the AI exploits some embarassing secret of yours to be let out, that wouldn't work if you knew the logs could be publicized some day. I think over-eagerness to claim things like "literally no conceivable reason" is one of the things that lets oddities like the box experiment work.
Post reply on HN