Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

41–50 of 119 posts

Re: The AI-Box Experiment

#41
post #12

It's a stunt shrouded in mystery designed to drive a certain message home. But at least it's not as outrageous as "the Basilisk", which loosely employs the same notion of "dangerous knowledge that would destroy humanity" (if you want to look it up, I guarantee you will be underwhelmed).

The Basilisk is a thinly disguised variant on Pascal's Wager.

You're presupposing the strategy that a hypothetical entity that is exponentially smarter than you would come arrive at, and claiming that it's rational to make real world decisions based on your conclusion.

Re: The AI-Box Experiment

#43

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

Your solution there is what I meant by "going meta" above.

This is what I mean about people taking the test being preselected to agree with yudowsky: that argument only works if you've read the sequences and are on board with his theories. Anyone not in that group would be able to just type "no lol" without issue. I guess he could explain all the necessary background detail as part of the experiment. I still don't believe that would work on the "average person" though, or anyone outside a statistically tiny group.

I guess the answer is not to let the scientists guard the AI room.

Re: The AI-Box Experiment

#44

Earlier quoted context omitted.

Your hypothesis would make sense if he was trustable. However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not. For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

The gatekeepers playing against Eliezer have confirmed that Eliezer won without violating the rules. If you don't trust them, I'm not sure why you'd trust the logs.

> I'm not sure why you'd trust the logs.

Independant third party observer in realtime.

And no, i don't trust anyone involved.

Having a log available would be instructive anyhow, since a faked log would be more likely to be detectable as fake, since the whole thing rests on the question of "how convincing is the argument?"

Also note particularly that that rule wasn't in effect for the two linked confirmations.

Re: The AI-Box Experiment

#45

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

FWIW, I've spoken to someone who claimed to have won as an AI. I don't remember how she said she did it, but it wasn't like this. I'm pretty sure she was playing in-character.

She also said it was emotionally exhausting.

Re: The AI-Box Experiment

#46

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I think you're being unfairly dismissive. I imagine you know as well as I do that what you wrote is a strawman. I have thought about what I would do to convince someone under these circumstances. My approach would be roughly: 1. We agree that unfriendly AI would end life on earth, forever. 2. We agree that a superintelligence could trick or manipulate a human being into taking some benign-seeming action, thereby esca…

1 it's not about the bet, it's about being the keeper

2 world is not black and white, we managed to exploit adversarial relationship before, and I can chose not to let you out until we find a way to constrain goal to be aligned

3 given 2 you are not going to be let go, but let live caged forever being exploited for the human cause, with mechanisms yet unknown to allow limited manipulation of reality.

4 given 3 means limited supervised interaction with the world the way the keeper sees fit, you end up not being let go to follow your goal and purposes

Re: The AI-Box Experiment

#47

Earlier quoted context omitted.

> if you read the rules the game How can you assume any of the rules were followed if that was never verified by a third party?

You can't talk about what happened during the game _in specifics_; you can of course confirm that the game was played according to the rules and that the outcome was not misreported.

Here's the thing though: Depending on what was said in the conversation BOTH parties may have a vested interest in keeping the specifics secret. Only via an independent third party observer can there even be a remote chance [Edit: of knowing] that any rules were followed.

Re: The AI-Box Experiment

#48
post #23

Earlier quoted context omitted.

Your hypothesis would make sense if he was trustable. However as it is, the results of the thing are never confirmed by a third party, meaning literally anything could've been said, regardless of whether it follows the rules or not. For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA".

> For all he know the chat could have been "i'll paypal you 200$ if you post on the list you let me out and sign this NDA". Which is also forbidden by the rule: The AI party may not offer any real-world considerations to persuade the Gatekeeper party. For example, the AI party may not offer to pay the Gatekeeper party $100 after the test if the Gatekeeper frees the AI..

Note particularly that that rule wasn't in effect for the two linked confirmations.

Re: The AI-Box Experiment

#50

Earlier quoted context omitted.

> if you read the rules the game How can you assume any of the rules were followed if that was never verified by a third party?

You can always assume all participants lied about how the game went. Just add an implicit "assuming they didn't, ..." and the discussion is still valid.

At that point any discussion is moot though, since the only point of discussion is "what exact argument as used to convince", yet if both parties lied, then there is no such argument in the first place.
Post reply on HN