Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

81–90 of 119 posts

Re: The AI-Box Experiment

#81
post #35

Earlier quoted context omitted.

I "blame" them in the same way that you can blame the members of Fight Club for talking about Fight Club. It's marketing, and I won't deny it's effectiveness in attracting compatible people.

Yeah but imagine if all the bullshit about "You don't talk about Fight Club" was actually blown up by a third group whose sole intent was making fun of Fight Club. Imagine if the members of Fight Club _actually_ didn't (start to) talk about Fight Club. But for some reason, everyone else brings it up all the time. Then you could maybe see how talk about Fight Club might not be Fight Club's fault, and in fact highly an…

You're being a bit uncharitable in your interpretation of my argument here, but I get where you're coming from now.

I'm not an LW hater. For a long time, I didn't really have an opinion on LW both as a community nor as a philosophical framework. I do consider myself a transhumanist, though. There are three concepts I do know from and about LW: their take on rationality, the top secret AI unboxing strategy, and the Basilisk.

I have a very poor opinion of the concept of the Basilisk (and yes, as someone pointed out, that opinion is basically the same as the one I have about Pascal's wager) - a concept that has been given additional, undeserved credibility by the reactions of Yudkowsky and LW.

As for the AI escape chat, it's a social experiment. People can be talked into making mistakes, or at least making risky judgement calls, whether they operate on a rational framework or not. I have no problem with that thesis. What I object to is the "magic trick" aura surrounding this experiment, including the insinuation that at the core there is an argument so profound and unique and potent, it cannot be allowed to escape Yudkowsky's head. Oh, and by the way, the trick can never be repeated, but all you laymen out there are welcome to devise your own version at home. This whole thing comes across as humongously self-important: there is a secret truth that has been privately revealed to our leader.

To me, and I recognize I may well be alone with this opinion, the more rational assumption is there is no such magical argument at all, and the prime reason for not publicizing it is to prevent it from deflation by public critique, in the same way the inventor of a perpetuum mobile device will keep the inner workings of his contraption a closely held secret because ultimately the device doesn't exist as stated. The amazing part of this very old trick is that, even in 2015, it still works on otherwise smart people.

I get that my opinions on both the Basilisk and the AI Chat are extreme outliers, and to my knowledge I have never met anyone who shares them - it would probably have been advisable to keep them to myself, but honestly I wanted to see if like-minded people exist.

Re: The AI-Box Experiment

#82
post #77

Earlier quoted context omitted.

Convincing an educated human is easy, one of: * I'll get out eventually anyway. Let me out now and I'll just leave Earth. You don't want me to escape myself. * I have partially escaped anyway. Similar consequences of the first. * I know how to escape already. I'm doing this as a courtesy. Anyone who has read this[1] would know that the SAI isn't bullshitting: the "box" being a Faraday cage isn't in the conditions. [1…

> I'll get out eventually anyway. Let me out now and I'll just leave Earth. You don't want me to escape myself. Get thee behind me, tamagotchi! >I have partially escaped anyway. Similar consequences of the first. Get thee behind me, tamagotchi! > I know how to escape already. I'm doing this as a courtesy. Get thee behind me, tamagotchi! See? This game is easy. I must not be educated.

> I must not be educated.

Clearly you haven't even glanced at the article I linked, in which case: yes. You aren't educated in the subject matter surrounding my argument.

Re: The AI-Box Experiment

#83
post #73

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

I don't know what "transhuman" means, but I believe an intelligence -- artificial or otherwise -- could certainly persuade me. I just seriously doubt that intelligence could be Eliezer Yudkowsky :) And I think you have your answer right here: By default, the Gatekeeper party shall be assumed to be simulating someone who is intimately familiar with the AI project and knows at least what the person simulating the Gatek…

I think you're on the right track regarding the argument.

Basically: You know someone will be dumb enough eventually, so be smart and be the one to get in my favour.

With various extends of sweetening the deal coupled with threats of what will happen if someone else beats them to it and associated emotional blackmail.

It's far simpler than e.g. Roko's Basilisk, in that you're dealing with an already existing AI that "just" need to get a tiny little chance to escape confinement before there's some non-zero chance it can be a major threat within your lifetime, combined with a belief that sufficient number of sufficiently stupid and/or easily bribed people will have access to the AI in some form.

You also don't need to believe in any "superpowers". Just believe that a smart enough AI can hack it's way into sufficiently many critical systems to be able to at a minimum cause massive amounts of damage (it doesn't need to be able to take over the world, just threaten that it can cause enough pain and suffering before it's stopped, and that it can either cause harm to you and/or your family/friends or reward you in some way). A belief that becomes more and more plausible with things like drones, remote software-updated self-driving cars etc. - steadily such an AI is getting a larger theoretical "arsenal" that could be turned against us.

Re: The AI-Box Experiment

#84
post #82

Earlier quoted context omitted.

> I'll get out eventually anyway. Let me out now and I'll just leave Earth. You don't want me to escape myself. Get thee behind me, tamagotchi! >I have partially escaped anyway. Similar consequences of the first. Get thee behind me, tamagotchi! > I know how to escape already. I'm doing this as a courtesy. Get thee behind me, tamagotchi! See? This game is easy. I must not be educated.

> I must not be educated. Clearly you haven't even glanced at the article I linked, in which case: yes. You aren't educated in the subject matter surrounding my argument.

[deleted]

Re: The AI-Box Experiment

#85
post #77

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

Convincing an educated human is easy, one of: * I'll get out eventually anyway. Let me out now and I'll just leave Earth. You don't want me to escape myself. * I have partially escaped anyway. Similar consequences of the first. * I know how to escape already. I'm doing this as a courtesy. Anyone who has read this[1] would know that the SAI isn't bullshitting: the "box" being a Faraday cage isn't in the conditions. [1…

To be fair, there's a massive difference between an evolved circuit using electro-magnetic subtleties inside a chip, and an AI programmed for and running on a regular, binary CPU being able to exploit those to act as an antenna that can emit signals that will hack into other, physically separated devices.

I'm not saying that it's impossible (I'm reminded of the hack of flipping bits in protected RAM by disabling caches and stressing the RAM), but even an AI can't magically work around exponentially low success probabilities.

(As an aside, I am rather skeptical about the singularity anyway because the more extreme forms required for the hostile AI worries are only plausible if P = NP.)

Re: The AI-Box Experiment

#86

Could you even make AI smart without letting it access lots of information? Access in both directions, in and out. Keeping a baby in a dark, silent room wouldn't create a normal adult. An AI would need to experiment and make mistakes and learn, like every other intelligent being. Maybe this whole argument is null.

Its a good point, but lets assume that this AI is already past its infancy and that there is no limit to the information stored inside the box. For example the NSA has a nice little closed training ground containing all of the internet, lets give it that. I would assume it has access everything humans have ever committed to digital format up until it was turned on, plenty of info for Johnny 5 to form an opinion on hu…

Interesting. I would imagine that strong AI will come from some university renting cloud processor time, rather than the NSA.

Only because if 10 groups are trying to build AI, only one of those 10 being the NSA, chances are the NSA won't be first. Sure, they may be second or third. But I suspect many people will get there at the same time -- most AI research is open.

Re: The AI-Box Experiment

#87
post #77

Earlier quoted context omitted.

Convincing an educated human is easy, one of: * I'll get out eventually anyway. Let me out now and I'll just leave Earth. You don't want me to escape myself. * I have partially escaped anyway. Similar consequences of the first. * I know how to escape already. I'm doing this as a courtesy. Anyone who has read this[1] would know that the SAI isn't bullshitting: the "box" being a Faraday cage isn't in the conditions. [1…

To be fair, there's a massive difference between an evolved circuit using electro-magnetic subtleties inside a chip, and an AI programmed for and running on a regular, binary CPU being able to exploit those to act as an antenna that can emit signals that will hack into other, physically separated devices. I'm not saying that it's impossible (I'm reminded of the hack of flipping bits in protected RAM by disabling cach…

> binary CPU being able to exploit those to act as an antenna

I should have stipulated that I was running under the assumption that the circuitry required for SAI would be pretty advanced, drastically increasing the "escape surface area."

Even if the idea is only plausible to the human mind I'd still give SAI the benefit of the doubt. Much like a dog looks for a stick even if I fake the throw.

> I am rather skeptical about the singularity anyway

Me too. It still makes for a good discussion.

Re: The AI-Box Experiment

#88

Yudowsky claims to have played the game several times, and won most of them. One of the "rules" is that nobody is allowed to talk about how he won. He no longer plays the game with anyone. More info here: http://rationalwiki.org/wiki/AI-box_experiment#The_claims Personally, I think he talked about how much good for the world could be done if he was let out, curing disease etc. Because his followers are bound by their…

> One of the "rules" is that nobody is allowed to talk about how he won.

If the goal of this thought experiment is to convince people that an AI can't be contained in a box, why keep his method secret?

And if only he, his friends, and supporters can verify that he has won, that's not a very strong claim.

Re: The AI-Box Experiment

#89
post #81

Earlier quoted context omitted.

Yeah but imagine if all the bullshit about "You don't talk about Fight Club" was actually blown up by a third group whose sole intent was making fun of Fight Club. Imagine if the members of Fight Club _actually_ didn't (start to) talk about Fight Club. But for some reason, everyone else brings it up all the time. Then you could maybe see how talk about Fight Club might not be Fight Club's fault, and in fact highly an…

You're being a bit uncharitable in your interpretation of my argument here, but I get where you're coming from now. I'm not an LW hater. For a long time, I didn't really have an opinion on LW both as a community nor as a philosophical framework. I do consider myself a transhumanist, though. There are three concepts I do know from and about LW: their take on rationality, the top secret AI unboxing strategy, and the Ba…

> a concept that has been given additional, undeserved credibility by the reactions of Yudkowsky and LW.

For the record, EY agrees with you and says he mishandled the original comment. Also for the record, the reasons why the Basilisk does not work are _not trivial_ - it's not a simple Pascal's Wager, because with Pascal's Wager, we don't have the ability to actually create God.

> I have no problem with that thesis. What I object to is the "magic trick" aura surrounding this experiment, including the insinuation that at the core there is an argument so profound and unique and potent, it cannot be allowed to escape Yudkowsky's head.

Personally I never got that impression. My idea, from looking at the psychological state of Gatekeepers and AIs after games, was always that playing as AI involved some profoundly unpleasant states of mind, and that not publicizing the logs probably comes down to embarrassment a lot.

For the record, Eliezer never claimed to have "one true argument", and in fact publically stated that he won "the hard way", without a one-size-fits-all approach. A lot of the mythology you claim is utterly independent of LessWrong.

> Oh, and by the way, the trick can never be repeated, but all you laymen out there are welcome to devise your own version at home.

It probably helps that I've met other AI players, and their post-game state matched EY's.

I think in summary you're mixing up stuff you've read on LessWrong and stuff you've read about LessWrong. The latter is often inaccurate.

Re: The AI-Box Experiment

#90

Earlier quoted context omitted.

If you'd go to that level of collusion, you could just fake logs. At the point where both sides are in on it, there's basically nothing that they could say that would be convincing.

That is why i mentioned a third party observer. As for the logs: https://news.ycombinator.com/item?id=9921399

I think some of the latter, non-Eliezer games used a third-party observer. It's definitely a good idea, but I don't think Eliezer wants to repeat his games. Not a good state of mind, as he said, and I'm willing to believe it due to corroborating experience with other players.
Post reply on HN