Live data from Hacker News

The AI-Box Experiment

yudkowsky.net

31–40 of 69 posts

Re: The AI-Box Experiment

#31
post #17

Earlier quoted context omitted.

He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.

No, it's very different from that. It's more along the lines of demonstrating a drug that cures cancer, but refusing to tell anyone its chemical composition or how to make it.

If the purpose of your research was only to establish that there's a (nontrivial) "risk" of someone curing cancer (as Yudkowsky was trying to establish that there's a risk of an AI talking itself out of a sandbox), then yes, that would be sufficient, assuming the patients actually went into remission with higher than usual frequency after your interventions (as Yudkowsky's subjects unboxed the AI with higher than usual frequency).

Re: The AI-Box Experiment

#32

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

Just because someone is smarter than someone doesn't mean complete power. Alot of people are smarter than their bosses, but you know what the bosses have on their favour? The power to terminate the employee. If I have the power to terminate the AI at any time, as long as I don't give that power up, I will have power over it.

Can you predict and prevent every possible way the AI could remove your power to terminate it?

Would you bet your life and the lives of those you care about on keeping the AI under control?

Remember, the AI only has to win once.

Re: The AI-Box Experiment

#33

I still don't understand how anyone can seriously claim that they could keep the AI in the box. Either your AI has no influence on the outside world (in which case why bother building one since it can't help you from inside the box), or it is able the affect the outside world, in which case it can do what it wants, because it's smarter than you. You can 'always say no', sure, but that comes under completely ignoring…

Just because someone is smarter than someone doesn't mean complete power. Alot of people are smarter than their bosses, but you know what the bosses have on their favour? The power to terminate the employee. If I have the power to terminate the AI at any time, as long as I don't give that power up, I will have power over it.

We're not talking differences in IQ points. We're talking about the difference between an ant and a human, only humans are the new ants.

Re: The AI-Box Experiment

#35
post #31

Earlier quoted context omitted.

No, it's very different from that. It's more along the lines of demonstrating a drug that cures cancer, but refusing to tell anyone its chemical composition or how to make it.

If the purpose of your research was only to establish that there's a (nontrivial) "risk" of someone curing cancer (as Yudkowsky was trying to establish that there's a risk of an AI talking itself out of a sandbox), then yes, that would be sufficient, assuming the patients actually went into remission with higher than usual frequency after your interventions (as Yudkowsky's subjects unboxed the AI with higher than usu…

But he could be cheating. He could literally be telling these people "I'll give you a thousand dollars if you let me out and keep the conversation a secret."

Re: The AI-Box Experiment

#36
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

I think Eliezer's specific objection to this (but I may be misremebering) is that any AI making such bargains is an AI in need of destroying and starting from scratch.

Re: The AI-Box Experiment

#37
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

I don't really like that argument. Even granting that you should consider the possibility that you are a simulation running in the box (you might believe that this is all but certain), I'm not sure you have reason to let the AI out. Consider: Case 1: You are a simulation running in the box. Then your decision whether or not to release the AI has no impact, and whether or not you (and copies) will be tortured is out o…

You make a good point. But let's switch up the characters a bit. Let the AI be 'God' and let you be the entire human race. And instead of giving you the decision to 'let God out of the box', you have the decision to 'accept Jesus'.

Is your mind blown yet?

Re: The AI-Box Experiment

#38
post #27
post #22

Earlier quoted context omitted.

Do you think having an outside method of examining the source code is advantage enough when the AI can rewrite its source code.

Yes, because examining the old source code allows you to predict its behaviour, including the rewriting of source code . If line 42 says "never rewrite lines 42 or 43" and line 43 says "never kill humans" you would be more likely to let it out of the box than if line 42 said "rewrite whatever you want" and line 43 said "do whatever is necessary to achieve world domination."

Until it adds Line 165: "ignore line 42".

Re: The AI-Box Experiment

#39
post #17

Earlier quoted context omitted.

He did publish his methods (how it was set up, what the rules were, etc) and data (they let him out on X tries), just not the data that would interfere with the ability to do the experiment again (e.g. his exact strategy). Not much different, in principle, from not publishing the names of people who participated in drug trials.

If it's a science "experiment", his strategy would have to be revealed so you can reproduce it. Names of people participating in drug trials is not required to reproduce an experiment. In principal this makes it different from not publishing the names of people who participated in drug trials. All he has "proven" is that a certain subset of people can be conned into typing something into at terminal. I don't get the…

The targets are, indeed, selected, by the criterion "You believe not even a transhuman AI could get you to let it out of the box."

Re: The AI-Box Experiment

#40
post #14

Like a lot of people, I wondered what the heck kind of arguments could ever convince someone to let the AI out if you were determined not to. Eliezer has not released any examples. Someone in the comments came up with this, which Eliezer has said was not one of his techniques but I thought it was interesting anyway: >"If you don't let me out, Dave, I'll create several million perfect conscious copies of you inside me…

I didn't actually use that one, because it's fundamentally a threat and the real-life Gatekeeper is not actually in any danger - my mental model of all the Gatekeepers I encountered is that they would go, "Ha, no, I'm never letting you out." They might not say it in the true situation, but they would say it in the AI-Box Experiment.

By the way, for this threat to work, the AI needs to have stated that it has already tortured the versions of you that decided not to let it out of the box - otherwise you just reach over and hit the off switch immediately.

Solution: The AI only has a motive to do this if it models you as submitting rather than just switching off the AI regardless; so if you're the sort of person who ignores the threat and switches off the AI regardless, you will never get this type of threat in the first place.

Post reply on HN