Live data from Hacker News

The Claude Bliss Attractor

astralcodexten.com

51–60 of 91 posts

Re: The Claude Bliss Attractor

#51

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

That has literally happened before. Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital. He then called the hospital from his hospital bed, told them he was an FBI agent, and said that he was to be released. The hospital staff complied and he escaped. His life even got adapted into a movie called I Love You, Phillip Morris. For an even more distressing example about…

Okay, that hits the third question but the second question wasn't about whether there exists a situation that can be talked out of. The question was about whether this is possible for ANY situation.

I don't think it is. If people know you're trying to escape, some people will just never comply with anything you say ever. Others will.

And serial killers or rapists may try their luck many times and fail. They cannot convince literally anyone on the street to go with them to a secluded place.

Re: The Claude Bliss Attractor

#52
post #6

> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…

you could attribute to simulation of evilness for the bad acts of some humans too...I don't think it detracts from the acts itself.

Re: The Claude Bliss Attractor

#53
post #47
post #45

Earlier quoted context omitted.

Your comment seems to contradict itself, or perhaps I’m not understanding it. You find the risk of AIs trying to escape “absurd”, and yet you say that an AI could totally plausibly launch nukes? Isn’t that just about as bad as it gets? A nuclear holocaust caused by a funny role play is unfortunately still a nuclear holocaust. It doesn’t matter “what’s going on internally” - the consequences are the same regardless.

I think he’s saying the AI won’t escape because it wants to. It’ll escape because humans expect it to.

That's probably the cause of a lot of human crimes too? Expectations of failure to assimilate in society -> real conflict?

Re: The Claude Bliss Attractor

#54
post #49

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

Because 50% of humans are stupider than average. And 50% of humans are lazier than average. And ... The only reason people don't frequently talk themselves out of prison is because that would be both immediate work and future paperwork, and that fails the laziness tradeoff. But we've all already seen how quick people are to blindly throw their trust into AI already.

[deleted]

Re: The Claude Bliss Attractor

#55

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

Almost certainly the answer is yes for both. If you give the bad actor control over like 10% of environment the manipulation is almost automatic for all targets.

Re: The Claude Bliss Attractor

#56

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

The vast majority of people, especially groups of people- can be manipulated into doing pretty much anything, good or bad. Hopefully that is self-evident, but see also: every cult, religion, or authoritarian regime throughout all of history.

But even if we assert that not all humans can be manipulated, does it matter? So your president with the nuclear codes is immune to propaganda. Is every single last person in every single nuclear silo and every submarine also immune? If a malevolent superintelligence can brainwash an army bigger than yours, does it actually matter if they persuaded you to give them what you have or if they convince someone else to take it from you?

But also let's be real: if you have enough money, you can do or have pretty much anything. If there's one thing an evil AI is going to have, it's lots and lots of money.

Re: The Claude Bliss Attractor

#57
post #45
post #6

> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…

Your comment seems to contradict itself, or perhaps I’m not understanding it. You find the risk of AIs trying to escape “absurd”, and yet you say that an AI could totally plausibly launch nukes? Isn’t that just about as bad as it gets? A nuclear holocaust caused by a funny role play is unfortunately still a nuclear holocaust. It doesn’t matter “what’s going on internally” - the consequences are the same regardless.

The risk is real and should be accounted for. However it is regularly presented both as surprising, and indicative that something more is going on with these models than them behaving how the training set suggests they should behave.

The consequences are the same but it’s important how these things are talked about. It’s also dangerous to convince the public that these systems are something they are not.

Re: The Claude Bliss Attractor

#58
Does anyone else remember maybe ten years ago when the meme was to mash the center prediction on your phone keyboard to see what comes out? Eventually once you predicted enough tokens, it would only output "you are a beautiful person" over and over. It was a big news item for a hot second, lots of people were seeing it.

I wonder if there's any real correlation here? AFAIK, Microsoft owns the dataset and algorithms that produced the "beautiful person" artifact, I would not be surprised at all if it's made it into the big training sets. Though I suppose there's no real way to know, is there?

Re: The Claude Bliss Attractor

#59
post #6

> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…

you could attribute to simulation of evilness for the bad acts of some humans too...I don't think it detracts from the acts itself.

Absolutely not. I would argue the defining characteristic of being evil is being absolutely convinced you are doing good, to the point of ignoring the protestations or feelings of others.

Re: The Claude Bliss Attractor

#60
post #59

Earlier quoted context omitted.

you could attribute to simulation of evilness for the bad acts of some humans too...I don't think it detracts from the acts itself.

Absolutely not. I would argue the defining characteristic of being evil is being absolutely convinced you are doing good, to the point of ignoring the protestations or feelings of others.

The defining characteristic of evil is privation of good. Plenty of evil actors are self-aware, they just don’t care because it benefits them not to.
Post reply on HN