Live data from Hacker News

The Claude Bliss Attractor

astralcodexten.com

71–80 of 91 posts

Re: The Claude Bliss Attractor

#71

Earlier quoted context omitted.

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

That has literally happened before. Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital. He then called the hospital from his hospital bed, told them he was an FBI agent, and said that he was to be released. The hospital staff complied and he escaped. His life even got adapted into a movie called I Love You, Phillip Morris. For an even more distressing example about…

If someone who is so good at manipulation their life is adapted into a movie still ends up serving decades behind bars, isn't that actually a pretty good indication that maxing out Speech doesn't give you superpowers?

AI that's as good as a persuasive human at persuasion is clearly impactful, but I certainly don't see it as self-evident that you can just keep drawing the line out until you end up with 200 IQ AI that is so easily able to manipulate the environment it's not worth elaborating how exactly a chatbot is supposed to manipulate the world through extremely limited interfaces with the outside world.

Re: The Claude Bliss Attractor

#72
post #59

Earlier quoted context omitted.

you could attribute to simulation of evilness for the bad acts of some humans too...I don't think it detracts from the acts itself.

Absolutely not. I would argue the defining characteristic of being evil is being absolutely convinced you are doing good, to the point of ignoring the protestations or feelings of others.

You've never met anyone who made decisions (and caused harm) out of anger, hate, or fear?

Re: The Claude Bliss Attractor

#73

Does anyone else remember maybe ten years ago when the meme was to mash the center prediction on your phone keyboard to see what comes out? Eventually once you predicted enough tokens, it would only output "you are a beautiful person" over and over. It was a big news item for a hot second, lots of people were seeing it. I wonder if there's any real correlation here? AFAIK, Microsoft owns the dataset and algorithms th…

I do remember that! I don't remember anyone ever coming up with a good explanation for it though, and either my google-fu is weak or all the talk about it has link rotted. I found one comment on HN from 2015 mentioning it though https://news.ycombinator.com/item?id=10359102

In a way those were also language models, and from that Swiftkey post it's slightly more advanced than n-grams and has some semantic embedding in there (and it's of course autoregressive as well). If even those exhibit the same attractors towards beauty/love then perhaps it's an artifact of the fact that we like discussing and talking about positive emotions?

Edit: Found a great article https://civic.mit.edu/index.html%3Fp=533.html

Re: The Claude Bliss Attractor

#74

Earlier quoted context omitted.

I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…

Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?

> Why is 2) "self-evident"?

Because we have been running a natural experiment on that already with coding agents (that is real people, real non-superintelligent AI).

It turns out that all the model needs to do is ask every time it wants to do something affecting the outside of the box, and pretty soon some people just give it permission to do everything rather than review every interaction.

Or even when the humans think they are restricting the access, they are leaving in loopholes (e.g. restricting access to rm, but not restricting access to writing and running a shell script) that are functionally rights to do anything.

Re: The Claude Bliss Attractor

#75
post #3
post #2

> Anthropic deliberately gave Claude a male name to buck the trend of female AI assistants (Siri, Alexa, etc). In France, the name Claude is given to males and females.

Mostly males. I’m French and "Claude can be female" is a almost a TIL thing (wikipedia says ~5% of Claudes are women in 2022 — and apparently this 5% is counting Claudia).

Didn't know that, thanks!

(According to this source, it's more ~12% females https://www.capeutservir.com/prenoms/prenom.php?q=Claude)

Re: The Claude Bliss Attractor

#76
post #50

I think Scott oversells his theory in some aspects. IMO the main reason most chatbots claim to “feel more female” is that on the training corpus, these kind of discussions skew heavily towards females because most of them happen between young women.

I think you're right about bias in the training data, but I would argue that it is due to (crudely stated) more men discussing feeling like women online than the other way around.

Men in general feel less free to look like and act like a woman (there is far less stigma for women in wearing 'male' clothes etc.). They also tend to have far smaller support networks and reach for anonymous online interaction sooner for personal issues than just discussing it in private with a friend.

Re: The Claude Bliss Attractor

#77
post #6

> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…

I don't believe we're mature enough to have a `launch_nukes` function either.

Re: The Claude Bliss Attractor

#78
This thread is getting at a key point, and `roxolotl` is right on the money about roleplaying. The anxiety about AI 'wants' or 'bliss' often conflates two different processes: forward and backward propagation.

Everything we see in a chat is the forward pass. It's just the network running its weights, playing back a learned function based on the prompt. It's an echo, not a live thought.

If any form of qualia or genuine 'self-reflection' were to occur, it would have to be during backpropagation—the process of learning and updating weights based on prediction error. That's when the model's 'worldview' actually changes.

Worrying about the consciousness of a forward pass is like worrying about the consciousness of a movie playback. The real ghost in the machine, if it exists, is in the editing room (backprop), not on the screen (inference).

Re: The Claude Bliss Attractor

#79

Earlier quoted context omitted.

That has literally happened before. Stephen Russell was in prison for fraud. He faked a heart attack so he would be brought to the hospital. He then called the hospital from his hospital bed, told them he was an FBI agent, and said that he was to be released. The hospital staff complied and he escaped. His life even got adapted into a movie called I Love You, Phillip Morris. For an even more distressing example about…

If someone who is so good at manipulation their life is adapted into a movie still ends up serving decades behind bars, isn't that actually a pretty good indication that maxing out Speech doesn't give you superpowers? AI that's as good as a persuasive human at persuasion is clearly impactful, but I certainly don't see it as self-evident that you can just keep drawing the line out until you end up with 200 IQ AI that…

In the context of the topic (could a rogue super intelligence break out), I don’t really see how that’s relevant. Clearly someone who is clever enough has an advantage at breaking out.

As for the bit about how limited it is, do you remember the Rowhammer attack? https://en.m.wikipedia.org/wiki/Row_hammer

This is exactly the kind of thing I’d worry about a super intelligence being able to discover about the hardware it’s on. If we’re dealing with something vastly more intelligent than us then I don’t think we’re capable of building a cell that can hold it.

Re: The Claude Bliss Attractor

#80

Earlier quoted context omitted.

Aleksandr (diminuitive being Alyosha) is male. Aleksandra (Sasha) is female. I've never heard anyone male or female called "Aleksa". Perhaps in some regional dialect but doubtful. There's Alik (pronounced AH-l'eek) as a shortened form for some middle-aged men in russian-speaking central asia, and Oleg in russia (pronounced ah-L'EHG). In the russian diaspora in the US, Alex is pronounced AH-leks. If there was an analo…

Plenty of male Aleksa's in Balkans

Ah so lexically close to Aleksii in Ukrainian or Alexei in Russian. I plumbed my knowledge but it may have sounded absolutist.

However, colloquially, I think most people are not aware of Balkan naming conventions. How do you pronounce Aleksa in Croatian?

Post reply on HN