> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…
I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…
The Claude Bliss Attractor
31–40 of 91 posts
Re: The Claude Bliss Attractor
#32> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…
I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…
Re: The Claude Bliss Attractor
#33Claude's increasing euphoria as a conversation goes can mislead me. I'll be exploring trade offs, and I'll introduce some novel ideas. Claude will use such enthusiasm that it will convince me that we're onto something. I'll be excited, and feed the idea back to a new conversation with Claude. It'll remind me that the idea makes risky trade offs, and would be better solved by with a simple solution. Try it out.
Re: The Claude Bliss Attractor
#34Claude does have an exuberant kind of “personality” where it feels like it wants to be really excited and interested about whatever subject. I wouldn’t describe it totally as sycophancy, more like panglossian. My least favorite AI personality of all is Gemma though, what a totally humorless and sterile experience that is.
'Perfect! I am now done with the totally zany solution that makes no sense, here it is!'
Re: The Claude Bliss Attractor
#35Earlier quoted context omitted.
Agreed that o3 can be brutally honest. If you ask it for direct feedback, even on personal topics, it will make observations that, if a person made them, would be borderline rude.
Isn't that what "direct feedback" means? I firmly believe you should be able to hit your fingers with a hammer, and in the process learn whether that's a good idea or not :)
Re: The Claude Bliss Attractor
#36Re: The Claude Bliss Attractor
#37> But in fact, I predicted this a few years ago. AIs don’t really “have traits” so much as they “simulate characters”. If you ask an AI to display a certain trait, it will simulate the sort of character who would have that trait - but all of that character’s other traits will come along for the ride. This is why the “omg the AI tries to escape” stuff is so absurd to me. They told the LLM to pretend that it’s a tortur…
I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…
So i’m now wondering, why are these researchers so bad at communicating? You explained this better than 90% of the blog posts i’ve read about this. They all focus on the “ai did x” instead of _why_ it’s concerning with specific examples.
Re: The Claude Bliss Attractor
#38Earlier quoted context omitted.
I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…
The surprise! Is what I’m surprised by though. They are incredible role players so when they role play “evil ai” they do it well.
Re: The Claude Bliss Attractor
#39> Anthropic deliberately gave Claude a male name to buck the trend of female AI assistants (Siri, Alexa, etc). In France, the name Claude is given to males and females.
Aleksa is male Slavic name
In the russian diaspora in the US, Alex is pronounced AH-leks. If there was an analogous Alexa (there isn't) the pronunciation would be ah-LEK-sa like the service.
Re: The Claude Bliss Attractor
#40Earlier quoted context omitted.
I think this reaction misses the point that the "omg the AI tries to escape" people are trying to make when it tries to escape. The worry among big AI doomers has never been that AI somehow inherently is resentful or evil or has something "going on internally" that makes it dangerous. It's a worry that stems from three seemingly self-evident axioms: 1) A sufficiently powerful and capable superintelligence, singlemind…
Why is 2) "self-evident"? Do you think it's a given that, in any situation, there's something you could say that would manipulate humans to get what you want? If you were smart enough, do you think you could talk your way out of prison?