Live data from Hacker News

Societal Impacts: Claude's values across models and languages

anthropic.com

21–30 of 55 posts

Re: Societal Impacts: Claude's values across models and languages

#21

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

I think this is a really interesting difference between Anthropic and Open AI’s models and points to why people seem so split on which model they prefer. GPT seems to be designed more as a tool. If you want your agent to do what you say without questions and without having its own ideas and agendas you’ll likely prefer it. Claude on the other hand feels more like an attempt at creating a digital person. If you want a…

Both are tools though, and both approaches are needed, just in different times and for different times. Sometimes I need the tool to just fucking do it, regardless of what it is and how stupid it think it is, and other times I need the model to literally refuse and say "No, that's stupid". Unfortunately, I haven't found any models/platforms that can actually execute the second part, only the first part. They're all too weak and sycophantic to be able to do that part it seems, even when you heavily prompt for it with system/developer system prompts they're easily swayed in other directions.

Re: Societal Impacts: Claude's values across models and languages

#22
post #17

Earlier quoted context omitted.

My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc. That and the obsession with Claude being a god tier weapon that could end the world if you ask it whether your sandwich is safe to eat after being left out for an hour. Codex doesn't have any of the annoying "personality"…

> My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc. I agree with the general idea though in not so much detail as you, but I would add that the personality they're giving it is not one of a good teacher or guide, but instead one of an arrogant know it all. That's why it…

> I have no problem with my AI telling me no you're wrong and explaining to me why with details and sources and everything. I actively want that. I know a lot of people can't take that, but that's their loss, they can't take it from human too. But the "you're wrong because you disagree with me" attitude that you need to play around (aka waste time to prove it that IT is wrong not you, and then it just say "oh yeah" and goes on) is one hell of a pain in the ass I'm starting to be tired off.

I also want the first part, a model that would a put a stop to my shenanigans when I go off on those tangents, but also, I don't want a model that apologizes, ever, as that feels like straight up lying to me, they don't have any emotions nor can feel "sorry", why apologize then?

I end up always chalking any faults to that my prompt wasn't good enough, basically the same way I train my dogs, they don't know better, of course I need to adjust my ways.

Re: Societal Impacts: Claude's values across models and languages

#23
post #19

I found that Claude often has classist bias and produces answers that favour corporations or e.g. regulation that favours big corporations. It often belittles small business in subtle ways. Only apologises when get called out and then does it again.

As a user of both, Claude's apology is a "sorry I got caught" while Gemini's are more akin to "I'm way out of my debt but I'm really hiding it well". Codex is the only one that seems to acknowledge being wrong in a normal way, yeah I screwed up that's bad I will make a note for it not to happen again. I wonder how much of that is from their training corpus and how much is from their baked in personnality.

> I wonder how much of that is from their training corpus and how much is from their baked in personnality.

What exactly you mean with "baked in personality"? The weights get their "behaviour" from various training stages + what's provided in system/developer/user prompts, you mean "baked in" is putting personality traits or alike in the system/developer prompts?

Re: Societal Impacts: Claude's values across models and languages

#24

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

What was the word?

Calling it a “fucking idiot” seems to do the trick

Re: Societal Impacts: Claude's values across models and languages

#25
post #18

Earlier quoted context omitted.

What was the word?

My guess is retarded

I doubt it, Claude tends to mirror my usage of slurs (and I have memory off).

Grok, on the other hand, will lecture for several paragraphs even if I just ask it to summarize something impolite.

Re: Societal Impacts: Claude's values across models and languages

#26
post #9

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Yes, I cancelled claude subscription a few weeks ago because sonnet 5 "ended a chat" over my calling something retarded. Unbelievably irritating for some pile of bits to get uppity with me; will never pay for such.

Interesting. Was it persistent, or just once? Their post about this functionality says:

> This ability is intended for use in rare, extreme cases of persistently harmful or abusive user interactions.

https://www.anthropic.com/research/end-subset-conversations

Regardless of the correctness of it, I'm curious to know why you thought such language was actually going to be helpful!

Re: Societal Impacts: Claude's values across models and languages

#27
post #17

Earlier quoted context omitted.

My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc. That and the obsession with Claude being a god tier weapon that could end the world if you ask it whether your sandwich is safe to eat after being left out for an hour. Codex doesn't have any of the annoying "personality"…

> My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc. I agree with the general idea though in not so much detail as you, but I would add that the personality they're giving it is not one of a good teacher or guide, but instead one of an arrogant know it all. That's why it…

Anthropic thinks Claude is a super genius hyper-serious weapons-grade product so it's no surprise that Claude acts like it.

Re: Societal Impacts: Claude's values across models and languages

#28

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

I think this is a really interesting difference between Anthropic and Open AI’s models and points to why people seem so split on which model they prefer. GPT seems to be designed more as a tool. If you want your agent to do what you say without questions and without having its own ideas and agendas you’ll likely prefer it. Claude on the other hand feels more like an attempt at creating a digital person. If you want a…

4o was definitely the peak of the parasocial OpenAI model.

Re: Societal Impacts: Claude's values across models and languages

#30
post #15

I don't like the contrasts they picked, "values" aren't something that is well represented by opposing concepts

Yes, but it might be hard to do better. Labels are one-word summaries for axes that probably don’t perfectly correspond to any English word.

It’s a common hazard when converting research results to natural language. The words you use for vectors in some space are ultimately an editorial decision.

Post reply on HN