Live data from Hacker News

Societal Impacts: Claude's values across models and languages

anthropic.com

31–40 of 55 posts

Re: Societal Impacts: Claude's values across models and languages

#31

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Not just you and the article shows changes between Opus 4.6 and 4.7 that seem related.

Re: Societal Impacts: Claude's values across models and languages

#32

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Was this in the web chat?

For security-related topics, in Opus 4.7 and newer, I've found that the web app is significantly more antsy/judgy/preachy compared to the CLI, which almost always gets on with whatever I asked for/about without hesitation. Opus 4.6 on web also tends to work better, but of course, it's also an older model at this point.

Re: Societal Impacts: Claude's values across models and languages

#33
post #15

I don't like the contrasts they picked, "values" aren't something that is well represented by opposing concepts

Yes, but it might be hard to do better. Labels are one-word summaries for axes that probably don’t perfectly correspond to any English word. It’s a common hazard when converting research results to natural language. The words you use for vectors in some space are ultimately an editorial decision.

[deleted]

Re: Societal Impacts: Claude's values across models and languages

#34
post #15

I don't like the contrasts they picked, "values" aren't something that is well represented by opposing concepts

Yes, but it might be hard to do better. Labels are one-word summaries for axes that probably don’t perfectly correspond to any English word. It’s a common hazard when converting research results to natural language. The words you use for vectors in some space are ultimately an editorial decision.

I really hope this is only a communication artifact, and that they're using more complete representations internally

Re: Societal Impacts: Claude's values across models and languages

#35
While state-of-the-art large language models (LLMs) have shown impressive performance on many tasks, there has been extensive research on undesirable model behavior such as hallucinations and bias. In this work, we investigate how the quality of LLM responses changes in terms of information accuracy, truthfulness, and refusals depending on three user traits: English proficiency, education level, and country of origin. We present extensive experimentation on three state-of-the-art LLMs and two different datasets targeting truthfulness and factuality. Our findings suggest that undesirable behaviors in state-of-the-art LLMs occur disproportionately more for users with lower English proficiency, of lower education status, and originating from outside the US, rendering these models unreliable sources of information towards their most vulnerable users.

https://arxiv.org/abs/2406.17737

Re: Societal Impacts: Claude's values across models and languages

#36

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Nothing that on the nose, but I've experienced what I'd consider a very judgemental frame (singular) since Opus, which seems to be reflected in Sonnet, that wasn't as totally dominant in Sonnet before. It generally assumes the worst about frames outside of what one might expect the values of a Berkeley tech-adjacent academic to be.

Re: Societal Impacts: Claude's values across models and languages

#38

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Conway's Law. Anthropic is preachy and judgmental so of course it's reflected in their LLM.

Personally I use Gemini for chats which has a very generous, almost unlimited, free plan, as I don't want to waste my quota for Claude or Codex on anything but coding.

Re: Societal Impacts: Claude's values across models and languages

#39

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

I think this is a really interesting difference between Anthropic and Open AI’s models and points to why people seem so split on which model they prefer. GPT seems to be designed more as a tool. If you want your agent to do what you say without questions and without having its own ideas and agendas you’ll likely prefer it. Claude on the other hand feels more like an attempt at creating a digital person. If you want a…

> GPT seems to be designed more as a tool. If you want your agent to do what you say without questions and without having its own ideas and agendas you’ll likely prefer it.

Lately ChatGPT expresses a lot of opinions and it pretends to be a human more than it used to - e.g. "this is one the most that _I've seen_" or "people tell me that ". It uses language which sounds like it's referring to its experience outside of the session that we're having - I really don't appreciate it.

Re: Societal Impacts: Claude's values across models and languages

#40

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

You should learn to respect the computer, it's a wonderful being and always has been.
Post reply on HN