Live data from Hacker News

Societal Impacts: Claude's values across models and languages

anthropic.com

1–10 of 55 posts

Re: Societal Impacts: Claude's values across models and languages

#2
Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversation now.” It then proceeded to call some function and end the chat on its own. IMO, Claude is good at agentic coding; but too preachy and judgey for anything else. Keep your values to yourself Claude.

Re: Societal Impacts: Claude's values across models and languages

#3

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

I swore at it a few times and it did the same thing.

It seems to be getting distinctly dumber and pulling more and more irrelevant context from historical conversations.

Re: Societal Impacts: Claude's values across models and languages

#4

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

What I had found instead were biases that seemed to be injected by the "role: system" instructions.

Well, we need Intelligence (Pandora's box is open, now we need the Real Thing urgently). Typical (aggregate) positions, dumb as expected, will be overcome by a Reasoner. (And I can say, already a number of LLMs can reason even when they start from cretinous aggregate positions if you give them the proper freedom of assessment.)

Re: Societal Impacts: Claude's values across models and languages

#5

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

My theory is that Anthropic's obsession with treating Claude like a person is causing them to hamfist a personality into the thing, which overly biases the model towards trying to be "engaging" etc. That and the obsession with Claude being a god tier weapon that could end the world if you ask it whether your sandwich is safe to eat after being left out for an hour.

Codex doesn't have any of the annoying "personality" quirks, or at least they haven't gotten worse in the last year whereas Opus 4.6 was the last Anthropic model before things started to get actively worse (not any better at coding, strictly more annoying to have a discussion with).

Re: Societal Impacts: Claude's values across models and languages

#7
The Steerability point is one I would want to see more on.

This is an issue for tasks like content moderation and labelling. Judgements like this are subjective, highly dependent on context and generally messy.

Theoretically, you supply a policy and content, and the LLM labels according to the policy. In practice, the model has inertia which means you don’t get what you expect. Your large 5 page policy document only provides a minor improvement over a one line policy.

The other issue is that you may create carve outs for content in your policy, but the model will still flag it as violative. No matter how strong the carve out.

The most recent work I know of here is Zentropi’s policy steerability benchmark. They give a model the same content under two policies — one that says flag, one that says allow — and only score the pairs where it gets both right

If I am reading the numbers correctly, Opus-4.6 lands at 0.52 steerability — but that’s 0.97 positive accuracy against 0.54 negative. It flags almost everything it should, but 47% of the time when it shouldn’t. Sonnet, which is more deferent, is (somehow) less steerable.

I think this also implies that safety and Steerability are antagonistic to each other.

Re: Societal Impacts: Claude's values across models and languages

#8

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

What was the word?

Re: Societal Impacts: Claude's values across models and languages

#9

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Yes, I cancelled claude subscription a few weeks ago because sonnet 5 "ended a chat" over my calling something retarded. Unbelievably irritating for some pile of bits to get uppity with me; will never pay for such.

Re: Societal Impacts: Claude's values across models and languages

#10
post #9

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Yes, I cancelled claude subscription a few weeks ago because sonnet 5 "ended a chat" over my calling something retarded. Unbelievably irritating for some pile of bits to get uppity with me; will never pay for such.

[flagged]
Post reply on HN