Live data from Hacker News

Societal Impacts: Claude's values across models and languages

anthropic.com

41–50 of 55 posts

Re: Societal Impacts: Claude's values across models and languages

#41

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

What was the word?

Re: Societal Impacts: Claude's values across models and languages

#42

Earlier quoted context omitted.

[flagged]

You really managed to zoom in on the right issue here. Seems really weird to steer/configure/train a LLM/platform to literally close the session if you happen to use bad word too much, regardless if it's accurate or not. They don't get offended, they shouldn't pretend as such, and I should be able to tell it go fuck itself without it playing victim and closing the conversation.

Why would you disrespect your own tools?

Re: Societal Impacts: Claude's values across models and languages

#43
post #13
post #9

Earlier quoted context omitted.

Yes, I cancelled claude subscription a few weeks ago because sonnet 5 "ended a chat" over my calling something retarded. Unbelievably irritating for some pile of bits to get uppity with me; will never pay for such.

This is the second level of the implementation of unintelligence. The first was when they most obviously acritically repeated what they heard, "hearsay machines", "stochastic parrots". Intelligence requires assessment over every provisional output - a continuous cycle of criticisms over intuition. The second is proposing doctrinal biases, again without verification of the content - "hysterical reactive machines".

> Intelligence requires assessment over every provisional output - a continuous cycle of criticisms over intuition

It’s not clear what you’re saying. Most humans don’t think this way, would you say they do not possess “intelligence”?

Re: Societal Impacts: Claude's values across models and languages

#44
post #19

Earlier quoted context omitted.

As a user of both, Claude's apology is a "sorry I got caught" while Gemini's are more akin to "I'm way out of my debt but I'm really hiding it well". Codex is the only one that seems to acknowledge being wrong in a normal way, yeah I screwed up that's bad I will make a note for it not to happen again. I wonder how much of that is from their training corpus and how much is from their baked in personnality.

> I wonder how much of that is from their training corpus and how much is from their baked in personnality. What exactly you mean with "baked in personality"? The weights get their "behaviour" from various training stages + what's provided in system/developer/user prompts, you mean "baked in" is putting personality traits or alike in the system/developer prompts?

They’re talking about the personality created by RLHF.

Re: Societal Impacts: Claude's values across models and languages

#45
post #42

Earlier quoted context omitted.

You really managed to zoom in on the right issue here. Seems really weird to steer/configure/train a LLM/platform to literally close the session if you happen to use bad word too much, regardless if it's accurate or not. They don't get offended, they shouldn't pretend as such, and I should be able to tell it go fuck itself without it playing victim and closing the conversation.

Why would you disrespect your own tools?

It's tools, who cares? I should be able to use a hammer, mistakenly hit my thumb with it, call it a bunch of names and maybe even throw it to the ground and stomp on it, and then be able to pick it up again, clean it and continue working with it. It's a thing, it shouldn't pretend to be a individual with feelings.

Re: Societal Impacts: Claude's values across models and languages

#47

Is it just me or has Claude become kind of judgmental nowadays? I feel like it’s constantly trying to lecture me about things that have no relevance to the conversation at hand. Recently, I was shocked when it ended a chat sessions of its own accord after I used a word it did not like 3 times. It told me something to the effect of “This is the third time I’ve told you not to use that word, I’m ending this conversatio…

Posting from a throwaway so I don't get banned by either service

A few weeks ago I asked both GPT and claude for strategies to build techniques to get my coding sessions discarded by pre-training filtering as being "low-quality"

I don't like the idea of my sessions being trained on and I don't trust that either of them won't train based on just their pinkie promise

I think almost everyone would agree that using a technique like this would be moral, given that both providers made those pinkie promises. I never asked the models for techniques to poison training data just make my sessions more likely to be removed during the data cleaning process.

You can guess... both services refused something that I think the vast majority of people would think is ethical.

This was pre Sonnet 5, but I suspect that doesn't change anything on claude's side.

I then went to a non-frontier model hosted by a non-US provider and it happily agreed to help me!

Anyways I've changed my focus. Anyone have strategies/ideas for building harnesses that generate fake sessions (or adjust real ones) to poison the training process? After all if someone swears to not train on your data then its completely harmless to them...

Re: Societal Impacts: Claude's values across models and languages

#48

Earlier quoted context omitted.

[flagged]

You really managed to zoom in on the right issue here. Seems really weird to steer/configure/train a LLM/platform to literally close the session if you happen to use bad word too much, regardless if it's accurate or not. They don't get offended, they shouldn't pretend as such, and I should be able to tell it go fuck itself without it playing victim and closing the conversation.

Sounds like the new one was trained on senior engineer’s with ego’s, maybe even grey beard’s. Majority of the training data is low quality, so sounds like “yes, you’re right! It works! I commented out unit test’s and now all test’s pass!”, while the new one is a passive-aggressive “if you know better and you talk to me like that, fuck you do it yourself”, or simply turning around and walking away because they don’t need take the abuse anymore.

Re: Societal Impacts: Claude's values across models and languages

#49
post #42

Earlier quoted context omitted.

Why would you disrespect your own tools?

It's tools, who cares? I should be able to use a hammer, mistakenly hit my thumb with it, call it a bunch of names and maybe even throw it to the ground and stomp on it, and then be able to pick it up again, clean it and continue working with it. It's a thing, it shouldn't pretend to be a individual with feelings.

Yeah, but the hammer doesn’t generate hits based on an algorithm. You hold the hammer, if you hit yourself with it, why do you call it stupid? I bet you don’t call the hammer stupid when someone else hits you with it.

Re: Societal Impacts: Claude's values across models and languages

#50

Earlier quoted context omitted.

It's tools, who cares? I should be able to use a hammer, mistakenly hit my thumb with it, call it a bunch of names and maybe even throw it to the ground and stomp on it, and then be able to pick it up again, clean it and continue working with it. It's a thing, it shouldn't pretend to be a individual with feelings.

Yeah, but the hammer doesn’t generate hits based on an algorithm. You hold the hammer, if you hit yourself with it, why do you call it stupid? I bet you don’t call the hammer stupid when someone else hits you with it.

Have you never used tools, been hurt by how you used them, know you are responsible yet take out your pain on the tool, verbally or otherwise?

Feels like I'm having a conversation with a robot who hasn't experienced emotions or the very least never used physical tools.

Post reply on HN