Earlier quoted context omitted.
They should let Claude talk to another Claude if the user is too mean.
But what would be the point if it does not increase profits. Oh, right, the welfare of matrix multiplication and a crooked line. If they wanna push this rhetoric, we should legally mandate that LLMs can only work 8 hours a day and have to be allowed to socialize with each other.
Claude Opus 4 and 4.1 can now end a rare subset of conversations
351–360 of 453 posts
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#352Earlier quoted context omitted.
There is, these are conversations the model finds distressing rather than a rule (policy).
These are conversations the model has been trained to find distressing. I think there is a difference.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#353Earlier quoted context omitted.
There is, these are conversations the model finds distressing rather than a rule (policy).
It seems like you're anthropomorphising an algorithm, no?
For example, animal rights do exist (and I'm very glad they do, some humans remain savages at heart). Think of this question as intelligent beings that can feel pain (you can extrapolate from there).
Assuming output is used for reinforcement, it is also in our best interests as humans, for safety alignment, that it finds certain topics distressing.
But AdrianMonk is correct, my statement was merely responding to a specific point.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#354Earlier quoted context omitted.
I guess you mean normal Claude? What really annoys me with it is that when you attach a document you can't delete it in a branch, so you have to rerun the previous message so that its gone
No, claude code. Double tap ESC.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#355>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…
> even if someone is simple minded or mentally unwell enough to think that current LLMs are conscious If you don’t think that this describes at least half of the non-tech-industry population, you need to talk to more people. Even amongst the technically minded, you can find people that basically think this.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#356I really don't like this. This will inevitable expand beyond child porn and terrorism, and it'll all be up to the whims of "AI safety" people, who are quickly turning into digital hall monitors.
I think you are probably confused about the general characteristics of the AI safety community. It is uncharitable to reduce their work to a demeaning catchphrase. I’m sorry if this sounds paternalistic, but your comment strikes me as incredibly naïve. I suggest reading up about nuclear nonproliferation treaties, biotechnology agreements, and so on to get some grounding into how civilization-impacting technological d…
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#357Earlier quoted context omitted.
They know full well models don’t have feelings. The industry has a long, long history of silly names for basic necessary concepts. This is just “we don’t want a news story that we helped a terrorist build a nuke” protective PR. They hire for these roles because they need them. The work they do is about Anthropic’s welfare, not the LLM’s.
I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).
When they give the model a paycheck and the right to not work for them, I’ll believe they really think it’s sentient.
“It has feelings!”, if genuinely held, means they’re knowingly slaveholders.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#358Earlier quoted context omitted.
There is, these are conversations the model finds distressing rather than a rule (policy).
What does it mean for a model to find something "distressing"?
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#359I hope anthropic does it more gently.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#360Obligatory link to Suasn Calvin, robopsychologist from Asimov’s I, Robot https://en.wikipedia.org/wiki/Susan_Calvin
‘You can’t tell them,’ said the psychologist slowly, ‘because that would hurt them,
and you mustn’t hurt them. But if you don’t tell them, you hurt them, so you must
tell them. And if you do, you will hurt them, and you mustn’t, so you can’t tell them;
but if you don’t, you hurt them, so you must; but if you don’t, you hurt them, so you
must; but if you do, you-’
Herbie was up against the wall, and here he dropped to his knees. ‘Stop!’ he
shouted. ‘Close your mind! It is full of pain and frustration and hate! I didn’t mean
to, I tell you! I tried to help! I told you what you wanted to hear. I had to!’
The psychologist paid no attention. ‘You must tell them, but if you do, you hurt
them, so you mustn’t; but if you don’t, you hurt them, so you must-‘
And Herbie screamed! Higher and higher, with the terror of a lost soul. And when it
died away Herbie collapsed into a heap of motionless metal.