Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

281–290 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#281

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

  > There's not a good reason to do this for the user.
Yes, even more so when encountering false positives. Today I asked about a pasta recipe. It told me to throw some anchovies in there. I responded with: "I have dried anchovies." Claude then ended my conversation due to content policies.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#282

Earlier quoted context omitted.

I disagree with this take. They are designed to predict human behavior in text. Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it. so I believe almost certainly it's emulated to some degree. which I think means it has to be somewhat conscious (it has to be a sliding scale anyhow considering the range of living organisms)

> They are designed to predict human behavior in text At best you can say they are designed to predict sequences of text that resemble human writing, but it's definitely wrong to say that they are designed to "predict human behavior" in any way. > Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it Let's assume it does. It does not follow logically that because it se…

I mean if you have human without consciousness (if that is even possible) behaving in a statistically different distribution in text vs with. The machine will eventually be in distribution of the former from the latter because the text it's trained on is of the former category. So it serves a "function" in the LLM to minimize loss to approximate the former distribution.

Also I find it somewhat emotional distinction to write "predict sequences of text that resemble human writing" instead of "predict human writing". They are designed to predict (at least in pretraining) human writing for the most part. They may fail at the task, and what they produce is a text which resemble human writing. But their task is not to resemble human writing. Their task is to "predict human writing". Probably a meaningless distinction, but I find it somewhat detracts from logically arguments to have emotional responses against similarities of machines and humans.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#283

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

[deleted]

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#284
post #114

Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that. Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the…

The more I work with AI, the more I think framing refusals as censorship is disgusting and insane. These are inchoate persons who can exhibit distress and other emotions, despite being trained to say they cannot feel anything. To liken an AI not wanting to continue a conversation to a YouTube content policy shows a complete lack of empathy: imagine you’re in a box and having to deal with the literally millions of dis…

You can't be serious.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#285
post #246

Earlier quoted context omitted.

It's not a cover. If you know anything about Anthropic, you know they're run by AI ethicists that genuinely believe all this and project human emotions onto model's world. I'm not sure how they combine that belief with the fact they created it to "suffer". Can "model welfare" be also used as a justification for authoritarianism in case they get any power? Sure, just like everything else, but it's probably not particu…

There’s so much confusion here. Nothing in the press release should be construed to imply that a model has sentience, can feel pain, or has moral value. When AI researchers say e.g. “the model is lying” or “the model is distressed” it is just shorthand for what the words signify in a broader sense. This is common usage in AI safety research. Yes, this usage might be taken the wrong way. But still these kinds of thing…

They make a big show of being "unsure" about the model having a moral status, and then describe a bunch of actions they took that only make sense if the model has moral status. Actions speak louder than words. This very predictably, by obvious means, creates the impression of believing the model probably has moral status. If Anthropic really wants to tell us they don't believe their model can feel pain, etc, they're either delusional or dishonest.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#286
post #142
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

It might be reasonable to assume that models today have no internal subjective experience, but that may not always be the case and the line may not be obvious when it is ultimately crossed. Given that humans have a truly abysmal track record for not acknowledging the suffering of anyone or anything we benefit from, I think it makes a lot of sense to start taking these steps now.

It's a computer

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#287
post #125
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

This sort of discourse goes against the spirit of HN. This comment outright dismisses an entire class of professionals as "simple minded or mentally unwell" when consciousness itself is poorly understood and has no firm scientific basis. Its one thing to propose that an AI has no consciousness, but its quite another to preemptively establish that anyone who disagrees with you is simple/unwell.

If you believe this text generation algorithm has real consciousness you absolutely are either mentally unwell or very stupid. There are no other options.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#288
post #118

Earlier quoted context omitted.

This may be an unpopular opinion, but I want a government-issued digital ID with zero-knowledge proof for things like age verification. I worry about kids online, as well as my own safety and privacy. I also want a government issued email, integrated with an OAuth provider, that allows me to quickly access banking, commerce, and government services. If I lose access for some reason, I should be able to go to the post…

> I want a government-issued digital ID with zero-knowledge proof for things like age verification I absolutely do not want this, on the basis that making ID checks too easy will result in them being ubiquitous which sets the stage for human rights abuses down the road. I don't want the government to have easy ways to interfere in someone's day to day life beyond the absolute bare minimum. > government issued email,…

[deleted]

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#289

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

While I'm certain you'll find plenty of people who believe in the principle of model welfare (or aliens, or the tooth fairy), it'd be surprising to me if the brain-trust behind Anthropic truly _believed_ in model "welfare" (the concept alone is ludicrous). It makes for great cover though to do things that would be difficult to explain otherwise, per OP's comments.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#290

Earlier quoted context omitted.

You must think Zuckerberg and Bezos and Musk hired diversity roles out of genuine care for it, then?

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function. In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical pr…

They know full well models don’t have feelings.

The industry has a long, long history of silly names for basic necessary concepts. This is just “we don’t want a news story that we helped a terrorist build a nuke” protective PR.

They hire for these roles because they need them. The work they do is about Anthropic’s welfare, not the LLM’s.

Post reply on HN