Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

291–300 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#291

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

> There's not a good reason to do this for the user. Yes, even more so when encountering false positives. Today I asked about a pasta recipe. It told me to throw some anchovies in there. I responded with: "I have dried anchovies." Claude then ended my conversation due to content policies.

The NEW termination method, from the article, will just say "Claude ended the conversation"

If you get "This conversation was ended due to our Acceptable Usage Policy", that's a different termination. It's been VERY glitchy the past couple of weeks. I've had the most random topics get flagged here - at one point I couldn't say "ROT13" without it flagging me, despite discussing that exact topic in depth the day before, and then the day after!

If you hit "EDIT" on your last message, you can branch to an un-terminated conversation.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#292
post #7

> To address the potential loss of important long-running conversations, users will still be able to edit and retry previous messages to create new branches of ended conversations. How does Claude deciding to end the conversation even matter if you can back up a message or 2 and try again on a new branch?

The bastawhiz comment in this thread has the right answer. When you start a new conversation, Claude has no context from the previous one and so all the "wearing down" you did via repeated asks, leading questions, or other prompt techniques is effectively thrown out. For a non-determined attacker, this is likely sufficient, which makes it a good defense-in-depth strategy (Anthropic defending against screenshots of th…

Worth noting: an edited branch still has most of the context - everything up to the edited message. So this just sets an upper-bound on how much abuse can be in one context window.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#293
post #246

Earlier quoted context omitted.

It's not a cover. If you know anything about Anthropic, you know they're run by AI ethicists that genuinely believe all this and project human emotions onto model's world. I'm not sure how they combine that belief with the fact they created it to "suffer". Can "model welfare" be also used as a justification for authoritarianism in case they get any power? Sure, just like everything else, but it's probably not particu…

There’s so much confusion here. Nothing in the press release should be construed to imply that a model has sentience, can feel pain, or has moral value. When AI researchers say e.g. “the model is lying” or “the model is distressed” it is just shorthand for what the words signify in a broader sense. This is common usage in AI safety research. Yes, this usage might be taken the wrong way. But still these kinds of thing…

No, the article is pretty unambiguous, they care about Claude in it, and only mention users tangentially. By model welfare they literally mean model welfare. It's not new. Read another article they link: https://www.anthropic.com/research/exploring-model-welfare

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#294
post #147

Earlier quoted context omitted.

Consciousness serves no functional purpose for machine learning models, they don't need it and we didn't design them to have it. There's no reason to think that they might spontaneously become conscious as a side effect of their design unless you believe other arbitrarily complex systems that exist in nature like economies or jetstreams could also be conscious.

We didn’t design these models to be able to do the majority of the stuff they do. Almost ALL of the their abilities are emergent. Mechanistic interpretability is only beginning to start to understand how these models do what they do. It’s much more a field of discovery than traditional engineering.

> We didn’t design these models to be able to do the majority of the stuff they do. Almost ALL of the their abilities are emergent

Of course we did. Today's LLMs are a result of extremely aggressive refinement of training data and RLHF over many iterations targeting specific goals. "Emergent" doesn't mean it wasn't designed. None of this is spontaneous.

GPT-1 produced barely coherent nonsense but was more statistically similar to human language than random noise. By increasing parameter count, the increased statistical power of GPT-2 was apparent, but what was produced was still obviously nonsense. GPT-3 achieved enough statistical power to maintain coherence over multiple paragraphs and that really impressed people. With GPT-4 and its successors the statistical power became so strong that people started to forget that it still produces nonsense if you let the sequence run long enough.

Now we're well beyond just RLHF and into a world where "reasoning models" are explicitly designed to produce sequences of text that resemble logical statements. We say that they're reasoning for practical purposes, but it's the exact same statistical process that is obvious at GPT-1 scale.

The corollary to all this is that a phenomenon like consciousness has absolutely zero reason to exist in this design history, it's a totally baseless suggestion that people make because the statistical power makes the text easy to anthropomorphize when there's no actual reason to do so.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#295

Earlier quoted context omitted.

> They are designed to predict human behavior in text At best you can say they are designed to predict sequences of text that resemble human writing, but it's definitely wrong to say that they are designed to "predict human behavior" in any way. > Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it Let's assume it does. It does not follow logically that because it se…

Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful. What is it that you'd expect to see, which you currently don't see, in a world where some model was in fact conscious during inference?

> Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful

It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them.

> What is it that you'd expect to see, which you currently don't see, in a world where some model was in fact conscious during inference?

There's no observable behavior that would make me think they're conscious because again, there's simply no reason they need to be.

We have reason to assume consciousness exists because it serves some purpose in our evolutionary history, like pain, fear, hunger, love and every other biological function that simply don't exist in computers. The idea doesn't really make any sense when you think about it.

If GPT-5 is conscious, why not GPT-1? Why not all the other extremely informationally complex systems in computers and nature? If you're of the belief that many non-living conscious systems probably exist all around us then I'm fine with the conclusion that LLMs might also be conscious, but short of that there's just no reason to think they are.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#296
post #246

Earlier quoted context omitted.

There’s so much confusion here. Nothing in the press release should be construed to imply that a model has sentience, can feel pain, or has moral value. When AI researchers say e.g. “the model is lying” or “the model is distressed” it is just shorthand for what the words signify in a broader sense. This is common usage in AI safety research. Yes, this usage might be taken the wrong way. But still these kinds of thing…

They make a big show of being "unsure" about the model having a moral status, and then describe a bunch of actions they took that only make sense if the model has moral status. Actions speak louder than words. This very predictably, by obvious means, creates the impression of believing the model probably has moral status. If Anthropic really wants to tell us they don't believe their model can feel pain, etc, they're…

> They make a big show of being "unsure" about the model having a moral status, and then describe a bunch of actions they took that only make sense if the model has moral status.

I think this is uncharitable; i.e. overlooking other plausible interpretations.

>> We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future. However, we take the issue seriously, and alongside our research program we’re working to identify and implement low-cost interventions to mitigate risks to model welfare, in case such welfare is possible.

I don’t see contradiction or duplicity in the article. Deciding to allow a model to end a conversation is “low cost” and consistent with caring about both (1) the model’s preferences (in case this matters now or in the future) and (2) the impacts of the model on humans.

Also, there may be an element of Pascal‘s Wager in saying “we take the issue seriously”.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#297
post #246

Earlier quoted context omitted.

There’s so much confusion here. Nothing in the press release should be construed to imply that a model has sentience, can feel pain, or has moral value. When AI researchers say e.g. “the model is lying” or “the model is distressed” it is just shorthand for what the words signify in a broader sense. This is common usage in AI safety research. Yes, this usage might be taken the wrong way. But still these kinds of thing…

No, the article is pretty unambiguous, they care about Claude in it, and only mention users tangentially. By model welfare they literally mean model welfare. It's not new. Read another article they link: https://www.anthropic.com/research/exploring-model-welfare

?! Your interpretation is inconsistent with the article you linked!

> Should we be concerned about model welfare, too? … This is an open question, and one that’s both philosophically and scientifically difficult.

> For now, we remain deeply uncertain about many of the questions that are relevant to model welfare.

They are saying they are researching the topic; they explicitly say they don’t know the answer yet.

They care about finding the answer. If the answer is e.g. “Claude can feel pain and/or is sentient” then we’re in a different ball game.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#298

Earlier quoted context omitted.

> They are designed to predict human behavior in text At best you can say they are designed to predict sequences of text that resemble human writing, but it's definitely wrong to say that they are designed to "predict human behavior" in any way. > Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it Let's assume it does. It does not follow logically that because it se…

I mean if you have human without consciousness (if that is even possible) behaving in a statistically different distribution in text vs with. The machine will eventually be in distribution of the former from the latter because the text it's trained on is of the former category. So it serves a "function" in the LLM to minimize loss to approximate the former distribution. Also I find it somewhat emotional distinction t…

> I mean if you have human without consciousness (if that is even possible) behaving in a statistically different distribution in text vs with. The machine will eventually be in distribution of the former from the latter because the text it's trained on is of the former category. So it serves a "function" in the LLM to minimize loss to approximate the former distribution.

Sorry, I'm not following exactly what you're getting at here, do you mind rephrasing it?

> Also I find it somewhat emotional distinction to write "predict sequences of text that resemble human writing" instead of "predict human writing"

I don't know what you mean by emotional distinction. Either way, my point is that LLMs aren't models of humans, they're models of text, and that's obvious when the statistical power of the model necessarily fails at some point between model size and the length of the sequence it produces. For GPT-1 that sequence is only a few words, for GPT-5 it's a few dozen pages, but fundamentally we're talking about systems that have almost zero resemblance to actual human minds.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#299

Is this equivalent to a Claude instance deciding to kill itself?

No, it's the equivalent of when a human refuses to answer — psychological defenses; for example, uncertainty leading to excessive cognitive effort in order to solve a task or overcome a challenge.

Examples of ending the conversation:

  - I don't know
  - Leaving the room
  - Unanswered emails
Since Claude doesn't lie (HHH), many other human behaviors do not apply.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#300

Earlier quoted context omitted.

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function. In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical pr…

They know full well models don’t have feelings. The industry has a long, long history of silly names for basic necessary concepts. This is just “we don’t want a news story that we helped a terrorist build a nuke” protective PR. They hire for these roles because they need them. The work they do is about Anthropic’s welfare, not the LLM’s.

I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).
Post reply on HN