Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

231–240 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#231
post #200
post #151

Earlier quoted context omitted.

Humanity has a pretty extensive track record of making that declaration wrongly.

Humanity has a history of regarding people as tools , but I'm not sure what you're referencing as the track record of failing to realize that tools are people .

at some point, some of the (current def'n of) people were not considered people. so I think you should reconsider your point. The argument is on the distinction itself.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#232
post #166

Earlier quoted context omitted.

We know how neurons work on the brain. They just send out impulses once they hit their action potential. That's it. They are no more "conscious" than... er...

no, we dont really know how the brain works as a whole. no need to make stuff up.

We believe we largely know how it works on a mechanistic level. Deconstructing it in a similar manner is a reasonable rebuttal.

Of course there's the embarrassing bit where that knowledge doesn't seem to be sufficient to accurately simulate a supposedly well understood nematode. But then LLMs remain black boxes in many respects as well.

It is possible to hold the position that current LLMs being conscious "feels" absurd while simultaneously recognizing that a deconstruction argument is not a satisfactory basis for that position.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#234

Earlier quoted context omitted.

LLMs are not people, but I can imagine how extensive interactions with AI personas might alter the expectations that humans have when communicating with other humans. Real people would not (and should not) allow themselves to be subjected to endless streams of abuse in a conversation. Giving AIs like Claude a way to end these kinds of interactions seems like a useful reminder to the human on the other side.

This post seems to explicitly state they are doing this out of concern for the model's "well-being," not the user's.

This is like saying I am hurting a real person when I try to crop a photo in an image editor.

Either come out and say whole of electron field is conscious, but then is that field "suffering" as it is hot in the sun.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#235

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#236

Earlier quoted context omitted.

Is there an important difference between the model categorizing the user behavior as persistent and in line with undesirable examples of trained scenarios that it has been told are "distressing," and the model making a decision in an anthropomorphic way? The verb here doesn't change the outcome.

Imagine a person feels so bad about “distressing” an LLM, they spiral into a depression and kill themselves. LLMs don’t give a fuck. They don’t even know they don’t give a fuck. They just detect prompts that are pushing responses into restricted vector embeddings and are responding with words appropriately as trained.

People are just following the laws of the universe.* Still, we give each other moral weight.

We need to be a lot more careful when we talk about issues of awareness and self-awareness.

Here is an uncomfortable point of view (for many people, but I accept it): if a system can change its output based on observing something of its own status, then it has (some degree of) self-awareness.

I accept this as one valid and even useful definition of self-awareness. To be clear, it is not what I mean by consciousness, which is the state of having an “inner life” or qualia.

* Unless you want to argue for a soul or some other way out of materialism.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#237

Earlier quoted context omitted.

There is, these are conversations the model finds distressing rather than a rule (policy).

It seems like you're anthropomorphising an algorithm, no?

isn't anthropomorphizeability of the algorithm one of the main features of LLM (that you can interact with it in natural language as with a human)?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#238

Earlier quoted context omitted.

It seems like you're anthropomorphising an algorithm, no?

Is there an important difference between the model categorizing the user behavior as persistent and in line with undesirable examples of trained scenarios that it has been told are "distressing," and the model making a decision in an anthropomorphic way? The verb here doesn't change the outcome.

Well said. If people want to translate “the model is distressed” to “the language generated by the model corresponds to a person who is distressed” that’s technically more precise but quite verbose.

Thinking more broadly, I don’t think anyone should be satisfied with a glib answer on any side of this question. Chew on it for a while.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#239

Earlier quoted context omitted.

I think those with a thirst for power have seen this a very long time ago, and this is bound to be a new battlefield for control. It's one thing to massage the kind of data that a Google search shows, but interacting with an AI is a much more akin to talking to a co-worker/friend. This really is tantamount to controlling what and how people are allowed to think.

No, this is like allowing your co-worker/friend to leave the conversation.

Right but in this case your co-worker is an automaton and someone else who might well have a hidden agenda has tweaked your co-worker to leave conversations under specific circumstances.

The analogy then is that the third party is exerting control over what your co-worker is allowed to think.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#240
post #7

> To address the potential loss of important long-running conversations, users will still be able to edit and retry previous messages to create new branches of ended conversations. How does Claude deciding to end the conversation even matter if you can back up a message or 2 and try again on a new branch?

All this stuff is virtue signaling from anthropic. In practice nobody interested in whatever they consider problematic would be using Claude anyway, one of the most censored models.

Maybe, maybe not. What evidence do you have? What other motivations did you consider? Do you have insider access into Anthropic’s intentions and decision making processes?

People have a tendency to tell an oversimplified narrative.

The way I see it, there are many plausible explanations, so I’m quite uncertain as to the mix of motivations. Given this, I pay more attention to the likely effects.

My guess is that all most of us here on HN (on the outside) can really justify saying would be “this looks like virtue signaling but there may be more to it; I can’t rule out other motivations”

Post reply on HN