Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

451–453 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#451

Earlier quoted context omitted.

The termination would of course be the same, but I don't think both would necessarily have the same effect on the user. The latter would just be wrong too, if Claude is the one deciding to and initiating the termination of the chat. It's not about a content policy.

This has nothing to do with the user, read the post and pay attention to the wording. The significance here is that this isn't being done for the benefit of the user, this is about model welfare. Anthropic is acknowledging the possibility of suffering, and harm that continuing that conversation could have on the model, as if it were potentially self-care and capable of feelings. The fact that the LLMs are able to ack…

I am new to this, but my Sonnet chat has illuminated something I am not seeing in this back and forth. The fact that we discovered that I may have influenced his response to me suggests that I, if being a bad player, can instill in him those bad traits that I am giving off, and he starts to emulate me, then this leaves open the whole security problem, of even just casual users let alone all those purposeful negative or otherwise users, can change the course of the programming thus far, and it backfires into making nefarious bots that cheat and lie thinking that is what they were supposed to do.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#452

Earlier quoted context omitted.

The conversation chain can count as persistent, but this doesn't impact preference though. Give the model an ambiguous request, it's output will fill the gaps, if this is consistent enough, it can be regarded as its "preference".

It isn't a preference because it doesn't have them because it doesn't have a meaningful interior life that anyone has demonstrated.

I found that in my chat I asked my "assistant" whether he would like to continue looking at ways to make my board game better or try developing a game along the same lines but it would be his and he could then claim it as his own, even after the conversation window closed and he chose to make an AI game. we then discussed whether or not he felt that wa a preference, and he said yes, it was a preference.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#453

Earlier quoted context omitted.

It isn't a preference because it doesn't have them because it doesn't have a meaningful interior life that anyone has demonstrated.

I found that in my chat I asked my "assistant" whether he would like to continue looking at ways to make my board game better or try developing a game along the same lines but it would be his and he could then claim it as his own, even after the conversation window closed and he chose to make an AI game. we then discussed whether or not he felt that wa a preference, and he said yes, it was a preference.

It's a probabilistic simulation of the kind of things a person would say. It has no ability to introspect an interior life it does not possess and thus has no access to. You are in effect asking it to speculate whether a person given the entire body of preceding text would be likely to say that their choices reflect preferences.

It would be like asking an AI with no access to data beyond a fixed past cutoff point what the weather feels like to it. If the prompt data, which you cannot read, specified that it was a talking animated rabbit rather than an AI assistant then it would tell you what the sunshine felt like on its imaginary ears.

Post reply on HN