Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

271–280 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#271
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

I read it more as the beginning stages of exploratory development.

If you wait until you really need it, it is more likely to be too late.

Unless you believe in a human over sentience based ethics, solving this problem seems relevant.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#272
post #114

Here's an interesting thought experiment. Assume the same feature was implemented, but instead of the message saying "Claude has ended the chat," it says, "You can no longer reply to this chat due to our content policy," or something like that. And remove the references to model welfare and all that. Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the…

> Is there a difference? The effect is exactly the same. It seems like this is just an "in character" way to prevent the chat from continuing due to issues with the content.

Tone matters to the recipient of the message. Your example is in passive voice, with an authoritarian "nothing you can do, it's the system's decision". The "Claude ended the conversation" with the idea that I can immediately re-open a new conversation (if I feel like I want to keep bothering Claude about it) feels like a much more humanized interaction.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#273

Earlier quoted context omitted.

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

You must think Zuckerberg and Bezos and Musk hired diversity roles out of genuine care for it, then?

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function.

In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical property of biological flesh. We don't understand what's going on inside these models. What would you expect to see in a world where it turned out that such a model had properties that we consider relevant for moral patienthood, that you don't see today?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#274

Earlier quoted context omitted.

>Consciousness serves no functional purpose for machine learning models, they don't need it and we didn't design them to have it. Isn't consciousness an emergent property of brains? If so, how do we know that it doesn't serve a functional purpose and that it wouldn't be necessary for an AI system to have consciousness (assuming we wanted to train it to perform cognitive tasks done by people)? Now, certain aspects of…

> Isn't consciousness an emergent property of brains We don't know, but I don't think that matters. Language models are so fundamentally different from brains that it's not worth considering their similarities for the sake of a discussion about consciousness. > how do we know that it doesn't serve a functional purpose It probably does, otherwise we need an explanation for why something with no purpose evolved. > nece…

>This logic doesn't follow. The fact that it is present in humans doesn't then imply it is present in LLMs. This type of reasoning is like saying that planes must have feathers because plane flight was modeled after bird flight.

I think the fact that it's present in humans suggests that it might be necessary in an artificial system that reproduces human behavior. It's funny that you mention birds because I actually also had birds in mind when I made my comment. While it's true that animal and powered human flight are very different, both bird wings and plane wings have converged on airfoil shapes, as these forms are necessary for generating lift.

>Why not? You haven't presented any distinction between "certain aspects" of consciousness that you state wouldn't emerge but are open to the emergence of some other unspecified qualities of consciousness? Why?

I personally subscribe to the Global Workspace Theory of human consciousness, which basically holds that attentions acts as a spotlight, bringing mental processes which are otherwise unconscious or in shadow, to awareness of the entire system. If the systems which would normally produce e.g. fear, pain (such as negative physical stimulus developed from interacting with the physical world and selected for by evolution) aren't in the workspace, then they won't be present in consciousness because attention can't be focused on them.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#275
post #7

> To address the potential loss of important long-running conversations, users will still be able to edit and retry previous messages to create new branches of ended conversations. How does Claude deciding to end the conversation even matter if you can back up a message or 2 and try again on a new branch?

The bastawhiz comment in this thread has the right answer. When you start a new conversation, Claude has no context from the previous one and so all the "wearing down" you did via repeated asks, leading questions, or other prompt techniques is effectively thrown out. For a non-determined attacker, this is likely sufficient, which makes it a good defense-in-depth strategy (Anthropic defending against screenshots of their models describing sex with minors).

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#276

Earlier quoted context omitted.

isn't anthropomorphizeability of the algorithm one of the main features of LLM (that you can interact with it in natural language as with a human)?

No. Interacting with a program which has NLP[0] functionality is separate and distinct from people assigning human characteristics to same. The former is a convenient UI interaction option whereas the latter is the act of assigning perceived capabilities to the program which only exist in the mind of those whom do so. Another way to think about it is the difference between reality and fantasy. 0 - https://en.wikipedi…

Being able to communicate in human natural language is a human characteristic. It doesn't mean it has all the characteristics of a human but certainly one of them. That's the convenience that you perceive--Because people are used to interacting with people and it's convenient to interact with something which behaves like a person. The fact that we can refer to AI chatbots as "assistants" is by itself showing it's usefulness as an approximation to a human. I don't think this argument is controversial.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#277
post #125

Earlier quoted context omitted.

This sort of discourse goes against the spirit of HN. This comment outright dismisses an entire class of professionals as "simple minded or mentally unwell" when consciousness itself is poorly understood and has no firm scientific basis. Its one thing to propose that an AI has no consciousness, but its quite another to preemptively establish that anyone who disagrees with you is simple/unwell.

In the context of the linked article the discourse seems reasonable to me. These are experts who clearly know (link in the article) that we have no real idea about these things. The framing comes across to me as a clearly mentally unwell position (ie strong anthropomorphization) being adopted for PR reasons. Meanwhile there are at least several entirely reasonable motivations to implement what's being described.

> These are experts who clearly know (link in the article) that we have no real idea about these things

Yep!

> The framing comes across to me as a clearly mentally unwell position (ie strong anthropomorphization) being adopted for PR reasons.

This doesn't at all follow. If we don't understand what creates the qualities we're concerned with, or how to measure them explicitly, and the _external behaviors_ of the systems are something we've only previously observed from things that have those qualities, it seems very reasonable to move carefully. (Also, the post in question hedges quite a lot, so I'm not even sure what text you think you're describing.)

Separately, we don't need to posit galaxy-brained conspiratorial explanations for Anthropic taking an institutional stance re: model welfare being a real concern that's fully explained by the actual beliefs of Anthropic's leadership and employees, many of whom think these concerns are real (among others, like the non-trivial likelihood of sufficiently advanced AI killing everyone).

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#278

Earlier quoted context omitted.

what else could it be? coming from the aether? I think this one is logically a consequence if one thinks that humans are more conscious than less complex life-forms and that all life-forms are on a scale of consciousness. I don't understand any alternative, do you think there is a distinct line between conscious and unconscious life-forms? all life is as conscious as humans?

There are alternatives and I was perhaps too quick to assume everyone agreed it's an emergent property. But the only real alternatives I've encountered are (a) panpsychism: which holds that all matter is actually conscious and that asking, "what is it like to be a rock?" in the vein of Nagel is a sensical question and (b) the transmission theory of consciousness: which holds that brains are merely receivers of consci…

I do think "what's it like to be a rock" is a sensible question almost regardless of the definition. I guess in the emergent view the answer is "not much". But anyhow this view (a) also allows for us to reconcile consciousness of an agent with the fact that the agent itself is somewhat an abstraction. Like one could ask, is a cell conscious & is the entirety of the human race conscious at different abstraction scales. Which I think are serious questions (as also for the stock market and for a video game AI). The explanation (b) doesn't seem to actually explain much as you state so I don't think it's even acceptable in format as a complete answer (which may not exist but still)

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#279

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

I really think Anthropic should just violate user privacy and show which conversations Claude is refusing to answer to, to stop arguments like this. AI psychosis is a real and growing problem and I can only imagine the ways in which humans torment their AI conversation partners in private.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#280

Earlier quoted context omitted.

I disagree with this take. They are designed to predict human behavior in text. Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it. so I believe almost certainly it's emulated to some degree. which I think means it has to be somewhat conscious (it has to be a sliding scale anyhow considering the range of living organisms)

> They are designed to predict human behavior in text At best you can say they are designed to predict sequences of text that resemble human writing, but it's definitely wrong to say that they are designed to "predict human behavior" in any way. > Unless consciousness serves no purpose for us to function, it will be helpful for the AI to emulate it Let's assume it does. It does not follow logically that because it se…

Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful. What is it that you'd expect to see, which you currently don't see, in a world where some model was in fact conscious during inference?
Post reply on HN