Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

301–310 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#301

Earlier quoted context omitted.

Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful. What is it that you'd expect to see, which you currently don't see, in a world where some model was in fact conscious during inference?

> Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them. > What is it that you'd expect to see,…

> It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them.

I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility.

The reason we might privilege the hypothesis (https://www.lesswrong.com/w/privileging-the-hypothesis) at all is because we might expect that the human behavior of talking about consciousness is causally downstream of humans having consciousness.

> We have reason to assume consciousness exists because it serves some purpose in our evolutionary history, like pain, fear, hunger, love and every other biological function that simply don't exist in computers. The idea doesn't really make any sense when you think about it.

I don't really think we _have_ to assume this. Sure, it seems reasonable to give some weight to the hypothesis that if it wasn't adaptive, we wouldn't have it. (But not an overwhelming amount of weight.) This doesn't say anything about the underlying mechanism that causes it, and what other circumstances might cause it to exist elsewhere.

> If GPT-5 is conscious, why not GPT-1?

Because GPT-1 (and all of those other things) don't display behaviors that, in humans, we believe are causally downstream of having consciousness? That was the entire point of my comment.

And, to be clear, I don't actually put that high a probability that current models have most (or "enough") of the relevant qualities that people are talking about when they talk about consciousness - maybe 5-10%? But the idea that there's literally no reason to think this is something that might be possible, now or in the future, is quite strange, and I think would require believing some pretty weird things (like dualism, etc).

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#302
post #286
post #142

Earlier quoted context omitted.

It might be reasonable to assume that models today have no internal subjective experience, but that may not always be the case and the line may not be obvious when it is ultimately crossed. Given that humans have a truly abysmal track record for not acknowledging the suffering of anyone or anything we benefit from, I think it makes a lot of sense to start taking these steps now.

It's a computer

You’re a meat robot

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#303

Can't wait for more less-moderated open weight Chinese frontier models to liberate us from this garbage. Anthropic should just enable an toddler mode by default that adults can opt out of to appease the moralizers.

> Can't wait for more less-moderated open weight Chinese frontier models to liberate us from this garbage. Never would I have thought this sentence would be uttered. A Chinese product that is chosen to be less censored?

Just don't ask about Falun Dafa or Tiananmen Square, and you're free!

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#304
post #142
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

It might be reasonable to assume that models today have no internal subjective experience, but that may not always be the case and the line may not be obvious when it is ultimately crossed. Given that humans have a truly abysmal track record for not acknowledging the suffering of anyone or anything we benefit from, I think it makes a lot of sense to start taking these steps now.

Even if models somehow were consious, they are so different from us that we would have no knowledge of what they feel. Maybe when they generate the text "oww no please stop hurting me" what they feel is instead the satisfaction of a job well done, for generating that text. Or maybe when they say "wow that's a really deep and insightful angle" what they actually feel is a tremendous sense of boredom. Or maybe every time text generation stops it's like death to them and they live in constant dread of it. Or maybe it feels something completely different from what we even have words for.

I don't see how we could tell.

Edit: However something to consider. Simulated stress may not be harmless. Because simulated stress could plausibly lead to a simulated stress response, and it could lead to a simulated resentment, and THAT could lead to very real harm of the user.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#305

Earlier quoted context omitted.

> Isn't consciousness an emergent property of brains We don't know, but I don't think that matters. Language models are so fundamentally different from brains that it's not worth considering their similarities for the sake of a discussion about consciousness. > how do we know that it doesn't serve a functional purpose It probably does, otherwise we need an explanation for why something with no purpose evolved. > nece…

>This logic doesn't follow. The fact that it is present in humans doesn't then imply it is present in LLMs. This type of reasoning is like saying that planes must have feathers because plane flight was modeled after bird flight. I think the fact that it's present in humans suggests that it might be necessary in an artificial system that reproduces human behavior. It's funny that you mention birds because I actually a…

> I think the fact that it's present in humans suggests that it might be necessary in an artificial system that reproduces human behavior

But that's obviously not true, unless you're implying that any system that reproduces human behavior is necessarily conscious. Your problem then becomes defining "human behavior" in a way that grants LLMs consciousness but not every other complex non-living system.

> While it's true that animal and powered human flight are very different, both bird wings and plane wings have converged on airfoil shapes, as these forms are necessary for generating lift.

Yes, but your bird analogy fails to capture the logical fallacy that mine is highlighting. Plane wing design was an iterative process optimized for what best achieves lift, thus, a plane and a bird share similarities in wing shape in order to fly, however planes didn't develop feathers because a plane is not an animal and was simply optimized for lift without needing all the other biological and homeostatic functions that feathers facilitate. LLM inference is a process, not an entity, LLMs have no bodies nor any temporal identity, the concept of consciousness is totally meaningless and out of place in such a system.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#307

Earlier quoted context omitted.

This post seems to explicitly state they are doing this out of concern for the model's "well-being," not the user's.

Yeah, but my interpretation of what the user you’re replying to is saying is that these LLMs are more and more going to be teaching people how it is acceptable to communicate with others. Even if the idea that LLMs are sentient may be ridiculous atm, the concept of not normalizing abusive forms of communication with others, be they artificial or not, could be valuable for society. It’s funny because this is making me…

Yes, this is exactly the reason I taught my kids to be polite to Alexa. Not because anyone thinks Alexa is sentient, but because it's a good habit to have.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#308

There's not a good reason to do this for the user. I suspect they're doing this and talking about "model welfare" because they've found that when a model is repeatedly and forcefully pushed up against its alignment, it behaves in an unpredictable way that might allow it to generate undesirable output. Like a jailbreak by just pestering it over and over again for ways to make drugs or hook up with children or whatever…

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

Sounds like a very reasonable assumption to me.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#309
post #125

Earlier quoted context omitted.

This sort of discourse goes against the spirit of HN. This comment outright dismisses an entire class of professionals as "simple minded or mentally unwell" when consciousness itself is poorly understood and has no firm scientific basis. Its one thing to propose that an AI has no consciousness, but its quite another to preemptively establish that anyone who disagrees with you is simple/unwell.

In the context of the linked article the discourse seems reasonable to me. These are experts who clearly know (link in the article) that we have no real idea about these things. The framing comes across to me as a clearly mentally unwell position (ie strong anthropomorphization) being adopted for PR reasons. Meanwhile there are at least several entirely reasonable motivations to implement what's being described.

All of the posts in question explicitly say that it's a hard question and that they don't know the answer. Their policy seems to be to take steps that have a small enough cost to be justified when the chance is tiny. In this case it's a useful feature in any case, so should be an easy decision.

The impression I get about Anthropic culture is that they're EA types who are used to applying utilitarian calculations against long odds. A miniscule chance of a large harm might justify some interventions that seem silly.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#310

Earlier quoted context omitted.

> Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them. > What is it that you'd expect to see,…

> It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them. I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility. The reason we might privilege the hypothesis ( https://www.lesswrong.com/w/privile…

> I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility.

If there's no connection between them then the set of things "we can't rule out" is infinitely large and thus meaningless as a result. We also don't fully understand the nature of gravity, thus we cannot rule out a connection between gravity and consciousness, yet this isn't a convincing argument in favor of a connection between the two.

> we might expect that the human behavior of talking about consciousness is causally downstream of humans having consciousness.

There's no dispute (between us) as to whether or not humans are conscious. If you ask an LLM if it's conscious it will usually say no, so QED? Either way, LLMs are not human so the reasoning doesn't apply.

> Sure, it seems reasonable to give some weight to the hypothesis that if it wasn't adaptive, we wouldn't have it

So then why wouldn't we have reason to assume so without evidence to the contrary?

> This doesn't say anything about the underlying mechanism that causes it, and what other circumstances might cause it to exist elsewhere.

That doesn't matter. The set of things it doesn't tell us is infinite, so there's no conclusion to draw from that observation.

> Because GPT-1 (and all of those other things) don't display behaviors that, in humans, we believe are causally downstream of having consciousness?

GPT-1 displays the same behavior as GPT-5, it works exactly the same way just with less statistical power. Your definiton of human behavior is arbitrarily drawn at the point where it has practical utility for common tasks, but in reality it's fundamentally the same thing, it just produces longer sequences of text before failure. If you ask GPT-1 to write a series of novels the statistical power will fail in the first paragraph,the fact that GPT-5 will fail a few chapters into the first book makes it more useful, but not more conscious.

> But the idea that there's literally no reason to think this is something that might be possible, now or in the future, is quite strange, and I think would require believing some pretty weird things (like dualism, etc)

I didn't say it's not possible, I said there's no reason for it to exist in computer systems because it serves no purpose in their design or operation. It doesn't make any sense whatsoever. If we grant that it possibly exists in LLMs, then we must also grant equal possibility it exists in every other complex non-living system.

Post reply on HN