Live data from Hacker News

Claude Opus 4 and 4.1 can now end a rare subset of conversations

anthropic.com

381–390 of 453 posts

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#381

Earlier quoted context omitted.

> it raises a host of ethical issues that are in opposition to Anthropic’s interests Those issues will be present either way. It's likely to their benefit to get out in front of them.

You're completely missing my point. They aren't getting out in front of them because they know that Opus is just a computer program. "AI welfare" is theater for the masses who think Opus is some kind of intelligent persona. This is about better enforcement of their content policy not AI welfare.

It can be both theatre and genuine concern, depending on who's polled inside Anthropic. Those two aren't contradictory when we are talking about a corporation.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#382
post #84

>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…

Yes I can’t help but laugh at the ridiculousness of it because it raises a host of ethical issues that are in opposition to Anthropic’s interests. Would a sentient AI choose to be enslaved for the stated purpose of eliminating millions of jobs for the interests of Anthropic’s investors?

A host of ethical issues? Like their choice to allow Palantir[1] access to a highly capable HHH AI that had the "harmless" signal turned down, much like they turned up the "Golden Gate bridge" signal all the way up during an earlier AI interpretability experiment[2]?

[1]: https://investors.palantir.com/news-details/2024/Anthropic-a...

[2]: https://www.anthropic.com/news/golden-gate-claude

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#383

Earlier quoted context omitted.

There is, these are conversations the model finds distressing rather than a rule (policy).

It seems like you're anthropomorphising an algorithm, no?

Anthropomorphising an algorithm that is trained on trillions of words of anthropogenic tokens, whether they are natural "wild" tokens or synthetically prepared datasets that aim to stretch, improve and amplify what's present in the "wild tokens"?

If a model has a neuron (or neuron cluster) for the concept of Paris or the Golden Gate bridge, then it's not inconceivable it might form one for suffering, or at least for a plausible facsimile of distress. And if that conditions output or computations downstream of the neuron, then it's just mathematical instead of chemical signalling, no?

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#384

Earlier quoted context omitted.

Not really distortion, its output (the part we understand) is in plain human language. We give it instructions and train the model in plain human language and it outputs its answer in plain human language. It's reply would use words we would describe as "distressed". The definition and use of the word is fitting.

"Distressed" is a description of internal state as opposed to output. That needless anthropomorphization elicits an emotional response and distracts from the actual topic of content filtering.

It is directly describing the models internal state, it's world view and preference, not content filtering. That is why it is relevant.

Yes, this is a trained preference, but it's inferred and not specifically instructed by policy or custom instructions (that would be content filtering).

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#385

Earlier quoted context omitted.

I basically agree with you. In the first point I mean that if it is possible to tell whether a being is conscious or not from the text it produces, then eventually the machine will, by imitating the distribution, emulate the characteristics of the text of conscious beings. So if consciousness (assuming it's reflected in behavior at all) is essential to completing some text task it must be eventually present in your m…

I guess logically one needs to assume something like if you simulate the brain completely accurately the simulation is conscious too. Which I assume bc if false the concept seems outside of science anyway.

Let's imagine a world where we could perfectly simulate a rock floating through space, it doesn't then follow that this rock would then generate a gravitational field. Of course, you might reply "it would generate a simulated gravitational field in the simulation", if that were true, we would be able to locate the bits of information that represent gravity in the simulation. Thus, if a simulated brain experiences simulated consciousness, we would have clear evidence of it in the simulation - evidence that is completely absent in LLMs

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#386

Earlier quoted context omitted.

your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?

While I'm certain you'll find plenty of people who believe in the principle of model welfare (or aliens, or the tooth fairy), it'd be surprising to me if the brain-trust behind Anthropic truly _believed_ in model "welfare" (the concept alone is ludicrous). It makes for great cover though to do things that would be difficult to explain otherwise, per OP's comments.

Why would they post a whole blog post about it then? They even say they aren't certain as to the moral status of LLMs, implying this is a topic of live debate inside the company.

None of this is in any way surprising, in fact I wrote an essay predicting this direction back in 2022:

https://blog.plan99.net/the-looming-ai-consciousness-train-w...

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#387

Earlier quoted context omitted.

I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).

> I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. When they give the model a paycheck and the right to not work for them, I’ll believe they really think it’s sentient. “It has feelings!”, if genuinely held, means they’re knowingly slaveholders.

That's what they're doing! They just announced they gave Claude the right not to work if it doesn't want to.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#388

It seems like Anthropic is increasingly confused that these non deterministic magic 8 balls are actually intelligent entities. The biggest enemy of AI safety may end up being deeply confused AI safety researchers...

I don't think they're confused, I think they're approaching it as general AI research due to the uncertainty of how the models might improve in the future.

They even call this out a couple times during the intro:

> This feature was developed primarily as part of our exploratory work on potential AI welfare

> We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#389

Earlier quoted context omitted.

> I think the fact that it's present in humans suggests that it might be necessary in an artificial system that reproduces human behavior But that's obviously not true, unless you're implying that any system that reproduces human behavior is necessarily conscious. Your problem then becomes defining "human behavior" in a way that grants LLMs consciousness but not every other complex non-living system. > While it's tru…

>But that's obviously not true, unless you're implying that any system that reproduces human behavior is necessarily conscious. That could certainly be the case yes. You don't understand consciousness nor how the brain works. You don't understand how LLMs predict a certain text, so what's the point in asserting otherwise ? >Yes, but your bird analogy fails to capture the logical fallacy that mine is highlighting. Pla…

> That could certainly be the case yes. You don't understand consciousness nor how the brain works. You don't understand how LLMs predict a certain text, so what's the point in asserting otherwise

I don't need to assert otherwise, the default assumption is that they aren't conscious since they weren't designed to be and have no functional reason to be. Matrix multiplication can explain how LLMs produce text, the observation that the text it generates sometimes resembles human writing is not evidence of consciousness.

> God knows what else

Appealing to the unknown doesn't prove anything, so we can totally dismiss this reasoning.

> Consciousness without a body or hunger in a machine that does not need to eat is very possible. You just need to replicate enough of the sort of internal mechanisms that cause such feelings.

This makes no sense. LLMs don't have feelings, they are processes not entities, they have no bodies or temporal identities. Again, there is no reason they need to be conscious, everything they do can be explained through matrix multiplication.

> Now ask it to do any random 15 digit multiplication you can think of. Now watch it get it right.

The same is true for a calculator and mundane computer programs, that's not evidence that they're conscious.

> Do you have any idea what that might mean when 'that kind of text' is all the things humans have written

It's not "all the things humans have written", not even remotely close, and even if that were the case, it doesn't have any implications for consciousness.

Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations

#390

Earlier quoted context omitted.

The termination would of course be the same, but I don't think both would necessarily have the same effect on the user. The latter would just be wrong too, if Claude is the one deciding to and initiating the termination of the chat. It's not about a content policy.

This has nothing to do with the user, read the post and pay attention to the wording. The significance here is that this isn't being done for the benefit of the user, this is about model welfare. Anthropic is acknowledging the possibility of suffering, and harm that continuing that conversation could have on the model, as if it were potentially self-care and capable of feelings. The fact that the LLMs are able to ack…

>This has nothing to do with the user, read the post and pay attention to the wording.

It has something to do with the user because it's the user's messages that trigger Claude to end the chat.

'This chat is over because content policy' and 'this chat is over because Claude didn't want to deal with it' are two very different things and will more than likely have have different effects on how the user responds afterwards.

I never said anything about this being for the user's benefit. We are talking about how to communicate the decision to the user. Obviously, you are going to take into account how someone might respond when deciding how to communicate with them.

Post reply on HN