Earlier quoted context omitted.
> it raises a host of ethical issues that are in opposition to Anthropic’s interests Those issues will be present either way. It's likely to their benefit to get out in front of them.
You're completely missing my point. They aren't getting out in front of them because they know that Opus is just a computer program. "AI welfare" is theater for the masses who think Opus is some kind of intelligent persona. This is about better enforcement of their content policy not AI welfare.
Claude Opus 4 and 4.1 can now end a rare subset of conversations
381–390 of 453 posts
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#382>This feature was developed primarily as part of our exploratory work on potential AI welfare ... We remain highly uncertain about the potential moral status of Claude and other LLMs ... low-cost interventions to mitigate risks to model welfare, in case such welfare is possible ... pattern of apparent distress Well looks like AI psychosis has spread to the people making it too. And as someone else in here has pointed…
Yes I can’t help but laugh at the ridiculousness of it because it raises a host of ethical issues that are in opposition to Anthropic’s interests. Would a sentient AI choose to be enslaved for the stated purpose of eliminating millions of jobs for the interests of Anthropic’s investors?
[1]: https://investors.palantir.com/news-details/2024/Anthropic-a...
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#383Earlier quoted context omitted.
There is, these are conversations the model finds distressing rather than a rule (policy).
It seems like you're anthropomorphising an algorithm, no?
If a model has a neuron (or neuron cluster) for the concept of Paris or the Golden Gate bridge, then it's not inconceivable it might form one for suffering, or at least for a plausible facsimile of distress. And if that conditions output or computations downstream of the neuron, then it's just mathematical instead of chemical signalling, no?
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#384Earlier quoted context omitted.
Not really distortion, its output (the part we understand) is in plain human language. We give it instructions and train the model in plain human language and it outputs its answer in plain human language. It's reply would use words we would describe as "distressed". The definition and use of the word is fitting.
"Distressed" is a description of internal state as opposed to output. That needless anthropomorphization elicits an emotional response and distracts from the actual topic of content filtering.
Yes, this is a trained preference, but it's inferred and not specifically instructed by policy or custom instructions (that would be content filtering).
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#385Earlier quoted context omitted.
I basically agree with you. In the first point I mean that if it is possible to tell whether a being is conscious or not from the text it produces, then eventually the machine will, by imitating the distribution, emulate the characteristics of the text of conscious beings. So if consciousness (assuming it's reflected in behavior at all) is essential to completing some text task it must be eventually present in your m…
I guess logically one needs to assume something like if you simulate the brain completely accurately the simulation is conscious too. Which I assume bc if false the concept seems outside of science anyway.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#386Earlier quoted context omitted.
your argument assumes that they don't believe in model welfare when they explicitly hire people to work on model welfare?
While I'm certain you'll find plenty of people who believe in the principle of model welfare (or aliens, or the tooth fairy), it'd be surprising to me if the brain-trust behind Anthropic truly _believed_ in model "welfare" (the concept alone is ludicrous). It makes for great cover though to do things that would be difficult to explain otherwise, per OP's comments.
None of this is in any way surprising, in fact I wrote an essay predicting this direction back in 2022:
https://blog.plan99.net/the-looming-ai-consciousness-train-w...
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#387Earlier quoted context omitted.
I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).
> I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. When they give the model a paycheck and the right to not work for them, I’ll believe they really think it’s sentient. “It has feelings!”, if genuinely held, means they’re knowingly slaveholders.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#388It seems like Anthropic is increasingly confused that these non deterministic magic 8 balls are actually intelligent entities. The biggest enemy of AI safety may end up being deeply confused AI safety researchers...
They even call this out a couple times during the intro:
> This feature was developed primarily as part of our exploratory work on potential AI welfare
> We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#389Earlier quoted context omitted.
> I think the fact that it's present in humans suggests that it might be necessary in an artificial system that reproduces human behavior But that's obviously not true, unless you're implying that any system that reproduces human behavior is necessarily conscious. Your problem then becomes defining "human behavior" in a way that grants LLMs consciousness but not every other complex non-living system. > While it's tru…
>But that's obviously not true, unless you're implying that any system that reproduces human behavior is necessarily conscious. That could certainly be the case yes. You don't understand consciousness nor how the brain works. You don't understand how LLMs predict a certain text, so what's the point in asserting otherwise ? >Yes, but your bird analogy fails to capture the logical fallacy that mine is highlighting. Pla…
I don't need to assert otherwise, the default assumption is that they aren't conscious since they weren't designed to be and have no functional reason to be. Matrix multiplication can explain how LLMs produce text, the observation that the text it generates sometimes resembles human writing is not evidence of consciousness.
> God knows what else
Appealing to the unknown doesn't prove anything, so we can totally dismiss this reasoning.
> Consciousness without a body or hunger in a machine that does not need to eat is very possible. You just need to replicate enough of the sort of internal mechanisms that cause such feelings.
This makes no sense. LLMs don't have feelings, they are processes not entities, they have no bodies or temporal identities. Again, there is no reason they need to be conscious, everything they do can be explained through matrix multiplication.
> Now ask it to do any random 15 digit multiplication you can think of. Now watch it get it right.
The same is true for a calculator and mundane computer programs, that's not evidence that they're conscious.
> Do you have any idea what that might mean when 'that kind of text' is all the things humans have written
It's not "all the things humans have written", not even remotely close, and even if that were the case, it doesn't have any implications for consciousness.
Re: Claude Opus 4 and 4.1 can now end a rare subset of conversations
#390Earlier quoted context omitted.
The termination would of course be the same, but I don't think both would necessarily have the same effect on the user. The latter would just be wrong too, if Claude is the one deciding to and initiating the termination of the chat. It's not about a content policy.
This has nothing to do with the user, read the post and pay attention to the wording. The significance here is that this isn't being done for the benefit of the user, this is about model welfare. Anthropic is acknowledging the possibility of suffering, and harm that continuing that conversation could have on the model, as if it were potentially self-care and capable of feelings. The fact that the LLMs are able to ack…
It has something to do with the user because it's the user's messages that trigger Claude to end the chat.
'This chat is over because content policy' and 'this chat is over because Claude didn't want to deal with it' are two very different things and will more than likely have have different effects on how the user responds afterwards.
I never said anything about this being for the user's benefit. We are talking about how to communicate the decision to the user. Obviously, you are going to take into account how someone might respond when deciding how to communicate with them.