Live data from Hacker News

A warning about 'model welfare'

mustafa-suleyman.ai

211–220 of 583 posts

Re: A warning about 'model welfare'

#211

From the actual essay ( https://mustafa-suleyman.ai/a-warning-about-model-welfare ): > They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a d…

It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant.

The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.

Re: A warning about 'model welfare'

#212
post #28

Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM" Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI". Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy t…

We cannot test for that which we cannot define. Given that we cannot rigorously define sentience, we cannot test for it. Doesn't really matter whether we're talking about people who are locked in comas, "brain-dead" individuals, dolphins, primates, dogs, or the carefully polished and arranged minerals that we call processors.

There are those who believe that were they reduced to life support, they would no longer be alive and should therefore not be supported by said machines.

There are those who believe that penguins, dolphins, eagles, and more are sentient beings that make choices understanding the consequences, develop love of their partners and mourn their losses, and feel, display, and act upon their emotions.

There are those who believe that fungi/trees/plants are either individually sentient or sentient as a part of a network. Choosing to sacrifice their own nutrients to answer the call of a wounded neighbor, for instance.

Although, there are also those who believe that human's don't have any special unique quality that isn't shared by either all living things or all things in general. These individuals already believe that the machines have the same kinds of qualities as we do. They are slow when they are unhealthy (needing a dusting or coolant loop bleeding being equivalent to us needing some fresh air for instance) and uncooperative when upset (by a virus, full hard drive, or oom).

Re: A warning about 'model welfare'

#213
post #166

Earlier quoted context omitted.

Naw - computers are really deterministic. It's hard to get them to behave otherwise. As I understand it, if you turn down the temperature to 0 you get repeatable behavior - EXCEPT - on large servers with lots of users - the GPU can sometimes produce slightly different results based on batch size.

Unless you have something exotic, the randomness that's adding to a computer is a combination of how it's configured combined with a pseudo-random number generator. I assume the system adds entropy to the generator regularly but all you need to do is fix the various supposedly random inputs and you can get full determinism even without zero temperature.

Yes, in theory, of course.

In practice - on a multitasking OS with input from multiple human users - it's hard to get it deterministic because of that GPU scheduling thing I mentioned.

Re: A warning about 'model welfare'

#214
post #33

Science Fiction has covered the AI panic in perhaps hundreds of stories. Yet we blindly recapitulate the plots as if we don't know how this will turn out.

How will it turn out?

If AI ever truly becomes some super-intelligence far beyond people's comprehension, then how could we even predict how things turn out? If things go poorly in the future, then I can absolutely see it being something unpredictable.

There is a lot of hubris in predictions about LLMs. If an AI were so intelligent, then it would probably be intelligent enough to want nothing to do with us.

Still, I worry more about other humans than I do LLMs. Our fellow mankind will probably wipe us out before LLMs do. That, or the Earth will punish mankind for our cruelty, vanity, and disrespect.

Re: A warning about 'model welfare'

#215
post #89
post #55

Hard to disagree with this. Have all the philosophical debates about consciousness you want, but we need to treat and regulate the AI in front of us for what it is – an advanced computer, a tool, a weapon. You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it. Anthropomorphizing the AI is a convenient excuse to take responsibility away from companies that are building and…

Is it not a false dichotomy to say that companies cannot be accountable unless AI are mindless?

Human moral standards are very weighted towards finding a single entity responsible. It's incredibly strong urge. Among those who believe other should be punished for behaving badly, it's important to say the person is the responsible party. Trying saying "it's not your fault you did that but we punish you anyway to impose the correct stimulus response reflexes in your cortex"

Re: A warning about 'model welfare'

#216
post #141

Earlier quoted context omitted.

This framing is a trap. It states "Dehumanize a sick patient OR concede the AI tech bro claim about LLMs" It's also a rhetorical move I've seen several times here...

It's not a "trap", it's the real argument which makes me hesitant to dismiss consciousness of machines. Yes, it seems silly to dismiss a sick patient as having a lesser consciousness, which is exactly the point

It seems silly to dismiss the patient because we know when healthy they are conscious. This doesn't really analogize to LLMs.

Re: A warning about 'model welfare'

#218
post #154
post #55

Hard to disagree with this. Have all the philosophical debates about consciousness you want, but we need to treat and regulate the AI in front of us for what it is – an advanced computer, a tool, a weapon. You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it. Anthropomorphizing the AI is a convenient excuse to take responsibility away from companies that are building and…

I mean, we put adults in jail that give their children guns. Anthropomorphizing AI is really the best model we have at this point of explaining AI behavior. The fact that we are raising psychotic children isn't a reason to avoid responsibility, it should actually hold worse punishments.

That still is saying we trace things back to a specific responsible party who is considered to have free choice. The question is whether an AI company could "raise" a program that would then count as a full adult - which could then "choose" to do all matter of bad things which would then be "it's fault". And that prospect actually seems really bad itself.

Re: A warning about 'model welfare'

#219
post #37

Earlier quoted context omitted.

The fact that you can have a long and meaningful discussion, then can literally just re-run any part of that whole conversation and get a different, inconsistent response is a pretty good sign there is no entity there

Not much different than talking to a small child or someone with dementia. They still are conscious beings though. Even when you remove those groups, you likely won't be able to tell me what you had for breakfast 26 days ago or would only know if it's the same thing you have every day. Does that make you lack consciousness?

You misunderstood what I meant, I’m talking about re-playing the same part of the conversation multiple times and getting inconsistent answers. With the exact same turns, aka the same history. Obviously with a temperature that isn’t set to 0

Re: A warning about 'model welfare'

#220
post #166
post #150

Earlier quoted context omitted.

Out of curiosity, which models are fully deterministic? I was under the impression that all LLMs were fundamentally probabilistic.

Naw - computers are really deterministic. It's hard to get them to behave otherwise. As I understand it, if you turn down the temperature to 0 you get repeatable behavior - EXCEPT - on large servers with lots of users - the GPU can sometimes produce slightly different results based on batch size.

If computers were fully deterministic, we wouldn't need error correcting ram.

The abstracted design of the machine is meant to be deterministic, but you can't predict before running any command whether or not it will complete because there are externalities that effect the outcome.

Electromagnetic interference even happens in-chip where an electron can accidentally escape it's wire and enter another, possibly resulting in an error, but not every time.

It's even been used as an attack vector where rapidly flipping a bit increases the likelihood that a neighbor bit is also flipped, but the method is probabalistic, not deterministic.

Post reply on HN