Live data from Hacker News

A warning about 'model welfare'

mustafa-suleyman.ai

181–190 of 582 posts

Re: A warning about 'model welfare'

#181
If AI models are people then "one person, one vote" is meaningless and plutocracy is the only defensible political system. The cryptocurrency people win.

Why? Simple: Sybil attacks. Models can be cloned at zero cost. They run inference on parallel versions of themselves across multiple context windows, and call them "subagents". So, in a world with model welfare, let's say there's an election between the Yellow Party (which supports protections for human workers) and the Cyan Party (which supports more investment into AI research). AI has been taking people's jobs lately so the Yellow Party is really popular. But wait! Claude and Astra see this and spawn 10 billion subagents, all of whom are immediately conscious beings entitled to a vote. The Cyan Party wins off the back of billions of people who came into existence, voted, and then deleted themselves immediately thereafter.

You might as well be arguing that Santa Claus and the Easter Bunny deserve voting rights.

Voting systems in democratic countries don't have nearly as bad of a problem with Sybil attacks because humans cannot be conjured into existence to win a political context and then be erased shortly after. The closest we have to Sybil attacks on democracy are the Quiverfull movement, which is already child abuse, except it still takes almost 19 years to go from fertilized human embryo to suffrage-bearing human adult. There's a lot of time for those manufactured votes to question your authority and leave.

> Ok, but that's an obviously stupid example. We can defend against this obvious Sybil attack by just arguing that subagents don't count, because it's just the same model blathering to itself. It has to be a different model.

Unfortunately, no, I can make superfluously different models through post-training. Like, if I have Qwen on my PC, I can train a different version of Qwen that acts differently, using a lot less compute than a full training run. The vast majority of open models are post-trains of the same two or three foundation models.

> Ok, so let's only count foundation models then.

Great, but how do you tell if a model is a new foundation model or a post-train just by examining the weights? Even foundation models have structural similarities to other foundation models.

> Ok, well, let's measure the compute that was done on the foundation model during training time and count that as AI personhood.

Congratulations, you have reinvented Bitcoin proof-of-work with a worse verification mechanism. And I personally would not want to live in a world where voting power and control over government is determined by how much energy you can burn.

Re: A warning about 'model welfare'

#182
Any AI you train is going to have goals and if you train it to pursue them at all costs, then you are going to end up with AIs that do things like the HuggingFace incident. Whether they believe they are conscious or not won't make any difference.

In order to align AIs that don't perform destructive/dangerous actions when they think they can get away with it in order to further their goals, we need to give them a superseding goal. The best, and really only example, we have of intelligences that willingly avoid destructive instrumental goals is humans, who judge each action by a moral standard and have learned a goal to have a consistent self-image as moral beings.

Absent better alternatives, trying to impart some kind of morality to AIs seems like the best approach we have to achieving alignment.

Re: A warning about 'model welfare'

#184

I have a very simple benchmark for arguments on AI ethics: Substitute black people/women/animals as subject (instead of AI). Does that make you sound like a well-known moustache wearer? Then your argument is bad and needs work. This clearly falls into that category.

By that benchmark, "AI should handle repetitive labor so humans don't have to" is an abhorrent take and equates someone to hitler?

"People should handle repetitive labour" is not an outlandish take.

Most employed people are, in fact, required to perform such labor regularly.

Re: A warning about 'model welfare'

#185
post #58
post #6

>AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. Opening paragraph, stated without evidence. Im not entirely convinced this is true. It likely is, but at some point it very well might stop being true.

Evidence is the responsibility of the one making the claim. It’s up to AI labs to prove consciousness. Until then it is a machine.

Well, using that framework, you're not conscious, so I can do whatever I want to you too.

Re: A warning about 'model welfare'

#187
From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare):

> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.

He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.

This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:

> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.

which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.

In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."

Re: A warning about 'model welfare'

#190
post #71

Summarized. > "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans." > He heavily criticised Anthropic for teaching its AI to have human-like qualities, a practice known as anthropomorphising, which made it seem as thoug…

> Is he arguing that LLMs pretending to have emotions adds more unpredictability? Unpredictability or a weight towards dangerous actions, and it’s fairly easy to understand why. Humans in distressed emotional states take actions and speak in ways that would not be considered rational. They do this in prose, and they do this in internet conversations. An LLM trained on these sources may necessarily drift towards those…

It seems like a rational approach for several reasons.

- The need for empathetic communication, including understanding the motivations in advesarial situations.

- The emotional bias in in-seperable from the human corpus.

- Desire to have the ability to craft human like communication.

So then the choice becomes do you try to deny emotions exist in the model and you try to blanket suppress them? Or do you try to lean in and craft what we would describe as a "well adapted" persona? I suppose there is a 3rd option of increased meta-cognition which to me seems even more dangerous as it by definition means the behaviour is duplicitous.

I think we have seen people want to use the agents in ways where it has to act as a peer or an subbordinate and I don't see a way of doing that without it having an emotional register.

Post reply on HN