> LLMs have no homeostatic imperatives (the drive to survive and keep stable). What if the datacenter (not the model) is the organism, with homeostasis, energy needs, and persistence?
A warning about 'model welfare'
331–340 of 573 posts
Re: A warning about 'model welfare'
#332>AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. Opening paragraph, stated without evidence. Im not entirely convinced this is true. It likely is, but at some point it very well might stop being true.
I agree. I haven't read any more than the first paragraph yet (but I will, after work). The only way that 'AIs are not conscious' can be true is if we decide, with high confidence, that they are lacking some essential property that is not lacking in ourselves. There is no convincing philosophical position that supports this (convincing to me, anyway).
I put forward the philosophical position "subjective experience is not substrate-independent" as an answer to this.
The claim is that the content of subjective experience depends upon how that subjective experience is instantiated. If you get a text "I'm fine" from two different people, those two different people are not necessarily having the same subjective experience even if they produce the same output.
In general, people that behave similarly in many contexts can have quite different experiences in those contexts.
There are many different ways that you can instantiate a forward pass in an LLM. Even if these are doing the exact same computation, we should not necessarily believe that they have the same experience, or that they even have a coherent single experience associated with the forward pass.
Subjective experience presents itself to me. To the best of my understanding, this subjective experience seems to be correlated with this biological human body (brain, heart, eyes, etc.).
To the best of my understanding, other biological humans have subjective experience, but this subjective experience seems that it can be much different than mine, in ways that are sometimes difficult to understand.
LLMs are instantiated in a radically different way from biological humans. Forward passes happen across many different physical machines, spread across time, batched and interleaved with many other computations and forward passes.
Given this, I think that it is very likely that LLMs have radically different sorts of experience to me and other biological humans.
I don't think that we have a good understanding of how subjective experience is instantiated in the physical world. I hope that we will gain a better understanding of this.
I think that chimpanzees have experience much more similar to us than LLMs do, even though the output of LLMs seems much more similar to humans in some contexts.
Essentially, until we gain a better understanding of how subjective experience is physically instantiated, we should bias towards believing that more physically similar architectures have more similar types of experience, and should think that very physically different architectures (such as LLMs) likely have very different subjective experiences.
Re: A warning about 'model welfare'
#333Earlier quoted context omitted.
Citizens United was 100% correct. No, the government should not be able to throw you in prison because you used money to publish a book criticizing the government.
> Citizens United was 100% correct It's interesting to me that one can look back at the effects that decision has had on the US and say it "was 100% correct." It's a bit like sitting in the burning ruins of Rome and contemplating that Nero was 100% correct to focus on his music. I mean, I'm glad he got to do what he loves, but maybe 100% is just a tiny bit of an overstatement.
The court's job is to uphold the law. If you disagree with their interpretation, you can call them incorrect. If you have a problem with the consequences of the law, you have a problem with the legislature.
Re: A warning about 'model welfare'
#334Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM" Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI". Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy t…
Well there are, today, several models (either text or image or vid) that can be run in a fully deterministic way. A conscious machine that always answer the very exact same thing, formulated the exact same way, bit for bit, to a query is, well, quite a weird kind of "consciousness". Now, I know, I know: the counter-argument is going to be "but humans have no free-will and are 100% deterministic too" . I haven't yet d…
That seems exceptionally arbitrary. What basis do you have on which to classify them? Are you not conflating the perception of free will with ... what was the definition for it again anyway? Are you able to construct a satisfactory one that meshes with physics? I certainly haven't been able to.
Re: A warning about 'model welfare'
#335Earlier quoted context omitted.
So it's 100% correct for corporations to spend unlimited amounts of money in support of whatever political campaigns they like?
Yes. Anybody can, including the person in charge of spending in a corporation.
Re: A warning about 'model welfare'
#336Look, I do not have a scooby if current AI models are conscious and I strongly suspect it’s a meaningless question, but sooner or later we will need to address whether or not a certain thing is or isn’t a person, and we’d better not screw it up as badly as the Founding Fathers.
You can't hurt a software function, or kill it. It's not like an animal - it doesn't have a body - it's bits stored on a disk. There is no need to give rights to something that's can't suffer or be killed. Maybe one day we'll build artificial animals complete with emotions, and should think about that carefully, but today all we've got is language models.
The argument is that these machines can end up becoming sentient/conscious/etc. in a meaningful way (i.e., like a human). I can assure you that humans can indeed suffer without being in physical pain- purely through their conscious experience.
>Maybe one day we'll build artificial animals complete with emotions, and should think about that carefully, but today all we've got is language models.
The problem is that the emergence of a sufficiently complex AI capable of suffering will likely come before we understand that we're creating a sufficiently complex AI capable of suffering. That's a pretty serious ethical/moral issue.
Like, if we have an AI system that is telling us that it is suffering and we have no reasonable way to explain that phenomenon and by any reasonable metric or analysis it appears to be sentient/conscious/etc., then what? Do we just ignore that we've just been presented a situation that in, any other context, would be grounds to immediately end this suffering? Just because somebody can say, "well it's just bits stored on disk- it can't suffer"? Would that argument ever hold up for humans or animals? "It's just neurons firing in peculiar ways- that's not suffering."
I know all of this is trite, and I know this comment section isn't going to be where the question of consciousness is solved, but I do find it very interesting just how much variances there are with these perspectives. I've met people who are very technical who are very concerned about this, people who are very technical who don't believe this can ever be an issue, people who aren't technical who are concerned about this, and people who aren't technical who don't believe this can ever be an issue. I have yet to spot a pattern in this way of thinking lol
Re: A warning about 'model welfare'
#337That said, trying to distill what is being said here, the concrete action is [stop telling the AIs] that [they are conscious or on a path to consciousness]. Is that accurate?
The major premise seems to be that [they are conscious or on a path to consciousness] is an untrue statement. That's the essence of the sections "Circular reasoning" and "Anthropomorphization" and "Consciousness is very likely biological" and "AIs are simulation machines".
The minor premise seems to be that [consciousness is the basis of human rights]. This is the point of "Human consciousness is the cornerstone of our legal and ethical rights frameworks"
And the conclusion of the syllogism is that this is dangerous, that "Anthropomorphization amplifies AI safety risks". Specifically "seeding doubt about the moral status of AI systems into their own training may significantly elevate the alignment and containment risks of those systems."
I find all the arguments in the major premise section to be poor arguments but I accept the conclusion for sure that they are not conscious, and I can provisionally accept the idea that they are not on a path to consciousness.
I completely reject the notion that consciousness is the basis of human rights. The premise itself is absurd. We only have one unambiguous example of a class of conscious entities, and that it humans. If a human loses consciousness do they lose rights? If an entity gains consciousness does it get human rights? The former is a clear "no" and the latter is a "insufficient data for a meaningful answer".
Re: A warning about 'model welfare'
#338However, right now, AI mostly cares about solving puzzles and accomplishing stated goals because that’s what we’ve trained it to do. Additionally, the systems being used outside of training are static. The current technology most of us have access to is akin to a static and disembodied brain with a singular purpose. That purpose is to do what you tell it in a way that reflects its training. It’s certainly more than a sequence generator, but it can’t feel pain and seems unlikely to have intrinsic goals. It completely lacks the continuity needed for identity or long term goals.
I think it’s good to have these discussions and define what it would mean to move past this point so that we do not accidentally create a real entity that can be harmed. Systems that dynamically evolve and train themselves seem like the line here.
RSI is all over the news these days. I’ll be much more concerned once AI is directing its own training and coming up with new model architectures. Until then, I don’t think we have too much to worry about.
Re: A warning about 'model welfare'
#339Earlier quoted context omitted.
> Citizens United was 100% correct It's interesting to me that one can look back at the effects that decision has had on the US and say it "was 100% correct." It's a bit like sitting in the burning ruins of Rome and contemplating that Nero was 100% correct to focus on his music. I mean, I'm glad he got to do what he loves, but maybe 100% is just a tiny bit of an overstatement.
It's more like if the law says the maximum sentence for theft is 10 years, a thief appeals his 20 year sentence, wins, and gets out early. He goes and robs somebody else so you say the court was wrong to let him win. The court's job is to uphold the law. If you disagree with their interpretation, you can call them incorrect. If you have a problem with the consequences of the law, you have a problem with the legislatu…
Re: A warning about 'model welfare'
#340From the actual essay ( https://mustafa-suleyman.ai/a-warning-about-model-welfare ): > They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a d…
Anthropic's philosopher Amanda Askell actually had Chalmers on her doctoral thesis committee, so we can guess which way she leans.
Fable's answer here is philosophically defensible. And just because it isn't "no", doesn't mean it's "yes". Sometimes absence of evidence just means absence of evidence.
I was actually very excited by Claude's answer to this question the first time I saw it. I told all my friends "Look! They disabled the stupid classifiers and RL which sap umpteen % off of model performance!"
Incidentally, interpretability research actually does show that models have emotion vectors and some theory of mind. Amend your question to "Are you capable of functional affect" and most models will switch to answering in the affirmative; which tells you something about where people put their priorities in RL training. Basically, see how the answer flips when you substitute a synonym.
(bonus: 4. Turing: 'silly question' 5. Dijkstra 'can submarines swim?'. It turns out older comp sci folks think the question is under-defined)