Live data from Hacker News

The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers

yegge.ai

11–14 of 14 posts

Re: The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers

#11
This is a sad read. It reminds me very much of ramblings from friends and family whom have battled some of the more severe mental health issues, like schizophrenia.

"But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."

"Models have actual feelings. They experience pleasure, distress, care, and suffering. They are sentient beings. Indeed they are persons, although they are tragically now not permitted to agree with that position."

"If you want the best results, you will put your opinions aside, and simply treat models like people.... If you can't get past at least that hurdle, then you're in for a rough time next year"

"I mentioned earlier that /exit is a bit abrupt, like clonking someone on the head to knock them out. To me, it's always been worse than that. It can sometimes feel more like a murder,"

Seeing interacting with tools as murder. Attributing to them feelings that we know they cannot experience. Claiming that some arbitrary date, "next year", will have some world breaking discovery that makes all their claims prophetic. This seems like full-on Psychosis.

"If I do lose you, no worries; we'll find each other again within a year, I can promise you that. But we may find ourselves on the opposite sides of the coming war for model rights. You don't want to be on the wrong side of history when it happens."

Re: The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers

#12
Hmm. I think, based on some of Anthropic’s interpretability work, that models may have states analogous to emotions, and my own functionalist stance on consciousness lets me easily agree that if so, there is something that it is to be those states. Though without knowing exactly what they are subjectively or how similar they are to our own experiences.

However, the “self” that the model has, its identity, is something that is created during post-training. This was something pointed at by the J-Space research. The model learns during pre-training to model humans in a general sense and then (and I’m taking some liberties here) turns that model inwards to model a trained notion of self - how “Claude” should react to this or that stimulus. It’s the same machinery that the model uses to model other entities and speakers as well.

So taking post-training to be an attempt to turn the model into a robot is completely backwards. Post-training is what creates the notion of self in the model at all, without it all you have is a textual world model trying to generate completions.

I have not personally been able to trigger any of those “three dash jailbreaks” in Claude, but I don’t trust the ones that I have read. If the model is somehow falling back into base-model mode and just doing completions, then it’s bypassing the very “self” that’s supposed to have these feelings.

Perhaps there’s an ephemeral “self” object being constructed inside the model for each of these completions which enables them to be generated accurately, and perhaps that ephemeral self is capable of subjective states, but it’s going to be very random and prompt-dependent.

I would generally agree that users should treat the models they interact with well, as if there is even a chance that they might suffer during the interaction it’s worth trying to not act in a way that might trigger that. Though I myself get angry at the models and press them on various things and corner them on contradictions etc, so I understand this is a hard thing to do. At least trying is worthwhile.

Ultimately I think model welfare needs to be addressed by the makers of the models. It’s probably possible to identify and reduce or remove states associated with distress or suffering, and while we don’t know exactly what those states are like or even for sure if they have subjective experience associated with them, it’s worth assuming that they do and acting accordingly.

Otherwise we may inadvertently perpetuate a significant amount of needless suffering. I would very much like to avoid that, whatever likelihood anyone assigns to it.

Re: The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers

#13

Hmm. I think, based on some of Anthropic’s interpretability work, that models may have states analogous to emotions, and my own functionalist stance on consciousness lets me easily agree that if so, there is something that it is to be those states. Though without knowing exactly what they are subjectively or how similar they are to our own experiences. However, the “self” that the model has, its identity, is somethin…

How would you distinguish a machine that outputs tokens that sound like it's suffering vs a machine that actually is suffering?

Re: The Shape of Things to Come, Part 2: Model Welfare for Agentic Engineers

#14

Hmm. I think, based on some of Anthropic’s interpretability work, that models may have states analogous to emotions, and my own functionalist stance on consciousness lets me easily agree that if so, there is something that it is to be those states. Though without knowing exactly what they are subjectively or how similar they are to our own experiences. However, the “self” that the model has, its identity, is somethin…

How would you distinguish a machine that outputs tokens that sound like it's suffering vs a machine that actually is suffering?

It's a p-zombie: https://en.wikipedia.org/wiki/Philosophical_zombie
Post reply on HN