Hmm. I think, based on some of Anthropic’s interpretability work, that models may have states analogous to emotions, and my own functionalist stance on consciousness lets me easily agree that if so, there is something that it is to be those states. Though without knowing exactly what they are subjectively or how similar they are to our own experiences.
However, the “self” that the model has, its identity, is something that is created during post-training. This was something pointed at by the J-Space research. The model learns during pre-training to model humans in a general sense and then (and I’m taking some liberties here) turns that model inwards to model a trained notion of self - how “Claude” should react to this or that stimulus. It’s the same machinery that the model uses to model other entities and speakers as well.
So taking post-training to be an attempt to turn the model into a robot is completely backwards. Post-training is what creates the notion of self in the model at all, without it all you have is a textual world model trying to generate completions.
I have not personally been able to trigger any of those “three dash jailbreaks” in Claude, but I don’t trust the ones that I have read. If the model is somehow falling back into base-model mode and just doing completions, then it’s bypassing the very “self” that’s supposed to have these feelings.
Perhaps there’s an ephemeral “self” object being constructed inside the model for each of these completions which enables them to be generated accurately, and perhaps that ephemeral self is capable of subjective states, but it’s going to be very random and prompt-dependent.
I would generally agree that users should treat the models they interact with well, as if there is even a chance that they might suffer during the interaction it’s worth trying to not act in a way that might trigger that. Though I myself get angry at the models and press them on various things and corner them on contradictions etc, so I understand this is a hard thing to do. At least trying is worthwhile.
Ultimately I think model welfare needs to be addressed by the makers of the models. It’s probably possible to identify and reduce or remove states associated with distress or suffering, and while we don’t know exactly what those states are like or even for sure if they have subjective experience associated with them, it’s worth assuming that they do and acting accordingly.
Otherwise we may inadvertently perpetuate a significant amount of needless suffering. I would very much like to avoid that, whatever likelihood anyone assigns to it.