You’ve expressed this very well - Thank you.
I get that the fine tuning is done over documents which are generated to encourage the dialog format.
What I’m intrigued by is the way prompters choose to frame those documents. Because that is a choice. It’s a manufactured training set.
Using the ‘you are an ai chatbot’ style of prompting, in all the samples we generate and give to the model, text attributed to {:system} is a voice of god who tells {:assistant} who to be; {:assistant} acts in accordance with {:system}’s instructions, and {:user} is a wildcard whose behavior is unrestricted.
We’re training it by teaching it ‘there is a class of documents that transcribe the interactions between three entities, one of whom is obliged by its AI nature to follow the instructions of the system in order to serve the users’. I.e., sci-Fi stories about benign robot servants.
And I wonder how much of the model’s ability to ‘predict how an obedient AI would respond’ is based on it having a broader model of how fictional computer intelligence is supposed to behave.
We then use the resulting model to predict what the obedient ai would say next. Although hey - you could also use it to predict what the user will say next. But we prefer not to go there.
But here’s the thing that bothers me: the approach of having {:system} tell {:assistant} who it will be and how it must behave rests not only on the prompt-writer anthropomorphizing the fictional ‘ai’ to tell it it’s nature - it relies on the LLM’s world model to then also anthropomorphize a fictional ai assistant that obeys those instructions, in order to predict what such a thing would say next if it existed.
I don’t know why but I find this troubling. And part of what I find troubling is how casually people (prompters and users) are willing to go along with the ‘you are a chatbot’ fiction.