Earlier quoted context omitted.
The question being originally asked is whether "next-token predictor" is the right mental model for an RL-trained model, and I think the answer is no - not only is it not technically correct, but it is a misleading mental model and will lead to incorrect expectations/explanations of what the model is doing. Calling the base model a next token predictor is accurate since it is literally making a prediction and being g…
> The question being originally asked is whether "next-token predicton" is the right mental model for an RL-trained model, Regardless, the statement being challenged here is "still next-token prediction, then". > and I think the answer is no - not only is it not technically correct It is correct. RL simply adjusts weights - with no effect beyond an equivalent adjustment to the corpus itself. Hence "next-token predict…
To me it is an accurate description of the algorithm. And a sufficient explanation for the behaviour. So I don't need it or anything else as a mental model.
I accept this does not suffice for people who cannot comprehend the huge amount of processing and data the empowers it. Lacking a factual understanding, they reach for any mental model as a kind of superstition.
It is sufficiently advanced technology which to many is indistinguishable from magic. This disguises its limitations and enables its limitless false promotion to the gullible, being the reason it is so dangerous to individuals and society.