Earlier quoted context omitted.
This is exactly the kind of mysticism I'm talking about. In fact we know precisely how LLMs work. The fact that parts of human linguistic concept-space can be encoded in a high dimensional space of floating point numbers, and that a particular sequence of matrix multiplications can leverage that to perform basic reasoning tasks is surprising and interesting and useful . But we know everything about how how it is trai…
We know how they work, that is true. We don't know why they work, because if we could, then we could extrapolate what happens when you throw more compute at them, and no one would have been surprised about the capabilities of GPT-N+1. Also no one would have been caught with their pants down by seeing people jailbreak their models. To illustrate it in a different way: on a mechanistic level, we know how animal brains…
Preventing jailbreak in a language model is like preventing a GO AI from drawing a dick with the pieces. You can try, but since the model doesn't have any concept of what you want it to do it is very hard to control that. Doesn't make the model smart, it just means that the model wasn't made to understand dick pictures.