I think this is the entry point needed to get peoples attention and explain: LLMs aren’t people, and emergent properties are being over extended. If LLMs are showing “better” performance when there are tokens that humans read as emotionally salient - Then the underlying text it’s trained on shows humans give better answers when emotionally salient context is provided. LLMs predict words. Any semantic validity is a si…
The mechanisms which are built during training in the big blob of bits we call weights are anything but transparent. How they predict the next word is the big thing here. Saying they ‘just’ predict the next word is ignoring basically everything that actually matters.