I think this is the entry point needed to get peoples attention and explain: LLMs aren’t people, and emergent properties are being over extended. If LLMs are showing “better” performance when there are tokens that humans read as emotionally salient - Then the underlying text it’s trained on shows humans give better answers when emotionally salient context is provided. LLMs predict words. Any semantic validity is a si…
The “statistical parrot” assertion is pretty thoroughly disproven by this point, but suppose we ignore the literature and just assume it’s true: what does it matter? “Real” people are time bombs too, for instance. Is there some predictive power that we gain by reducing LLM skills to mere token production side effects?
If those skills were real, why do they fizzle out on production data ?