Earlier quoted context omitted.
> LLMs clearly develop internal representations, this is an empirical fact. you've offered an anecdote. Some LLMs (more generally, this type of neural network model) will generate configurations that can be understood as a representation; others will not. The fact that the authors were able to find an apparent model in a heavily rule-based system is not incredibly surprising but offers little clue about whether this…
Okay, I think we are in agreement. Let's say current LLM architectures are clearly capable of developing internal representations, not just learning surface statistics, and it can execute algorithmic computation as complex as computing Othello board states, and it can completely generalize out of distribution thanks to such algorithmic computation. (One experiment was to completely eliminate any Othello games startin…
LLMs are trained on the "rules" of human written language, which are sufficiently distinct from anything in the actual world that find world modelling would, indeed, be a surprise to me.