> an ad-hoc created, and later updated or discarded model of a situation in which exist only the elevator, some abstract structure around it, and the laws of physics as I know them from knowledge and experience.
LLMs can do all of this. In fact, multimodality specifically can be shown to improve their physical intuition.
> The difference is not in how sensory input is gathered. The difference is in what that input represents. For the LLM the token represents...the token. That's it. There is nothing else. The token exists for its own sake, and has no information other than itself. It isn't something from which an abstract concept is built, it IS the concept.
The token has structure. The photons have structure. We conjecture that the photons represent real objects. The LLM conjectures (via reinforcement learning) that the tokens represent real objects. It's the exact same concept.
> As a consequence, an language model doesn't understand whether statements are false or nonsensical.
Neither do humans, we just error out at higher complexities. No human has access to the platonic truth of statements.
> So in a language models world a wrong statement can still somehow be "less wrong" than another wrong statement.
Of course, but so with humans? I have no idea what you're trying to say here. As with humans, in a LLM token improbability can derive from lots of different reasons, including world model violation, in-context rule violation, prior improbability and grammatical nonsense. In fact, their probability calibration is famously perfect, until RLHF ruins it. :)
> Bear in mind when I say all this, I don't mean to say (and I think I made that clear elsewhere in the thread) that this mimickry of reasoning isn't useful.
I fundamentally do not believe there is such a thing as "mimickry of reason". There is only reason, done more or less well. To me, it's like saying that a pocket calculator merely "mimicks math" or, as the quote goes, whether a submarine "mimicks swimming". Reason is a system of rules. Rules cannot be "applied fake"; they can only be computed. If the computation is correct, the medium or mechanism are irrelevant.
To quote gwern, if you'll allow me the snark:
> We should pause to note that a Clippy² still doesn’t really think or plan. It’s not really conscious. It is just an unfathomably vast pile of numbers produced by mindless optimization starting from a small seed program that could be written on a few pages. It has no qualia, no intentionality, no true self-awareness, no grounding in a rich multimodal real-world process of cognitive development yielding detailed representations and powerful causal models of reality; it cannot ‘want’ anything beyond maximizing a mechanical reward score, which does not come close to capturing the rich flexibility of human desires, or historical Eurocentric contingency of such conceptualizations, which are, at root, problematically Cartesian. When it ‘plans’, it would be more accurate to say it fake-plans; when it ‘learns’, it fake-learns; when it ‘thinks’, it is just interpolating between memorized data points in a high-dimensional space, and any interpretation of such fake-thoughts as real thoughts is highly misleading; when it takes ‘actions’, they are fake-actions optimizing a fake-learned fake-world, and are not real actions, any more than the people in a simulated rainstorm really get wet, rather than fake-wet. (The deaths, however, are real.)