This highlights one of the types of muddled thinking around LLMs. These tasks are used to test theory of mind because for people, language is a reliable representation of what type of thoughts are going on in the person's mind. In the case of an LLM the language generated doesn't have the same relationship to reality as it does for a person. What is being demonstrated in the article is that given billions of tokens o…
Seems like its nailing it to me. You ask about a scenario and it gives an appropriate answer.
We have evidence that LLMs build models of the things they are learning about. Have a look at this paper:
Do Large Language Models learn world models or just surface statistics?
https://thegradient.pub/othello/
previously discussed https://news.ycombinator.com/item?id=34474043