LLMs don't learn to simulate or mimic, that's just a byproduct. They learn to predict the training corpus. There is absolutely nothing about the act of prediction that necessitates an upper bound of intelligence on the corpus itself. https://www.pnas.org/doi/full/10.1073/pnas.2016239118 They found representations on fundamental properties of proteins such as secondary structure, contacts, and biological activity in a…
To your point, prediction may not be bound by what the data explicitly shows, but it is necessarily bound by the data and its implicit and explicit structure. But to argue that the LLM can make any discovery a person can, or interact fluidly with the world as it exists, would require one to believe that every property of the world is either explicitly captured in human-written text or in the structure of that text, which I think would be pretty difficult to argue.