Earlier quoted context omitted.
Like GP said, the LLM has no chance at knowing what a cat is, regardless of how much data it ingests, because a cat is not made of data. It's not like you're getting closer and closer to knowing what a "Mæw" is. You were at the same remote distance all the time. This is called the "grounding problem" in AI. As for how you would test it, I think one-shot learning would get one closer to proving understanding.
The grounding problem is an intelligence problem, not an artificial intelligence problem. How would you envision a test based on one-shot learning working?
As for one-shot learning, what I was driving at, is that a truly intelligent system should not need to consume millions of documents in order to predict that, say, driving at night puts larger demands on one's vision than driving during the day. Or any other common sense fact. These systems require ingesting the whole frickin' internet in order to maybe kinda sometimes correctly answer some simple questions. Even for questions restricted to the narrow range where the system is indeed grounded: the world of symbols and grammar.