This is a terrible eval. Do not update your beliefs on whether LLMs have Theory of Mind based on this paper. The eval is a weird, noisy visual task (picture of astronaut with “care packages”). Their results are hopelessly narrow. A better eval is to use actual scientifically tested psychology test on text (the native and strongest domain for LLMs), for example the sort of scenarios used to gauge when children develop…
Nothing I’ve seen shows evidence of any sort of abstract concepts in there.