Live data from Hacker News

Large language models lack deep insights or a theory of mind

arxiv.org

1–10 of 270 posts

Re: Large language models lack deep insights or a theory of mind

#3
The fun question is whether human cognition similarly lacks deep insights or said theory of mind.

I perceive a moving of the goalposts as machine intelligence improves. Once we'd have been happy with smarter than an especially stupid person, now I think we're aiming at smarter than the smartest person.

Re: Large language models lack deep insights or a theory of mind

#8

Do humans have that as well ? I read studies that suggest we make up consciousness a half second after something happened.

Bad liars seem to have difficulty with theory of mind. Sometimes ChatGPT comes across somewhat like this.

Re: Large language models lack deep insights or a theory of mind

#9
I appreciate this paper for relatively clearly stating what "human-like" might entail, which in this case involves "reasoning about the causes behind other people's behavior" which is "critical to navigate the social world" as outlined in this citation:

https://www.sciencedirect.com/science/article/abs/pii/S00100...

I get frustrated often when people argue "well, it isn't really intelligent" and then give examples that are clearly dependent on our brain's chemical state and our bodies' existence in-the-physical-world.

I get the feeling that when/if we are all enslaved by a super-intelligent AI that we do not understand its motives, we will still argue that it is not intelligent because it doesn't get hungry and it can't prove to us that it has Qualia.

This paper argues that gpts are bad at understanding human risk/reward functions, which seems like a much more explicit way to talk about this, and also casts it in a way that could help reframe the debate about how human evolution and our physical beings might be significantly responsible for the structure of our rational minds.

Re: Large language models lack deep insights or a theory of mind

#10

Another paper in a long series that confuses "our tests against currently available LLMs tuned for specific tasks found that they didn't perform well on our task" with "LLMs are architecturally unsuitable for our task".

It's a weird title anyway. I was expecting worse results but GPT-4V is close to or matching Human median performance on most of the tests besides the multimodal "Intuitive Psychology" tests.
Post reply on HN