LLMs, Theory of Mind, and Cheryl's Birthday
21–30 of 150 posts
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#22The majority of humans in flesh can't solve the problem - so we need alternate measures for judging theory of mind capabilities in LLMs
What about the difference that the human knows what they don't know? In contrast, the LLM knows nothing, but confidently half regurgitates correlational text that it is seen before.
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#23I think the test is better than many other commenters are giving credit. It reminds me of responses to the river crossing problems. The reason people do tests like this is because we know the answer a priori or can determine the answer. Reasoning tests are about generalization, and this means you have to be able to generalize based on the logic. So the author knows that the question is spoiled, because they know that…
>I think the test is better than many other commenters are giving credit. The test is fine. The conclusion drawn from it, not so much. If humans fail your test for x and you're certain humans have x then you're not really testing for x. x may be important to your test for sure but you're testing for something else too. Or maybe humans don't have x after all. Either conclusion is logically consistent at least. It's th…
> The conclusion drawn from it, not so much. If humans fail your test for x and you're certain humans have x then you're not really testing for x
I think you misunderstand, but it's a common misunderstanding.Humans have the *ability* to reason. This is not equivalent to saying that humans reason at all times (this was also started in my previous comment)
So it's none of: "humans have x", "humans don't have x", nor "humans have x but f doesn't have x because humans perform y on x and f performs z on x".
It's correct to point out that not all humans can solve this puzzle. But that's an irrelevant fact because the premise is not that human always reason. If you'd like to make the counter argument that LLMs are like humans in that they have the ability to reason but don't always, then you got to provide strong evidence (just like you need to provide strong evidence that LLMs can reason). But this (both) is quite hard to prove because humans aren't entropy minimizers trained on petabytes of text. It's easier to test humans because we generally have a much better idea of what they've been trained on and we can also sample from different humans that have been trained on different types of data.
And here's a real kicker, when you've found a human that can solve a problem (meaning not just state the answer but show their work) nearly all of them can adapt easily to novel augmentations.
So I don't know why you're talking about trickery. The models are explicitly trained to solve problems like these. There's no slight of hand. There's no magic tokens, no silly or stage wording that would be easily misinterpreted. There's a big difference between a model getting an answer wrong and a promoter tricking the model.
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#24I think the test is better than many other commenters are giving credit. It reminds me of responses to the river crossing problems. The reason people do tests like this is because we know the answer a priori or can determine the answer. Reasoning tests are about generalization, and this means you have to be able to generalize based on the logic. So the author knows that the question is spoiled, because they know that…
Humans are different from LLMs. LLMs are giving it 100%, every time. "Proof by self-incompetence" is an exact example of this, where the goal is to explain a prior belief. But fitting data is easy, explaining data is hard (von Neumann's Elephant).
Ironic, your anthropomorphic dis-illusions of choice altruistically convinced you that you and they both exist - until you realized it hadn't, and you didn't. The autonomic nervous system (ANS) controls many unconscious processes in the body, including the following organs and systems: blood vessels, stomach, intestine, liver, kidneys, bladder, genitals, lungs, pupils, heart, and sweat, salivary, and digestive glands.
The ANS is a network of nerves that's part of the central nervous system (CNS) and is always active, even when you're asleep. It's essential for life - the war against entropy is ceaseless.Re: LLMs, Theory of Mind, and Cheryl's Birthday
#25This question was posed to o1, it is able to reason through it - but now I wonder if that is because the model is already aware of the puzzle. https://x.com/d_feldman/status/1834313124058726894
Cheryl's birthday puzzle is a logic problem where Albert and Bernard are trying to figure out Cheryl's birthday based on certain clues.
Cheryl provides them with ten possible dates: May 15, May 16, May 19, June 17, June 18, July 14, July 16, August 14, August 15, and August 17.
Here’s the reasoning:
1. Albert knows the month and Bernard knows the day.
2. Albert says he knows Cheryl’s birthday, meaning May and June can be eliminated because they contain unique days (May 19 and June 18). If Albert had been told May or June, he wouldn’t know for sure.
3. Bernard, knowing this, says he now knows Cheryl’s birthday. This eliminates the remaining dates with unique days (July 14 and August 14).
4. Albert then confirms that he also knows the birthday, meaning Cheryl’s birthday must be in July or August, but on a date with no unique days left: July 16, August 15, or August 17.
Thus, Cheryl's birthday is *July 16*.
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#26Of course he is hardly the only offender: arrogant disregard for psychology is astonishingly common among LLM researchers. Maybe they should turn off ChatGPT and read a book.
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#27AI researchers need to learn what terms like "theory of mind" actually mean before they write dumb crap like this. Theory of mind is about attributing mental states to others, not information. What Norvig has done here is present a logic puzzle, one that works equally well when the agents are Prolog programs instead of clever children. There's no "mind" in this puzzle at all. Norvig is being childishly ignorant to ca…
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#28The majority of humans in flesh can't solve the problem - so we need alternate measures for judging theory of mind capabilities in LLMs
Re: LLMs, Theory of Mind, and Cheryl's Birthday
#29Deducing things from the inability of an LLM to answer a specific question seemed doomed by the "it will be able to on the next itteration" principle. It seems like the only way you could systematic chart the weaknesses of an LLM is by having a class of problems that get harder for LLMs at a steep rate, so a small increase in problem complexity requires a significant increase in LLM power.
> Deducing things from the inability of an LLM to answer a specific question seemed doomed by the "it will be able to on the next itteration" principle. That's orthogonal. If we are pointing in the right direction(s) then yes, next iteration could resolve all problems. If we are not pointing in the right direction(s) then no, next iteration will not resolve these problems. Given LLMs rapid improvement in regurgitatin…
Thus observers of the LLM space like us need to keep finding novel “Bellweather problems” that we think will evaluate a model’s ability to reason, knowing that once we start talking about it openly the problem will no longer be a useful Bellweather.
By their nature as “weird-shaped” problems, these aren’t the kind of thing we’re guaranteed to have an infinite supply of. As the generations move on it will become more and more difficult to discern “actual improvements in reasoning” from “the model essentially has the solution to your particular riddle hard-coded”.