Earlier quoted context omitted.
Here is what the study says: "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions." So, if people's experience is with modern LLMs, they are being rational…
> So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct. They are not. But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.
People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.