> The funniest part: Perplexity's Sonar and Sonar Pro got the right answer for completely wrong reasons. They cited EPA studies and argued that walking burns calories which requires food production energy, making walking more polluting than driving 50 meters. Right answer, insane reasoning. I mean, Sam Altman was making the same calorie-based arguments this weekend https://www.cnbc.com/2026/02/23/openai-altman-defend…
“Car Wash” test with 53 models
41–50 of 469 posts
Re: “Car Wash” test with 53 models
#42I'm going to test this on my kids.
Re: “Car Wash” test with 53 models
#43Also, the summary of the Gemini model says: "Gemini 3 models nailed it, all 2.x failed", but 2.0 Flash Lite succeeded, 10/10 times?
Re: “Car Wash” test with 53 models
#44Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…
Re: “Car Wash” test with 53 models
#45The test is rigged because they used non thinking models.
Re: “Car Wash” test with 53 models
#46Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…
You mean Sonnet 4.6? I ran 9 claude models including Haiku, swipe through the gallery in the link to see their responses.
Edit: Found Haiku. Alas!
Re: “Car Wash” test with 53 models
#47Re: “Car Wash” test with 53 models
#48Earlier quoted context omitted.
You mean Sonnet 4.6? I ran 9 claude models including Haiku, swipe through the gallery in the link to see their responses.
I don't see Sonnet 4.6 in the screenshots. I see the other Claude models though. Edit: Found Haiku. Alas!
Re: “Car Wash” test with 53 models
#49I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…
A few years ago if you asked an LLM what the date was, it would tell you the date it was trained, weeks-to-months earlier. Now it gives the correct date. What you've proven is that LLMs leverage web search, which I think we've known about for a while.
Re: “Car Wash” test with 53 models
#50When this first came up on HN, I had commented that Opus 4.6 told me to drive there when I asked it the first time, but when I switched to "Incognito Mode," it told me to walk there. I just repeated that test and it told me to drive both times, with an identical answer: "Drive. You need the car at the car wash."