Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

41–50 of 469 posts

Re: “Car Wash” test with 53 models

#41

> The funniest part: Perplexity's Sonar and Sonar Pro got the right answer for completely wrong reasons. They cited EPA studies and argued that walking burns calories which requires food production energy, making walking more polluting than driving 50 meters. Right answer, insane reasoning. I mean, Sam Altman was making the same calorie-based arguments this weekend https://www.cnbc.com/2026/02/23/openai-altman-defend…

This was a weird one for sure.

Re: “Car Wash” test with 53 models

#44
post #21

Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…

I tested this with Opus the day 4.6 came out and it failed then, still fails now. There were a lot of jokes I've seen related to some people getting a 'dumber' model, and while there's probably some grain of truth to that I pay for their highest subscription tier so at the very least I can tell you it's not a pay gate issue.

Re: “Car Wash” test with 53 models

#46
post #21

Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…

You mean Sonnet 4.6? I ran 9 claude models including Haiku, swipe through the gallery in the link to see their responses.

I don't see Sonnet 4.6 in the screenshots. I see the other Claude models though.

Edit: Found Haiku. Alas!

Re: “Car Wash” test with 53 models

#47
To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

Re: “Car Wash” test with 53 models

#48
post #46

Earlier quoted context omitted.

You mean Sonnet 4.6? I ran 9 claude models including Haiku, swipe through the gallery in the link to see their responses.

I don't see Sonnet 4.6 in the screenshots. I see the other Claude models though. Edit: Found Haiku. Alas!

Yea good catch Sonnet 4.6 is not part of the test.

Re: “Car Wash” test with 53 models

#49

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

A few years ago if you asked an LLM what the date was, it would tell you the date it was trained, weeks-to-months earlier. Now it gives the correct date. What you've proven is that LLMs leverage web search, which I think we've known about for a while.

Gemini now "knows the time", I was using it in December and it was still lost about dates/intervals...

Re: “Car Wash” test with 53 models

#50

When this first came up on HN, I had commented that Opus 4.6 told me to drive there when I asked it the first time, but when I switched to "Incognito Mode," it told me to walk there. I just repeated that test and it told me to drive both times, with an identical answer: "Drive. You need the car at the car wash."

I mean the n is only 10, so it could still be different for you
Post reply on HN