Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

51–60 of 469 posts

Re: “Car Wash” test with 53 models

#51
post #31

What I find odd about all the discourse on this question is that no one points out that you have to get out of the car to pay a desk agent at least in most cases. Therefore there's a fundamental question of whether it's worth driving 50m parking, paying, and then getting back in the car to go to the wash itself versus instead of walking a little bit further to pay the agent and then moving your car to the car wash.

You pay at the car wash where I live.

Re: “Car Wash” test with 53 models

#52

Except for a few models, the selected ones were non-reasoning models. Naturally, without reasoning enabled, the reasoning performance will be poor. This is not a surprising result. I asked GPT-5.2 10x times with thinking enabled and it got it right every time.

Thinking or extended thinking?

Re: “Car Wash” test with 53 models

#53

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

Thing is, it's not a riddle or a subliminal message. Everything needed to answer the question is contained therein.

Re: “Car Wash” test with 53 models

#54
post #11

Since the conclusion is that context is important, I expected you’d redo the experiment with context. Just add the sentence “The car I want to wash is here with me.” Or possibly change it to “should I walk or drive the dirty car”. It’s interesting that all the humans critiquing this assume the car isn’t at the car to be washed already, but the problem doesn’t say that.

Agreed, even for humans, context-free logic is a challenge.

Re: “Car Wash” test with 53 models

#55

Earlier quoted context omitted.

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

Presumably it did an actual search and summarized the results and neither answered "off the cuff" by following gradients to reproduce the text it was trained on nor by following gradients to reproduce the "logic" of reasoning. [1] [1] e.g. trained on traces of a reasoning process

[deleted]

Re: “Car Wash” test with 53 models

#56
That human baseline is wild. Either the rapid data test is methodologically flawed or the entire premise of the question is invalid and people are much stupider than even I, a famed misanthrope, think.

Re: “Car Wash” test with 53 models

#57

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

Gemini was a good laugh as well:

    Silas: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
    Gemini:
    ….
    That is a classic “efficiency vs. logic” dilemma.
    Strictly speaking, you should drive. Here is the breakdown of why driving wins this specific round, despite the short distance:
    ...
    * The “Post-Wash” Logic: If you walk there, you’ll eventually have to walk back, get the car, and drive it there anyway. You’re essentially suggesting a pre-wash stroll.
    When should you walk?
    …
    3. You’ve decided the car is too dirty to be seen in public and you’re going to buy a tarp to cover your shame.

Re: “Car Wash” test with 53 models

#58

Earlier quoted context omitted.

A few years ago if you asked an LLM what the date was, it would tell you the date it was trained, weeks-to-months earlier. Now it gives the correct date. What you've proven is that LLMs leverage web search, which I think we've known about for a while.

Gemini now "knows the time", I was using it in December and it was still lost about dates/intervals...

Yeah, the chat log they saved had the correct date. What's your point?

Re: “Car Wash” test with 53 models

#59

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

If you were forced to answer either or, which one would you pick? I think that's where the interesting dynamic comes from. Most humans would pick drive, also seen in the human control, even if it is lower that I thought it'd be

Re: “Car Wash” test with 53 models

#60
post #53

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

Thing is, it's not a riddle or a subliminal message. Everything needed to answer the question is contained therein.

If you want to argue that, then you could also argue that everything needed to challenge the questions’ motives and its validity is also contained therein.

This reminds me of people who answer with “Yes” when presented with options where both can be true but the expected outcome is to pick one. For example, the infamous: “Will you be paying with cash or credit sir?” then the humorous “Yes.”

Post reply on HN