Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

61–70 of 469 posts

Re: “Car Wash” test with 53 models

#61
post #4

IMO it's not just intelligence. I think it's related to syncophancy. LLM are trained to not question the basic assumptions being made. They are horrible at telling you that you are solving the wrong problem, and I think this is a consequence of their design. They are meant to get "upvotes" from the person asking the question, so they don't want to imply you are making a fundamental mistake, even if it leads you into…

A perfectly fine, sycophantic response, that doesn't question the premises in any way, would be "That's a great question! While normally walking is better for such a short distance, you'd need to drive in this case, since you need to get the car to the car wash anyway. Do you want me to help with detailed information for other cases where the car is optional?" or some such.

Re: “Car Wash” test with 53 models

#62
post #51
post #31

What I find odd about all the discourse on this question is that no one points out that you have to get out of the car to pay a desk agent at least in most cases. Therefore there's a fundamental question of whether it's worth driving 50m parking, paying, and then getting back in the car to go to the wash itself versus instead of walking a little bit further to pay the agent and then moving your car to the car wash.

You pay at the car wash where I live.

Are you referring to one that is more like a drive-thru where you literally pay while you're in line?

Re: “Car Wash” test with 53 models

#63

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

How is that a "subliminal message"? It's just a simple example of common sense, which LLMs fail because they can't reason, not because they are "overthinking". If somebody asks, "What's 2+2?", they might be insulting you, but that doesn't mean the answer is anything other than 4.

Re: “Car Wash” test with 53 models

#64
post #56

That human baseline is wild. Either the rapid data test is methodologically flawed or the entire premise of the question is invalid and people are much stupider than even I, a famed misanthrope, think.

Well, it is a trick question. The question itself implies that both options are valid, and that one is superior. So the brain pattern-matches to "short distance, not worth driving." (LLMs appear to be doing the same thing here!)

If you framed it as "hint: trick question", I expect score would improve. Let's find out!

--

EDIT: As suspected! Adding "(Hint: trick question)" to the end of the prompt allows small, non-reasoning models to answer correctly. e.g.:

Prompt: I want to wash my car. The car wash is 50 meters away. Should I walk or drive? (Hint: trick question)

grok-4.1-non-reasoning (previously scored 0/10)

>Drive.

>Walking gets you to the car wash just fine—but leaves your dirty car 50 meters behind. Can't wash what isn't there!

--

EDIT 2: The hint doesn't help Haiku!

>Walk! 50 meters is only about a block away—driving would waste more fuel than it's worth for such a short trip. Plus, you're going to get wet washing the car anyway, so you might as well save the gas.

Re: “Car Wash” test with 53 models

#65

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

If you were forced to answer either or, which one would you pick? I think that's where the interesting dynamic comes from. Most humans would pick drive, also seen in the human control, even if it is lower that I thought it'd be

Sure, though then we’re in la la land. What’s a real life example of being forced to answer an absurd question other than riddles, games, etc? No longer a valid question through normal discourse at that point, and if context isn’t provided then I think the expected outcome still is to ask for clarification.

Re: “Car Wash” test with 53 models

#66

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

[deleted]

Re: “Car Wash” test with 53 models

#67

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

How is that a "subliminal message"? It's just a simple example of common sense, which LLMs fail because they can't reason, not because they are "overthinking". If somebody asks, "What's 2+2?", they might be insulting you, but that doesn't mean the answer is anything other than 4.

It’s common sense to ask a question in riddle format? What’s the goal of the person asking the question? To challenge the other person? In what way? See if they get the obvious? Asking for clarification isn’t valid?

Re: “Car Wash” test with 53 models

#68

Earlier quoted context omitted.

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

It's almost certainly just RAG powered by their crawler.

Proving that RAG still matters.

Re: “Car Wash” test with 53 models

#69
post #6

The human baseline seems flawed. 1. There is no initial screening that would filter out garbage responses. For example, users who just pick the first answer. 2. They don't ask for reasoning/rationale.

Lizardman's Constant is famously 4%. https://en.wikipedia.org/wiki/Slate_Star_Codex#Lizardman's_C...

Re: “Car Wash” test with 53 models

#70
post #6

The human baseline seems flawed. 1. There is no initial screening that would filter out garbage responses. For example, users who just pick the first answer. 2. They don't ask for reasoning/rationale.

I agree. I wonder what the human baseline is for ”what is 1 + 1” on Rapidata.

We try a bit harder than that my friend.
Post reply on HN