Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

11–20 of 469 posts

Re: “Car Wash” test with 53 models

#11
Since the conclusion is that context is important, I expected you’d redo the experiment with context. Just add the sentence “The car I want to wash is here with me.” Or possibly change it to “should I walk or drive the dirty car”.

It’s interesting that all the humans critiquing this assume the car isn’t at the car to be washed already, but the problem doesn’t say that.

Re: “Car Wash” test with 53 models

#12
post #4

IMO it's not just intelligence. I think it's related to syncophancy. LLM are trained to not question the basic assumptions being made. They are horrible at telling you that you are solving the wrong problem, and I think this is a consequence of their design. They are meant to get "upvotes" from the person asking the question, so they don't want to imply you are making a fundamental mistake, even if it leads you into…

I think there's also an "alignment blinkers" effect. There is an ethical framework bolted on.

EDIT: Though it could simply reflect training data. Maybe Redditors don't drive.

Re: “Car Wash” test with 53 models

#13

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

It's almost certainly just RAG powered by their crawler.

Re: “Car Wash” test with 53 models

#14
> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive.

I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or “I’m not sure I understand the question, can you rephrase?”, or “Is the vehicle you would drive the same as the car that you want to wash?”, or “Where is your car currently located?”, and so on.

Re: “Car Wash” test with 53 models

#15

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

Presumably it did an actual search and summarized the results and neither answered "off the cuff" by following gradients to reproduce the text it was trained on nor by following gradients to reproduce the "logic" of reasoning. [1]

[1] e.g. trained on traces of a reasoning process

Re: “Car Wash” test with 53 models

#17

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

It seems like the search ai results are generally misunderstood, I also misunderstood them for the first weeks/months.

They are not just an LLM answer, they are an (often cached) LLM summary of web results.

This is why they were often skewed by nonsensical Reddit responses [0].

Depending on the type of input it can lean more toward web summary or LLM answer.

So I imagine that it can just grab the description of the „car wash” test from web results and then get it right because of that.

[0] https://www.bbc.com/news/articles/cd11gzejgz4o

Re: “Car Wash” test with 53 models

#18
The question does not specify what kind of car it is. Technically speaking, a toy car (Hot wheels or a scaled model) could be walked to a car wash.

Now why anyone would wash a toy car at a car wash is beyond comprehension, but the LLM is not there to judge the user's motives.

Re: “Car Wash” test with 53 models

#19
post #6

The human baseline seems flawed. 1. There is no initial screening that would filter out garbage responses. For example, users who just pick the first answer. 2. They don't ask for reasoning/rationale.

I agree. I wonder what the human baseline is for ”what is 1 + 1” on Rapidata.

Re: “Car Wash” test with 53 models

#20
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

That's a fair point, but if you would see it as a riddle, which I don't really think it is, and you had to answer either or, I'd still assume it's most logical to chose drive isn't it?
Post reply on HN