Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

271–280 of 469 posts

Re: “Car Wash” test with 53 models

#271
Funny how we now see AI go through developmental phases similar to what we see in young child development. In a weird convoluted way. Strawberry spelling and car wash aren't particularly intuitive as cognitive developmental stages.

E.g. well known mirror-test [1], passed by kids from age 1.5-2

Or object permanence [2], children knowing by age 2 that things that are not in sight do not disappear from existence.

[1] https://en.wikipedia.org/wiki/Mirror_test [2] https://en.wikipedia.org/wiki/Object_permanence

Re: “Car Wash” test with 53 models

#272
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

It feels more like a question on english linguistic conventions than logic.

If someone asked me the same question and I wanted to give a smartass reply, I'd tell them "You want to wash your car, good to know. Now, about your question, unless you tell me where you wanna go I can't really help you".

Re: “Car Wash” test with 53 models

#273

Earlier quoted context omitted.

We try a bit harder than that my friend.

I actually didn't mean to criticize Rapidata. I just think that a forced-choice question like this begs for low-effort answers. At least the respondents should have had the opportunity to explain their reasoning, like the LLMs did.

All good ^^, its a fair point, we have come up with some fun ways to track peoples reliability over time. But the validation sets contain plenty of forced-choice questions, those that have an empirical true can be used directly to calculate a reliability, those that are subjective need to be re-asked after sometime to ensure consistency. People that don't pass thresholds would not be part of the 10'000 here.

But of course. If every human was told to take 3 minutes to deeply think about it and told that its a trick question, then they most likely will all get it right. But its the same with the LLMs, if you ask them like that they will get it right most of the time. The low effort is kinda the point here.

Re: “Car Wash” test with 53 models

#274

I'm doubting the 29-ish percent of people submitting 'walk' are actually human. Is it not obvious that you need a car to wash? Are they using LLM to answer?

it is surprising, but give this question to some random people on the street without context and you would be surprised

Re: “Car Wash” test with 53 models

#275
The article claims that every Claude model other than Opus 4.6 reliably fails. This is not true, Sonnet 3.5 answers correctly around half of the time, even though it's such an old model it's not even available on the main API anymore.

Re: “Car Wash” test with 53 models

#276

If you speak French to Mistral, it gets it right everytime: Je veux laver ma voiture. La station de lavage est à 50 mètres. J'y vais à pied ou en voiture ?

I've been gone from France too long. I've never heard "station de lavage" before.

Very awkward and formal. Anyone would call it lavage auto, lave-auto or simply lavage if the context is clear.

Re: “Car Wash” test with 53 models

#277

The test is rigged because they used non thinking models.

Testing some subset X does not mean the test is rigged unless they failed to disclose that. But also: GPT 5.2 Thinking, Standard Effort: Walk - https://chatgpt.com/share/699d38cb-e560-8012-8986-d27428de8a... I'm assuming "GPT 5.2 Thinking" is, in fact, a thinking model?

The problem is you haven't used the API, but you have used your ChatGPT subscriptions with personality, memories and possible customization. I can see for instance that your ChatGPT answers with emojis, while my ChatGPT subscription never does.

If you ask GPT 5.2 with high reasoning efforts in the API, you get 10 out of 10: drive.

Re: “Car Wash” test with 53 models

#278

Earlier quoted context omitted.

I don’t think it’s under specified. You are clearly stating “I want to wash my car”, then asking how you should get there. It’s an easy logical step to know that, in this context, you need your car with you to wash it, and so no matter the distance you should drive. You can ask the human race the simplest, most logical question ever, and a percentage of them will get it wrong.

1. When do you want to wash your car? Tomorrow? Next year? In 50 years? 2. Where is the car now? Is it already at the car wash waiting for you to arrive? I can see why an LLM might miss this. I think any good software engineer would ask clarifying questions before giving an answer. The next step for an LLM is to either ask questions before giving a definitive answer for uncertain things or to provide multiple answers…

3. Is the car broken somewhere? Does it have wheels on?

4. Does the car have enough fuel?

Jokes asides, all of those questions are unnecessary. There's no more context to this.

Re: “Car Wash” test with 53 models

#279

Opus 4.6 was getting this wrong only last week.

Oh wow, Sonnet still isn't handling it well: Opus 4.6: Drive ( https://claude.ai/share/d57fef01-df32-41f2-b1dc-07de7916bdc7 ) Opus 4.5: Drive ( https://claude.ai/chat/a590cac1-100a-490b-b0a2-df6676e1ae99 ) Opus 3.0: Walk ( https://claude.ai/chat/372c144c-d6eb-43f5-b7ea-fd4c51c681db ) Sonnet 4.6: Walk ( https://claude.ai/share/1f2a80f3-4741-40a5-8a05-7349ea1a17e5 ) Sonnet 4.5: Walk ( https://claude.ai/share/905afeb6-f…

This is because it is without thinking enabled. Of course the results are disappointing.

Re: “Car Wash” test with 53 models

#280

Funny how we now see AI go through developmental phases similar to what we see in young child development. In a weird convoluted way. Strawberry spelling and car wash aren't particularly intuitive as cognitive developmental stages. E.g. well known mirror-test [1], passed by kids from age 1.5-2 Or object permanence [2], children knowing by age 2 that things that are not in sight do not disappear from existence. [1] ht…

Enable reasoning effort and the results are completely different.
Post reply on HN