Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

161–170 of 469 posts

Re: “Car Wash” test with 53 models

#162

To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times. My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them. This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification im…

Great observation. Seems like we're back to prompt abracadabra. My little experiment gave me: No added hint 0/3 hint added at the end 1.5/3 hint added at the beginning 3/3 .5 because it stated "Walk" and then convinced it self that "Drive" is the better answer.

If you change the order of the sentences, Sonnet gets it right 3/3: The car wash is 50 meters away. I want to wash my car. Should I walk or drive?

That trick didn't help Mistral Le Chat.

Re: “Car Wash” test with 53 models

#163

Opus 4.6 was getting this wrong only last week.

Oh wow, Sonnet still isn't handling it well:

Opus 4.6: Drive (https://claude.ai/share/d57fef01-df32-41f2-b1dc-07de7916bdc7)

Opus 4.5: Drive (https://claude.ai/chat/a590cac1-100a-490b-b0a2-df6676e1ae99)

Opus 3.0: Walk (https://claude.ai/chat/372c144c-d6eb-43f5-b7ea-fd4c51c681db)

Sonnet 4.6: Walk (https://claude.ai/share/1f2a80f3-4741-40a5-8a05-7349ea1a17e5)

Sonnet 4.5: Walk (https://claude.ai/share/905afeb6-ffc9-4b4b-a9ee-4481e5cfd527)

Favorite answer, using my default custom instructions: "Drive. Walking there means... leaving your car at home? Walk it there on a leash? Walk if you want the exercise, but you're bringing the car either way."

Re: “Car Wash” test with 53 models

#164

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

I don't think 30% of people can't reason. I think 30% of people will fail fairly simple trick questions on any given attempt. That's not at all the same thing.

Some people love riddles and will really concentrate on them and chew them over. Some people are quickly burning through questions and just won't bother thinking it through. "Gotta go to a place, but it's 50 feet away? Walk. Next question, please." Those same people, if they encountered this problem in real life, or if you told them the correct answer was worth a million bucks, would almost certainly get the answer right.

Re: “Car Wash” test with 53 models

#165

Go ask 53 Americans. I’m willing to bet less than 11 get it right.

Don't bet too much, from the linked article ...

They ran the exact same question with the same forced choice between "drive" and > "walk," no additional context, past 10,000 real people through their human feedback platform.

71.5% said drive.

Re: “Car Wash” test with 53 models

#166

To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.

I would love to see LLMs start to ask clarifying questions. That feels like it would be a step up similar to reasoning

Claude Code has an entire tool for the LLM to asking clarifying questions - it'll give you three pre-written responses or you can respond with your own text.

Re: “Car Wash” test with 53 models

#167
post #140

What I find wild is the presumption that with a prompt as simple as “I want to wash my car. My car is 50m away. Should I walk or drive?”, everyone here seems to assume “washing your car” means “taking your car to the car wash”, while what I pictured was “my car is in the driveway, 50m away from me, next to a water hose”, in which case I 100% need to drive.

Critically, that's not the question that was asked. It's not "My car is 50m away", it's "The Car Wash Is 50 Meters Away"

Which hopefully explains why everyone is assuming that "washing your car" does in fact mean "taking your car to the car wash"

Re: “Car Wash” test with 53 models

#168

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

You left out the first half of the prompt: “I want to wash my car”.

Yeah I see this argument being made that it’s ambiguous for humans. Uh, no? Why on earth would I walk to the car wash when I want to wash my car?

Re: “Car Wash” test with 53 models

#170

The test is rigged because they used non thinking models.

Testing some subset X does not mean the test is rigged unless they failed to disclose that.

But also:

GPT 5.2 Thinking, Standard Effort: Walk - https://chatgpt.com/share/699d38cb-e560-8012-8986-d27428de8a...

I'm assuming "GPT 5.2 Thinking" is, in fact, a thinking model?

Post reply on HN