“Car Wash” test with 53 models
161–170 of 469 posts
Re: “Car Wash” test with 53 models
#162To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times. My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them. This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification im…
Great observation. Seems like we're back to prompt abracadabra. My little experiment gave me: No added hint 0/3 hint added at the end 1.5/3 hint added at the beginning 3/3 .5 because it stated "Walk" and then convinced it self that "Drive" is the better answer.
That trick didn't help Mistral Le Chat.
Re: “Car Wash” test with 53 models
#163Opus 4.6 was getting this wrong only last week.
Opus 4.6: Drive (https://claude.ai/share/d57fef01-df32-41f2-b1dc-07de7916bdc7)
Opus 4.5: Drive (https://claude.ai/chat/a590cac1-100a-490b-b0a2-df6676e1ae99)
Opus 3.0: Walk (https://claude.ai/chat/372c144c-d6eb-43f5-b7ea-fd4c51c681db)
Sonnet 4.6: Walk (https://claude.ai/share/1f2a80f3-4741-40a5-8a05-7349ea1a17e5)
Sonnet 4.5: Walk (https://claude.ai/share/905afeb6-ffc9-4b4b-a9ee-4481e5cfd527)
Favorite answer, using my default custom instructions: "Drive. Walking there means... leaving your car at home? Walk it there on a leash? Walk if you want the exercise, but you're bringing the car either way."
Re: “Car Wash” test with 53 models
#164The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…
It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.
Some people love riddles and will really concentrate on them and chew them over. Some people are quickly burning through questions and just won't bother thinking it through. "Gotta go to a place, but it's 50 feet away? Walk. Next question, please." Those same people, if they encountered this problem in real life, or if you told them the correct answer was worth a million bucks, would almost certainly get the answer right.
Re: “Car Wash” test with 53 models
#165Go ask 53 Americans. I’m willing to bet less than 11 get it right.
They ran the exact same question with the same forced choice between "drive" and > "walk," no additional context, past 10,000 real people through their human feedback platform.
71.5% said drive.
Re: “Car Wash” test with 53 models
#166To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.
I would love to see LLMs start to ask clarifying questions. That feels like it would be a step up similar to reasoning
Re: “Car Wash” test with 53 models
#167What I find wild is the presumption that with a prompt as simple as “I want to wash my car. My car is 50m away. Should I walk or drive?”, everyone here seems to assume “washing your car” means “taking your car to the car wash”, while what I pictured was “my car is in the driveway, 50m away from me, next to a water hose”, in which case I 100% need to drive.
Which hopefully explains why everyone is assuming that "washing your car" does in fact mean "taking your car to the car wash"
Re: “Car Wash” test with 53 models
#168The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…
You left out the first half of the prompt: “I want to wash my car”.
Re: “Car Wash” test with 53 models
#169Re: “Car Wash” test with 53 models
#170The test is rigged because they used non thinking models.
But also:
GPT 5.2 Thinking, Standard Effort: Walk - https://chatgpt.com/share/699d38cb-e560-8012-8986-d27428de8a...
I'm assuming "GPT 5.2 Thinking" is, in fact, a thinking model?