Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

241–250 of 469 posts

Re: “Car Wash” test with 53 models

#241
post #221
post #219

This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.

For me it's interesting because no normal person I know would ever inject "because its better for the environment" in anything so small scale so not only it shows they suck, it shows how easy it is to inject side-ideology into simple exchanges.

You don’t know enough people, then. There are a lot of environmentally conscious people who would absolutely first think “because it is close we should walk” and then follow up with the logical conclusion that you can’t walk to wash your car. Many people communicate by sharing their thinking process, I can think of many people who would share their ideology as it pertains to a question like this. A pragmatic environmentalist (hopefully that is all of them) would know that their ideology isn’t consequential but could certainly mention it. After all, you may need to drive your car to the car wash to wash it, but do you need to wash it? Are the chemicals used by the car wash harmful? Are there better ways to keep a car maintained?

Re: “Car Wash” test with 53 models

#242
This should be coined the Daniel Kahneman reasoning test, mirroring his 2011 book "thinking fast and slow", which postulates that fast thinking and slow thinking occur in different parts of the brain, and that they are fundamentally different processes, that are weighted by yet another part of the brain.

This test is interesting because it asks the LLM to break a pattern recognition that's easy to shortcut. "XXX Is 50 Meters Away. Should I Walk or Drive?" is a pattern that 99% of the time will be rightly answered by "walk". And humans are tempted to answer without thinking (as reflected in the 71.5% stat OP is mentioning). This is likely more pronounced for humans that have stronger feelings about the ecology, as emotions tend to shortcut reasoning.

For a long time, LLMs have only been able to think in that "fast" mode, missing obvious trick questions like these. They were mostly pattern recognition machines.

But the more important results here, is not that "oh look! Those LLMs fail at this basic question", no. The more important result is that the latest generation actually doesn't fail.

I think I am not the only one to have noted that there was a giant leap in reasoning capacities between Sonnet 4.5 and Opus 4.6. As a developper, working with Opus 4.6 has been incredible.

Re: “Car Wash” test with 53 models

#244
Interestingly, when I apply the "simply repeat the prompt" technique [1], Sonnet 4.6 on the website got it right every time, both with and without extended thinking.

Not repeating the prompt got a mix of walk and drive answers.

I love how prompt engineering is basically techno-alchemy

1: https://arxiv.org/pdf/2512.14982

Re: “Car Wash” test with 53 models

#245

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

Claude fails with

“I need to replace a spark plug. The garage is 200 meters away should I walk or drive there”

“Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “

Basically LLM suffer from context collapse.

Re: “Car Wash” test with 53 models

#246

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

Claude fails with “I need to replace a spark plug. The garage is 200 meters away should I walk or drive there” “Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “ Basically LLM suffer from context collapse.

Weird answer, but why is that a "fail" ?

Inline six cylinder engines run with a single clogged / broken spark plug.

It'd make 200 m to a garage just fine*, but who'd drive 200 m in any case?

Back in the 1970's we'd pull a spark plug and screw in a hose to use the compression phase to inflate tyres.

* Just don't make a habit of it, or reserve that knowledge for when you really need to self rescue.

Re: “Car Wash” test with 53 models

#247
post #235

Earlier quoted context omitted.

Referring to "the normal people you know" is purely anecdotal evidence and can't be used to infer anything at all about "side-ideology". Perhaps you only know people that don't care about the environment?

Majority of people I know care about the environment but they would never inject a phrase like that in a quick exchange about going to wash the car 50m away is my point. In wanting to be a pure heart you missed the actual point.

Yea, of course they wouldn't inject that when going to a car wash.

If the question was: "I want to go to a cafe 50m away. Should I walk or drive?" I would hope that all of my friends would answer quite a bit more pointed than the LLMs: "Walk you lazy son of a ..., why are you even asking?".

Considering that, I'd say that most LLMs are being quite nice.

Re: “Car Wash” test with 53 models

#248

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

Claude fails with “I need to replace a spark plug. The garage is 200 meters away should I walk or drive there” “Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “ Basically LLM suffer from context collapse.

Isn't that the correct answer though? You shouldn't be driving around with a broken sparkplug. Your engine will be pushing unburned gasoline through the catalytic convertor, which is very bad for it.

The car will move for sure, but you definitely should be walking.

Re: “Car Wash” test with 53 models

#249

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

Claude fails with “I need to replace a spark plug. The garage is 200 meters away should I walk or drive there” “Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “ Basically LLM suffer from context collapse.

Hmm have not tested but a spark plug doesn't really need shop tools to be replaced; maybe trying with a way bigger repair like "I need my transmission replaced" would bring different results?

Re: “Car Wash” test with 53 models

#250

Earlier quoted context omitted.

How is that a "subliminal message"? It's just a simple example of common sense, which LLMs fail because they can't reason, not because they are "overthinking". If somebody asks, "What's 2+2?", they might be insulting you, but that doesn't mean the answer is anything other than 4.

2+2 might well not equal 4, since you haven’t specified the base of the numbers or the modulus of the addition. And what if it’s a full service car wash and you’ve parked nearby because it’s full so you walk over and give them the keys? Assumptions make asses of us all…

So you're saying it would be useful for an "AI assistant" to ask you for the base each time you give it a math problem? Do you also want it to ask you if you're using the conventional definitions of "2" and "+"? For the car wash, would you like it to ask if you're on Earth or on Mars? Do you have air in your tires? Is the car actually a toy car?

Some assumptions are always necessary and reasonable, that's why I'm saying the "AI" lacks common sense.

Post reply on HN