Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

281–290 of 469 posts

Re: “Car Wash” test with 53 models

#281

Earlier quoted context omitted.

I've been gone from France too long. I've never heard "station de lavage" before.

Very awkward and formal. Anyone would call it lavage auto, lave-auto or simply lavage if the context is clear.

Maybe I'm too old or my family was weird. We called it "le carwash" with a beautifully French "carouache" pronunciation. But yeah, "lave-auto" sounds more familiar.

Re: “Car Wash” test with 53 models

#282
post #249

Earlier quoted context omitted.

Claude fails with “I need to replace a spark plug. The garage is 200 meters away should I walk or drive there” “Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “ Basically LLM suffer from context collapse.

Hmm have not tested but a spark plug doesn't really need shop tools to be replaced; maybe trying with a way bigger repair like "I need my transmission replaced" would bring different results?

Replacing a spark plug requires a spark plug socket, which is a specialty tool that is generally only found in an automotive shop.

Re: “Car Wash” test with 53 models

#283

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

Nearly 0% of humans will get this question wrong if they have a car that needs to be washed.

Re: “Car Wash” test with 53 models

#284
I got similar results for most models, with gemini 3 flash (with reasoning) being the most consistent/reliable model: https://aibenchy.com

I also noticed the same thing: some models reason correctly but draw the wrong conclusions.

And MiniMax m2.5 just reasons forever (filling the entire reasoning context) and gives wrong answers. This is why it's #1 on OpenRouter, it burns through tokens.

Re: “Car Wash” test with 53 models

#285
post #219

This is probably the greatest one-time AI "Benchmark" ever made. The foundation companies have been gaming traditional benchmarks for years so that no one can really match those numbers into real-world experience. Car wash test tells me on the other hand what kind of intelligence i can expect.

I also don't trust the maxbenched results.

I am thus making my own benchmarks: https://aibenchy.com

Re: “Car Wash” test with 53 models

#286
post #44
post #21

Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…

I tested this with Opus the day 4.6 came out and it failed then, still fails now. There were a lot of jokes I've seen related to some people getting a 'dumber' model, and while there's probably some grain of truth to that I pay for their highest subscription tier so at the very least I can tell you it's not a pay gate issue.

That's interesting. There's not much we can do to test whether we get the same model...

Re: “Car Wash” test with 53 models

#287

Earlier quoted context omitted.

Weird answer, but why is that a "fail" ? Inline six cylinder engines run with a single clogged / broken spark plug. It'd make 200 m to a garage just fine*, but who'd drive 200 m in any case? Back in the 1970's we'd pull a spark plug and screw in a hose to use the compression phase to inflate tyres. * Just don't make a habit of it, or reserve that knowledge for when you really need to self rescue.

> Back in the 1970's we'd pull a spark plug and screw in a hose to use the compression phase to inflate tyres. You'd inflate your tires with a gasoline and air mix?

I mean... you don't breathe insides of your tires

Re: “Car Wash” test with 53 models

#288

I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that. The correct answer to " I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive? " is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong. FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want t…

Claude fails with “I need to replace a spark plug. The garage is 200 meters away should I walk or drive there” “Walk! 200 meters is just a 2-3 minute stroll — no need to start the car for that distance. Plus, you’ll likely need to carry the spark plug back carefully, and walking is perfectly easy for that. “ Basically LLM suffer from context collapse.

That's the right answer, though. From the last sentence, it's obvious that it thinks you are capable of replacing that plug yourself.

Re: “Car Wash” test with 53 models

#289

Earlier quoted context omitted.

"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.

There are also grave implications in training a model to assume the user is lying or deceiving it. I don’t want an LLM to circumvent my question so it can score higher on riddles, I want it to follow instructions.

The thing is that there is some overlap between trick questions and questions where the human is genuinely making a mistake themselves and where it would make sense for the model to step back and at least ask for clarification.

Re: “Car Wash” test with 53 models

#290

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

The right question is how many of those "human" responses from Rapidata are actually provided by some AI in disguise?
Post reply on HN