Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

151–160 of 469 posts

Re: “Car Wash” test with 53 models

#151
post #24

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.

I mean their whole existence is about token prediction, so they just want to do their things :)

Re: “Car Wash” test with 53 models

#152

To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times. My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them. This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification im…

Great observation. Seems like we're back to prompt abracadabra.

My little experiment gave me:

No added hint 0/3

hint added at the end 1.5/3

hint added at the beginning 3/3

.5 because it stated "Walk" and then convinced it self that "Drive" is the better answer.

Re: “Car Wash” test with 53 models

#153
Sites wanting to block AI scraping should simply ask questions like these, instead of furthering the complexity-driven monopoly of Big Tech by requiring specifically sanctioned software and hardware. This is how you determine human intelligence, and not mindless compliance.

Re: “Car Wash” test with 53 models

#154
post #126
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

It highlights a general problem with LLMs, that they always jump to answering, whereas humans will often ask clarifying questions first.

I wonder if anyone has any research on this field. I've often seen this myself (too often) where LLMs make assumptions and run off with the wrong thing.

"This is how you do " or "This is why is impossible!". Ffs man, just ask for info! A human wouldn't need to - they'd get the context - but LLMs apparently don't?

Re: “Car Wash” test with 53 models

#157

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

> 30% of people simply lack the ability to reason

While I’m sure it’s more than 0%, seems more likely that somewhere between 0% and 30% don’t feel obligated to give the inquiry anything more than the most cursory glance.

How do incentives align differently with LLMs?

Re: “Car Wash” test with 53 models

#160
Did AI write the post?

First section says "The models that passed the car wash test: ...Gemini 2.0 Flash Lite..."

A section or 2 down it says: "Single-Run Results by Model Family: Gemini 3 models nailed it, all 2.x failed"

In the section below that about 10 runs it says: 10/10 — The Only Reliable AI Models ... Gemini 2.0 Flash Lite ..."

So which it is? Gemini 2.x failed (2nd section) or it succeeded (1st and 3rd) section. Or am I mis-understanding

Post reply on HN