Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

301–310 of 469 posts

Re: “Car Wash” test with 53 models

#302

Funny how we now see AI go through developmental phases similar to what we see in young child development. In a weird convoluted way. Strawberry spelling and car wash aren't particularly intuitive as cognitive developmental stages. E.g. well known mirror-test [1], passed by kids from age 1.5-2 Or object permanence [2], children knowing by age 2 that things that are not in sight do not disappear from existence. [1] ht…

Also strawberry spelling isn't any real test for current LLMs as they have no concept of letters, they work on tokens which may be several characters including punctuation and numerals. To have any hope of getting that question right tokens would have to have the granularity of individual letters, massively ballooning model size and training time, or the LLM needs to be able to call out to an external tool that will return the result (and needs sufficient examples in the training data to prime that trigger to fire).

Re: “Car Wash” test with 53 models

#303

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

What if 30% lack the ability to fill out forms and surveys?

Re: “Car Wash” test with 53 models

#305

Earlier quoted context omitted.

2+2 might well not equal 4, since you haven’t specified the base of the numbers or the modulus of the addition. And what if it’s a full service car wash and you’ve parked nearby because it’s full so you walk over and give them the keys? Assumptions make asses of us all…

So you're saying it would be useful for an "AI assistant" to ask you for the base each time you give it a math problem? Do you also want it to ask you if you're using the conventional definitions of "2" and "+"? For the car wash, would you like it to ask if you're on Earth or on Mars? Do you have air in your tires? Is the car actually a toy car? Some assumptions are always necessary and reasonable, that's why I'm say…

Seems like you’re the one not applying common sense now!

Re: “Car Wash” test with 53 models

#306

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

I don't think this is quite right. It's not that the question is inherently underspecified, it's that the context of being asked a question is itself information that we use to help answer the question. If someone asks "should I walk or drive" to do X, we assume that this is a question that a real human being would have about an actual situation, so even if all available information provided indicates that driving is…

You are only touching on a far bigger and deeper issue around this seemingly “simple prompt”. There is an inherent malicious nature also baked into this prompt that is both telling and very human; a spiteful nature, which usually says more about the humans than anything else.

Your perspective on the meta-question about why such a question would need to be asked in the first place is just the first layer, and most people seem to not even get to that point.

PS: I for one would just like to quickly note for posterity that I do not participate in or am supportive of malicious deception, manipulation, and abuse of AI.

Re: “Car Wash” test with 53 models

#307
> The question has been making the rounds online as a simple logic test, the kind any human gets instantly, but most AI models don't.

...

> They ran the exact same question with the same forced choice between "drive" and "walk," no additional context, past 10,000 real people through their human feedback platform.

> 71.5% said drive.

Well that's a bit embarrassing.

That implies that some models are just better than humans.

I don't think the technology needs to live up to some expectation of perfection, just beat out the human average to have benefit (often, sadly, not to workers themselves).

Re: “Car Wash” test with 53 models

#308
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

This reminds me of a Uni exam that was soooo broken that answering “correctly” entailed guessing how exactly the professor designing the questions misunderstood the topic of his own lectures.

An interesting parallel to that is the "What's the next number in this sequence?" sort of questions.

If four numbers are provided, one can calculate the coefficients of a a quartic polynomial, for x values of 0, 1, 2 and 3, and then solve for x=4. Which does indeed provide a defensible "next number". And by similar reasoning, there are an infinite number of answers to this question.

Even worse. You could in fact provide any number as an answer, because there is always a quintic polynomial that fits the four initial numbers AND your arbitrary fifth number.

So these questions are actually not about what the next number is, but trying to imagine what the person who set the question thought was a "cool" answer, for some curious definition of "cool", for some person who isn't smart enough to realize that the premise on which the question is based is flawed.

Re: “Car Wash” test with 53 models

#309
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

I agree. If the LLM were truly an intelligence, it would be able to ask about this nonsense question. It would be able to ask "Why is walking even an option? Can you please explain how you imagine that would work? Do you mean hand-washing the car at home, instead?" (etc, etc) Real people can ask for clarification when things are ambiguous or confusing. Once something is clarified, they can work that into their unders…

And the corollary: if LLMs were truly intelligent, they would also be able to respond to such questions sarcastically.

Re: “Car Wash” test with 53 models

#310
post #32

Earlier quoted context omitted.

I don’t agree that the question as written would qualify as a riddle. If anything, the riddle is what the intention of the asker is. One can always ask stupid questions with an artificially limited set of answering options; that doesn’t mean it makes sense.

I don't think it qualifies as a stupid question either, it does make sense

It is TOTALLY a stupid question, because OBVIOUSLY you should drive. It is based on the false premise that there is actually a choice. If somebody were to sincerely ask me this question, actually believing that walking was an option, I'm not sure I could resist the temptation to say "walk", just to see what happens next.

Only slightly evil, because the worst-case consequences are an unnecessary 100m walk. I think I could get that past an ethics committee, if I wanted to run an experiment to see what percentage of human responders would ACTUALLY walk to the car wash.

Post reply on HN