I was interested in the human results, so I had an llm build a visualization for them: https://codepen.io/lovasoaaa/pen/QwKWGBd You can see that 17% of answers come from India alone and that software developers got below average results, for instance.
“Car Wash” test with 53 models
381–390 of 469 posts
Re: “Car Wash” test with 53 models
#382Earlier quoted context omitted.
"Where is your car?" is not a clarifying question, any more than "Do you hold a valid driver license?" or "Are you a spotted leopard?" Implicit in the question "Should I walk or drive?" is that walking and driving are not strictly impossible choices.
If walking is an option, then your car is already at the car wash. If your car was not at the car wash, then this wouldn't be a question
The only good answers to the car wash questions are either a) "well, duh, drive, since you're gonna need your car there to wash it" (or just "drive", recognizing this as a logic/gotcha puzzle, with no explanation required), or b) "is there something you are not telling me here that makes walking, leaving your car at home, a viable option when the goal is to have your car at the car wash to wash it?".
Re: “Car Wash” test with 53 models
#383The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…
It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.
You can't really infer that from survey data, and particularly from this question. A few criticisms that I came up with off the top of my head:
- What if the number were actually 60% but half guessed right and half guessed wrong?
- Assuming the 30% is a failure of reasoning, it's possible that those 30% were lacking reason at that moment and it's not a general trend. How many times have you just blanked on a question that's really easy to answer?
- A larger percentage than you expected maybe never went to a car wash or don't know what one is?
- Language barrier that leaked through vetting? (Would be a small %, granted)
- Other obvious things like a fraction will have lied just because it's funny, were suspicious, weren't paying attention and just clicked a button without reading the question.
I do agree that the question isn't framed particularly badly, however. I'm just focusing on cognitive impairment, which I don't think is necessarily true all of the time.
Re: “Car Wash” test with 53 models
#384Re: “Car Wash” test with 53 models
#385Re: “Car Wash” test with 53 models
#386Re: “Car Wash” test with 53 models
#387This doesn’t look like a reasoning ceiling. It looks like a decision reliability problem. The unstable tier is the key result. Models that get it right 70–80% of the time are not “almost correct.” They are nondeterministic decision functions. In production that’s worse than being consistently wrong. A single sampled output is just a proposal. If you treat it as a final decision, you inherit its variance. If you treat…
> This doesn’t look like a reasoning ceiling. It looks like a decision reliability problem. This doesn’t look like a human comment. It looks like a LLM response.
Re: “Car Wash” test with 53 models
#388Earlier quoted context omitted.
By the same reasoning, why on earth would a person sincerely ask you that question unless the car that they want to wash is either already at the car wash, or that someone is bringing it to them there for some reason? If it's as unambiguous as you say, then the natural human response to that question isn't "you should drive there". It's "why are you fucking with me?" Or maybe "have you recently suffered a head injury…
How could the car already be at the car wash if you have the option to drive it there?
Re: “Car Wash” test with 53 models
#389Earlier quoted context omitted.
> This doesn’t look like a reasoning ceiling. It looks like a decision reliability problem. This doesn’t look like a human comment. It looks like a LLM response.
Fair I cleaned up the wording with ChatGPT with my review prompt. The substance matters more than the style. If a model flips 3/10 times on a trivial constraint, that’s a reliability issue, not a reasoning ceiling.
I have reviewed your previous comments, and you have consistently written: that's instead of that’s. So what I read is still some LLM output, even though I think there is some kind of human behind the LLM.
Re: “Car Wash” test with 53 models
#390Earlier quoted context omitted.
I don't think this is quite right. It's not that the question is inherently underspecified, it's that the context of being asked a question is itself information that we use to help answer the question. If someone asks "should I walk or drive" to do X, we assume that this is a question that a real human being would have about an actual situation, so even if all available information provided indicates that driving is…
We could probably test this. I wonder if the results shift if the question is prefaced with something like "Here is a trick question: ...".