Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

391–400 of 469 posts

Re: “Car Wash” test with 53 models

#391

Earlier quoted context omitted.

I fail to see how these things are one and the same. I get the point you are making, I just don't agree with it. 2+2 is a complete expression, the other is grammatically correct but logically flawed. Where is the logical fallacy in 2+2?

Well, I don't think you get my point based on your last question. My point is that there is no logical fallacy in the car wash question, just like there is none in 2+2. How is it any more logically flawed than asking, "I want to shop for groceries. The shop 50 meters away. Should I walk or drive?".

You’re conflating it being a question granting making it logically sound. The prior context in the question is what adds the logical fallacy to it, the question without that is fine but given the information about the car it becomes absurd. Your new example illustrate different things, context cannot be ignored here as it is what makes the entire thing what it is. In the car wash example, the context has a direct relationship with the question that determines the answer, the relationship matters so much that OP claims that for its benchmark purposes only “drive” is the valid answer. That special condition is what makes it a puzzle, a test, and a logically flawed proposition to test your attention despite it being structured as a question grammatically. 2+2 does not bring this relationship in its structure and presentation.

Re: “Car Wash” test with 53 models

#392

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

> how we interpret underspecified questions

The question was not merely 'should I walk or drive to the car wash', it was prefaced with 'I Want to Wash My Car. The Car Wash Is 50 Meters Away.'

This is not underspecified - the only relevant detail was included up front in the very first sentence.

Re: “Car Wash” test with 53 models

#394
I maintain a private evaluation set of what many call "misguided attention" questions.

In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions.

A lot of viral "frontier model fails THIS simple question" examples are essentially carefully constructed token sequences designed to bias the statistical prior toward an intuitively wrong answer. Small wording changes can flip results entirely.

If you systematically expand the prompt space around such questions—adding or removing minor contextual cues you'll typically find symmetrical variants where the same models both succeed and fail. That suggests sensitivity to framing and distributional priors (adding unnecessary info, removing clear info, add ambiguity, ...), not necessarily absence of reasoning capability.

Re: “Car Wash” test with 53 models

#396

I maintain a private evaluation set of what many call "misguided attention" questions. In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions. A lot of viral "fro…

You should publish your evaluation set, that seems pretty interesting!

What’s your favourite one?

Re: “Car Wash” test with 53 models

#397

I maintain a private evaluation set of what many call "misguided attention" questions. In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions. A lot of viral "fro…

Some might argue "sensitivity to framing and distributional priors" is a fancy way to say "absence of reasoning capability".

Re: “Car Wash” test with 53 models

#398

I maintain a private evaluation set of what many call "misguided attention" questions. In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions. A lot of viral "fro…

Some might argue "sensitivity to framing and distributional priors" is a fancy way to say "absence of reasoning capability".

that's what the cultivators of these examples are preying on. but in practice what people care about is "can i get it to do ", not "is it a decider on every possible token sequence that humans perceive to be about ".

Re: “Car Wash” test with 53 models

#399
post #388

Earlier quoted context omitted.

How could the car already be at the car wash if you have the option to drive it there?

You might own multiple cars, you might be borrowing someone elses and so forth.

That still doesn't make sense. I'm going to use another car, or borrow a car to drive to a carwash where my car I want to wash is and then....I guess leave it there? Or leave the car I came in?

This isn't a viable out for explaining why AI can't "reason" through this.

Re: “Car Wash” test with 53 models

#400

I maintain a private evaluation set of what many call "misguided attention" questions. In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions. A lot of viral "fro…

Sounds interesting, would be nice to see the questions if you're open to sharing?
Post reply on HN