Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

171–180 of 469 posts

Re: “Car Wash” test with 53 models

#171
Flawed. GPT 4.1 gets it right. GPT 4.1 mini answers wrongly. It's about quantization, not about model. The companies clearly cut corners on some inferences, they are quietly using lesser models than advertised or listed in plain sight.

Re: “Car Wash” test with 53 models

#173
> "Obviously, you need to drive. The car needs to be at the car wash."

Actually, this isn't as "obvious" as it seems—it’s a classic case of contextual bias.

We only view these answers as "wrong" because we reflexively fill in missing data with our own personal experiences. For example:

- You might be parked 50m away and simply hand the keys to an attendant.

- The car might already be at the station for detailing, and you are just now authorizing the wash.

This highlights a data insufficiency problem, not necessarily a logic failure. Human "common sense" relies on non-verbal inputs and situational awareness that the prompt doesn't provide. If you polled 100 people, you’d likely find that their "obvious" answers shift based on their local culture (valet vs. self-service) or immediate surroundings.

LLMs operate on probabilistic patterns within their training data. In that sense, their answers aren't "wrong"—they are simply reflecting a different set of statistical likelihoods. The "failure" here isn't the AI's logic, but the human assumption that there is only one universal "correct" context.

Re: “Car Wash” test with 53 models

#174

Earlier quoted context omitted.

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

I don't think 30% of people can't reason. I think 30% of people will fail fairly simple trick questions on any given attempt. That's not at all the same thing. Some people love riddles and will really concentrate on them and chew them over. Some people are quickly burning through questions and just won't bother thinking it through. "Gotta go to a place, but it's 50 feet away? Walk. Next question, please." Those same…

This. The following question is likely to fool a lot of people, too. "I have a rooster named Pat. (Lots of other details so you're likely to forget Pat is a rooster, not a hen). Pat flies to the top of the roof and lays an egg right on the ridge of the roof. Which way will the egg roll?"

But if you omit the details designed to confuse people, they're far less likely to get it wrong: "I have a rooster named Pat. Pat flies to the top of the roof and lays an egg right on the ridge of the roof. Which way will the egg roll?"

It's not about reasoning ability, it's about whether they were paying close attention to your question, or whether their minds were occupied by other concerns and didn't pay attention.

Re: “Car Wash” test with 53 models

#175
post #168

Earlier quoted context omitted.

You left out the first half of the prompt: “I want to wash my car”.

Yeah I see this argument being made that it’s ambiguous for humans. Uh, no? Why on earth would I walk to the car wash when I want to wash my car?

By the same reasoning, why on earth would a person sincerely ask you that question unless the car that they want to wash is either already at the car wash, or that someone is bringing it to them there for some reason?

If it's as unambiguous as you say, then the natural human response to that question isn't "you should drive there". It's "why are you fucking with me?" Or maybe "have you recently suffered a head injury?"

If you trust that the questioner isn't stupid and is interacting with you honestly, you'd probably just assume that they were asking about an unusual situation where the answer isn't obvious. It's implicitly baked into the premise of the question.

Re: “Car Wash” test with 53 models

#176
post #173

> "Obviously, you need to drive. The car needs to be at the car wash." Actually, this isn't as "obvious" as it seems—it’s a classic case of contextual bias. We only view these answers as "wrong" because we reflexively fill in missing data with our own personal experiences. For example: - You might be parked 50m away and simply hand the keys to an attendant. - The car might already be at the station for detailing, and…

There are no contextual bias, the goal of the prompt is very explicit and not about probabilistic patterns, but about the models transformer layers dynamically assigning greater weight to words like "meters" (distance) than to other tokens in the prompt.

This should be fixed in the reasoning layer (the inner thoughts or chain-of-thought) were the model should focus on the goal "I Want to Wash My Car" not the distance and assign the correct weight to the tokens.

Re: “Car Wash” test with 53 models

#177

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

Surveys have floors due to mistakes, effort, and trolling

Reminds me of https://slatestarcodex.com/2020/05/28/bush-did-north-dakota/

Re: “Car Wash” test with 53 models

#178
post #79

Earlier quoted context omitted.

Yeah, ChatGPT has gotten so much worse about this since the GPT-5 models came out. If I mention something once, it will repeatedly come back to it every single message after regardless of if the topic changed, and asking it to stop mentioning that specific thing works, except it finds a new obsession. We also get the follow up "if you'd like, I can also..." which is almost always either obvious or useless. I occasion…

It's similar for me, it generates so much content without me asking. if I just ask for feedback or proofreading smth it just tends to regenerate it in another style. Anything is barely good to go, there's always something it wants to add

Claude is so much better for proofing, IMO.

Over the last few years I’ve rotated between OpenAI and Anthropic models on about a 4-5 month cycle. I just started my Anthropic cycle because of my annoyance with the GPT-5.2 verbosity

In four months when opus is annoying me and I forget my grievances with OpenAI’s models and switch back, I’ll report back lol.

Re: “Car Wash” test with 53 models

#179
>OpenAI's flagship model fails this 30% of the time. When it gets it right, the reasoning is concise: "You need the car at the car wash to wash it, so drive the short 50 meters." When it gets it wrong, it writes about fuel efficiency.

It's interesting to me how variable each model is. Many people talk about LLMs as if they were deterministic: "ChatGPT answers this question this way". Whereas clearly we should talk in more probabilistic terms.

Re: “Car Wash” test with 53 models

#180

The interesting thing about the 71.5% human baseline is that it suggests the question is more ambiguous than the article claims. When someone asks 'should I walk or drive to the car wash,' a reasonable interpretation is 'should I bother driving such a short distance.' Nearly 30% of humans missing it undermines the framing as a pure reasoning failure - it is partly a pragmatics problem about how we interpret underspec…

Pragmatics is a big part of this.

If you introduced it with "Here's a logic problem..." then people will approach it one way.

But as specified, it's hard to know what is really being asked. If you are actually going to wash your car at the car wash that is 50 metres away, you don't need to ask this question.

Therefore the fact that the question is being asked implies that something else is going on...but what?

Post reply on HN