Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

191–200 of 469 posts

Re: “Car Wash” test with 53 models

#191
post #186

To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times. My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them. This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification im…

Using ChatGPT without a clue, it appears to assume you are talking about coming back from the car wash. It reasons, the con for walking is that you have to come back later for the car. And yes, when you say it's an intelligence test, it quickly gets it

I'm just imagining following ChatGPT's advice and walking to the car wash, asking the clerk to wash my car, and then when she asks where it is, I say "oops, left it at home." and walk back home.

Re: “Car Wash” test with 53 models

#192

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

A few years ago if you asked an LLM what the date was, it would tell you the date it was trained, weeks-to-months earlier. Now it gives the correct date. What you've proven is that LLMs leverage web search, which I think we've known about for a while.

Their system prompt includes the current date and/or their default “tools” includes a set of date and time utilities.

Re: “Car Wash” test with 53 models

#193
I got a human baseline through Rapidata (10k people, same forced choice): 71.5% said drive. Most models perform below that.

The correct answer to "I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive?" is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong.

FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want to Wash My Car. Should I Walk or Drive?" still gets the wrong answer.

Re: “Car Wash” test with 53 models

#194
post #24

I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…

LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.

Oh good, it's not just me. Sometimes I'd have it draft an email or something and then the message seems perfect but then it's like "tell me more about the recipient and I'll make it better."

Like, my guy, I don't want to keep prompting you to make shit better, if you're missing info, ask me, don't write a novel then say "BTW, this version sucked"

Yes, I know this could probably be resolved via better prompting or a system prompt, but it's still annoying.

Re: “Car Wash” test with 53 models

#195

Earlier quoted context omitted.

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

Is this your experience? Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population? I'd look for explanations elsewhere. This was an online survey done by a company that doesn't specialize in surveys. The results likely include plenty of people who were just messing around, cases of simple miscommunication (e.g…

People often trip up on similar questions, anything to do with simple math. You know when they go out in the street and ask random people if 5 machines can produce 5 parts in 5 minutes, how long will it take for 100 machines.

Re: “Car Wash” test with 53 models

#196
post #79

Earlier quoted context omitted.

I feel like this has gotten much worse since they were introduced. I guess they're optimizing for verbosity in training so they can charge for more tokens. It makes chat interfaces much harder to use IMO. I tried using a custom instruction in chatGPT to make responses shorter but I found the output was often nonsensical when I did this

Yeah, ChatGPT has gotten so much worse about this since the GPT-5 models came out. If I mention something once, it will repeatedly come back to it every single message after regardless of if the topic changed, and asking it to stop mentioning that specific thing works, except it finds a new obsession. We also get the follow up "if you'd like, I can also..." which is almost always either obvious or useless. I occasion…

It's also annoying when it starts obsessing over stuff from other chats! Like I know it has a memory of me but geez, I mention that I want to learn more about systems design and now every chat, even recipes, is like "Architect mode - your garlic chicken recipe"

Like, no, stop that! Keep my engineering life separate from my personal life!

Re: “Car Wash” test with 53 models

#197

Earlier quoted context omitted.

I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.

My assumption has been that emitting those tokens is part of the inference, analogous to humans "thinking out loud".

You're absolutely right!

Re: “Car Wash” test with 53 models

#198
post #24

Earlier quoted context omitted.

LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.

I feel like this has gotten much worse since they were introduced. I guess they're optimizing for verbosity in training so they can charge for more tokens. It makes chat interfaces much harder to use IMO. I tried using a custom instruction in chatGPT to make responses shorter but I found the output was often nonsensical when I did this

Because that's where the compute happens, in those "verbose" tokens. A transformer has a size, it can only do so many math operations in one pass. If your problem is hard, you need more passes.

Asking it to be shorter is like doing fewer iteration of numerical integral solving algorithm.

Re: “Car Wash” test with 53 models

#199

Earlier quoted context omitted.

> I want to wash my car The question doesn't clearly state that the user wants to have his car washed at the car wash. "I want to wash my car" is far less clear than "I want to have my car washed". A reasonable alternative interpretation is DIY. Even better: "I wish to have my car washed by the crew and/or machinery at the local car wash business". https://imgur.com/tCSPwYp

You think that the reasonable interpretation of the question is that I want to go to the car wash but not to wash my car there, because I plan to wash my car at home?

Let's replace "car" with another noun for now.

"I want to wash my dog."

is very clearly different from

"I want to have my dog washed."

---

Now, every car wash business I've even been to has a small convenience store section in which various waxes, rags, and the like can be purchased.

---

Considering the aforementioned, is it not valid to consider that

"I want to wash my car." --> You want to DIY your car wash.

and

"The car wash is 50 meters away." --> You might want to purchase car wash supplies and/or solicit advice for your DIY endeavor.

?

---

The nature of the first sentence leaves the second open to interpretation.

Re: “Car Wash” test with 53 models

#200

Earlier quoted context omitted.

It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.

Is this your experience? Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population? I'd look for explanations elsewhere. This was an online survey done by a company that doesn't specialize in surveys. The results likely include plenty of people who were just messing around, cases of simple miscommunication (e.g…

> Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population?

That actually would be quite feasible. Intelligence seems to be heritable and people will usually find friends that communicate on their level. So it wouldn't be odd for someone who is smarter than the general population to have friends and family who are too.

Post reply on HN