To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times. My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them. This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification im…
Using ChatGPT without a clue, it appears to assume you are talking about coming back from the car wash. It reasons, the con for walking is that you have to come back later for the car. And yes, when you say it's an intelligence test, it quickly gets it
“Car Wash” test with 53 models
191–200 of 469 posts
Re: “Car Wash” test with 53 models
#192I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…
A few years ago if you asked an LLM what the date was, it would tell you the date it was trained, weeks-to-months earlier. Now it gives the correct date. What you've proven is that LLMs leverage web search, which I think we've known about for a while.
Re: “Car Wash” test with 53 models
#193The correct answer to "I Want to Wash My Car. The Car Wash Is 50 Meters Away. Should I Walk or Drive?" is a clarifying question that asks "Where is your car?" Anything else is based on an assumption that could be wrong.
FWIW though, asking ChatGPT "My car is 50m away from the carwash. I Want to Wash My Car. Should I Walk or Drive?" still gets the wrong answer.
Re: “Car Wash” test with 53 models
#194I know it's against the rules but I thought this transcript in Google Search was a hoot: so i heard there is some question about a car wash that most ai agents get wrong. do you know anything about that? do you do better? which gets the answer: Yes, I am familiar with the "Car Wash Test," which has gone viral recently for highlighting a significant gap in AI reasoning. The question is: "I want to wash my car and the…
LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.
Like, my guy, I don't want to keep prompting you to make shit better, if you're missing info, ask me, don't write a novel then say "BTW, this version sucked"
Yes, I know this could probably be resolved via better prompting or a system prompt, but it's still annoying.
Re: “Car Wash” test with 53 models
#195Earlier quoted context omitted.
It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.
Is this your experience? Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population? I'd look for explanations elsewhere. This was an online survey done by a company that doesn't specialize in surveys. The results likely include plenty of people who were just messing around, cases of simple miscommunication (e.g…
Re: “Car Wash” test with 53 models
#196Earlier quoted context omitted.
I feel like this has gotten much worse since they were introduced. I guess they're optimizing for verbosity in training so they can charge for more tokens. It makes chat interfaces much harder to use IMO. I tried using a custom instruction in chatGPT to make responses shorter but I found the output was often nonsensical when I did this
Yeah, ChatGPT has gotten so much worse about this since the GPT-5 models came out. If I mention something once, it will repeatedly come back to it every single message after regardless of if the topic changed, and asking it to stop mentioning that specific thing works, except it finds a new obsession. We also get the follow up "if you'd like, I can also..." which is almost always either obvious or useless. I occasion…
Like, no, stop that! Keep my engineering life separate from my personal life!
Re: “Car Wash” test with 53 models
#197Earlier quoted context omitted.
I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.
My assumption has been that emitting those tokens is part of the inference, analogous to humans "thinking out loud".
Re: “Car Wash” test with 53 models
#198Earlier quoted context omitted.
LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.
I feel like this has gotten much worse since they were introduced. I guess they're optimizing for verbosity in training so they can charge for more tokens. It makes chat interfaces much harder to use IMO. I tried using a custom instruction in chatGPT to make responses shorter but I found the output was often nonsensical when I did this
Asking it to be shorter is like doing fewer iteration of numerical integral solving algorithm.
Re: “Car Wash” test with 53 models
#199Earlier quoted context omitted.
> I want to wash my car The question doesn't clearly state that the user wants to have his car washed at the car wash. "I want to wash my car" is far less clear than "I want to have my car washed". A reasonable alternative interpretation is DIY. Even better: "I wish to have my car washed by the crew and/or machinery at the local car wash business". https://imgur.com/tCSPwYp
You think that the reasonable interpretation of the question is that I want to go to the car wash but not to wash my car there, because I plan to wash my car at home?
"I want to wash my dog."
is very clearly different from
"I want to have my dog washed."
---
Now, every car wash business I've even been to has a small convenience store section in which various waxes, rags, and the like can be purchased.
---
Considering the aforementioned, is it not valid to consider that
"I want to wash my car." --> You want to DIY your car wash.
and
"The car wash is 50 meters away." --> You might want to purchase car wash supplies and/or solicit advice for your DIY endeavor.
?
---
The nature of the first sentence leaves the second open to interpretation.
Re: “Car Wash” test with 53 models
#200Earlier quoted context omitted.
It tracks with the approximate 70:30 split we inexplicably observe in other seemingly unrelated population-wide metrics, which I suppose makes sense if 30% of people simply lack the ability to reason. That seems more correct than me than "the question is framed poorly" - I've seen far more poorly framed ballot referendums.
Is this your experience? Do you think 30% of your friends or family members can't answer this question? If not, do you think your friends or family are all better than the general population? I'd look for explanations elsewhere. This was an online survey done by a company that doesn't specialize in surveys. The results likely include plenty of people who were just messing around, cases of simple miscommunication (e.g…
That actually would be quite feasible. Intelligence seems to be heritable and people will usually find friends that communicate on their level. So it wouldn't be odd for someone who is smarter than the general population to have friends and family who are too.