Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

101–110 of 469 posts

Re: “Car Wash” test with 53 models

#102
post #21

Would be interesting to see Sonnet (4.6*). It's fair bit smaller than Opus but scores pretty high on common sense, subjectively. I'm also curious about Haiku, though I don't expect it to do great. -- EDIT: Opus 4.6 Extended Reasoning > Walk it over. 50 meters is barely a minute on foot, and you'll need to be right there at the car anyway to guide it through or dry it off. Drive home after. Weird since the author says…

Interesting. I wonder if that's related to the phenomenon mentioned in the Opus 4.6 model card[1], where increased reasoning effort leads to 4.6 overthinking and convincing itself of the wrong answer on many questions. It seems to be unique to 4.6; I guess they fried it a bit too much during RL training.

[1] https://www.anthropic.com/claude-opus-4-6-system-card

Re: “Car Wash” test with 53 models

#103

Earlier quoted context omitted.

I wonder to what extent the Google search LLM is getting smarter, or simply more up-to-date on current hot topics.

Presumably it did an actual search and summarized the results and neither answered "off the cuff" by following gradients to reproduce the text it was trained on nor by following gradients to reproduce the "logic" of reasoning. [1] [1] e.g. trained on traces of a reasoning process

[dead]

Re: “Car Wash” test with 53 models

#104
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

I agree. If the LLM were truly an intelligence, it would be able to ask about this nonsense question. It would be able to ask "Why is walking even an option? Can you please explain how you imagine that would work? Do you mean hand-washing the car at home, instead?" (etc, etc) Real people can ask for clarification when things are ambiguous or confusing. Once something is clarified, they can work that into their unders…

Gemini's responses come very close to doing that when they make fun of the question (see other posts in the thread). If the model had been RL'ed to ask follow-up questions, it seems likely that it would meet your criterion.

Re: “Car Wash” test with 53 models

#105
post #93
post #62

Earlier quoted context omitted.

Are you referring to one that is more like a drive-thru where you literally pay while you're in line?

You drive up to the car wash, there's a little terminal with a screen and a card reader. You pick the program, pay for it and drive into the machine. Can't remember the last time I got out of my car when getting it washed.

Fair. I guess I'm remembering the old full service wash places where people would wash the inside as well. Maybe those barely exist anymore. I live in a city and don't have a car so my intuition is off. Not as far off as a model that has never walked, driven, or been to a car wash tho.

Re: “Car Wash” test with 53 models

#106
post #4

IMO it's not just intelligence. I think it's related to syncophancy. LLM are trained to not question the basic assumptions being made. They are horrible at telling you that you are solving the wrong problem, and I think this is a consequence of their design. They are meant to get "upvotes" from the person asking the question, so they don't want to imply you are making a fundamental mistake, even if it leads you into…

A perfectly fine, sycophantic response, that doesn't question the premises in any way, would be "That's a great question! While normally walking is better for such a short distance, you'd need to drive in this case, since you need to get the car to the car wash anyway. Do you want me to help with detailed information for other cases where the car is optional?" or some such.

AI syncophancy isn't just polite or even obsequious language, it's also "yes man" responses.

Do you want me to track down some research that shows people think information is more likely to be correct of they agree with it?

Re: “Car Wash” test with 53 models

#107
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

Thank you for saying this. It reminds me of class tests where you always had to wonder if something was a trick question and you never really knew... it was always after the teacher. Which frankly is fine in open-ended questions where you can explain your rationale or how different interpretations would lead you to different paths but a terrible situation when it comes to multiple choice. I remember being very frustrated by those

Re: “Car Wash” test with 53 models

#108
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

I agree. If the LLM were truly an intelligence, it would be able to ask about this nonsense question. It would be able to ask "Why is walking even an option? Can you please explain how you imagine that would work? Do you mean hand-washing the car at home, instead?" (etc, etc) Real people can ask for clarification when things are ambiguous or confusing. Once something is clarified, they can work that into their unders…

LLMs like the ones from Claude can ask questions and even have you pick from multiple choices or provide your own answer…

Re: “Car Wash” test with 53 models

#110
post #24

Earlier quoted context omitted.

LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.

I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.

Because inference costs are negligible compared to training costs
Post reply on HN