> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…
“Car Wash” test with 53 models
111–120 of 469 posts
Re: “Car Wash” test with 53 models
#112Earlier quoted context omitted.
I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.
Because they don't yet know how to "just stop emitting so much hot air" without also removing their ability to do anything like "thinking" (or whatever you want to call the transcript mode), which is hard because knowing which tokens are hot air is the hard problem itself. They basically only started doing this because someone noticed you got better performance from the early models by straight up writing "think step…
It would actually take more work to condense that long response into a terse one, particularly if the condensing was user specific, like "based on what you know about me from our interactions, reduce your response to the 200 words most relevant to my immediate needs, and wait for me to ask for more details if I require them."
Re: “Car Wash” test with 53 models
#113Earlier quoted context omitted.
LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.
I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.
Re: “Car Wash” test with 53 models
#114To me the only acceptable answer would be “what do you mean?” or “can you clarify?” if we were to take the question seriously to begin with. People don’t intentionally communicate with riddles and subliminal messages unless they have some hidden agenda.
How is that a "subliminal message"? It's just a simple example of common sense, which LLMs fail because they can't reason, not because they are "overthinking". If somebody asks, "What's 2+2?", they might be insulting you, but that doesn't mean the answer is anything other than 4.
And what if it’s a full service car wash and you’ve parked nearby because it’s full so you walk over and give them the keys?
Assumptions make asses of us all…
Re: “Car Wash” test with 53 models
#115Re: “Car Wash” test with 53 models
#116The question does not specify what kind of car it is. Technically speaking, a toy car (Hot wheels or a scaled model) could be walked to a car wash. Now why anyone would wash a toy car at a car wash is beyond comprehension, but the LLM is not there to judge the user's motives.
I think if surveyed at least 90% of native English speakers would understand "I want to wash my car" to mean a full size automobile. The next largest group would probably ask a clarifying question, rather than assume a toy car.
The question doesn't clearly state that the user wants to have his car washed at the car wash.
"I want to wash my car" is far less clear than "I want to have my car washed". A reasonable alternative interpretation is DIY.
Even better: "I wish to have my car washed by the crew and/or machinery at the local car wash business".
Re: “Car Wash” test with 53 models
#117This is a beautiful example of a little prompt engineering going a long way I asked Gemini and it got it wrong, then on a fresh chat I asked it again but this time asked it to use symbolic reasoning to decide. And it got it! The same applies to asking models to solve problems by scripting or writing code. Models won’t use techniques they know about unprompted - even when it’ll result in far better outcomes. Current m…
Interesting, which Gemini model? And how did you ask for symbolic reasoning, just added it to the prompt?
Re: “Car Wash” test with 53 models
#118Re: “Car Wash” test with 53 models
#119Re: “Car Wash” test with 53 models
#120Earlier quoted context omitted.
I think if surveyed at least 90% of native English speakers would understand "I want to wash my car" to mean a full size automobile. The next largest group would probably ask a clarifying question, rather than assume a toy car.
> I want to wash my car The question doesn't clearly state that the user wants to have his car washed at the car wash. "I want to wash my car" is far less clear than "I want to have my car washed". A reasonable alternative interpretation is DIY. Even better: "I wish to have my car washed by the crew and/or machinery at the local car wash business". https://imgur.com/tCSPwYp