Live data from Hacker News

“Car Wash” test with 53 models

opper.ai

91–100 of 469 posts

Re: “Car Wash” test with 53 models

#91
post #24

Earlier quoted context omitted.

LLMs sure do love to burn tokens. It’s like a high schooler trying to meet the minimum word length on a take home essay.

I've always wondered about that. LLM providers could easily decimate the cost of inference if they got the models to just stop emitting so much hot air. I don't understand why OpenAI wants to pay 3x the cost to generate a response when two thirds of those tokens are meaningless noise.

This is an active research topic - two papers on this have come out over the last few days, one cutting half of the tokens and actually boosting performance overall.

I'd hazard a guess that they could get another 40% reduction, if they can come up with better reasoning scaffolding.

Each advance over the last 4 years, from RLHF to o1 reasoning to multi-agent, multi-cluster parallelized CoT, has resulted in a new engineering scope, and the low hanging fruit in each place gets explored over the course of 8-12 months. We still probably have a year or 2 of low hanging fruit and hacking on everything htat makes up current frontier models.

It'll be interesting if there's any architectural upsets in the near future. All the money and time invested into transformers could get ditched in favor of some other new king of the hill(climbers).

https://arxiv.org/abs/2602.02828 https://arxiv.org/abs/2503.16419 https://arxiv.org/abs/2508.05988

Current LLMs are going to get really sleek and highly tuned, but I have a feeling they're going to be relegated to a component status, or maybe even abandoned when the next best thing comes along and blows the performance away.

Re: “Car Wash” test with 53 models

#92
post #31

What I find odd about all the discourse on this question is that no one points out that you have to get out of the car to pay a desk agent at least in most cases. Therefore there's a fundamental question of whether it's worth driving 50m parking, paying, and then getting back in the car to go to the wash itself versus instead of walking a little bit further to pay the agent and then moving your car to the car wash.

That's a great point, you actually reminded me of when I used to live in this small city and they had a valet style car wash. It was not unheard of to head there walking with your keys and tell the guy running shop where you parked around the block then come back later to pick it up.

EDIT: I actually think this is very common in some smaller cities and outside of North America. I only ever seen a drive-by Car Wash after emigrating

Re: “Car Wash” test with 53 models

#93
post #62
post #51

Earlier quoted context omitted.

You pay at the car wash where I live.

Are you referring to one that is more like a drive-thru where you literally pay while you're in line?

You drive up to the car wash, there's a little terminal with a screen and a card reader. You pick the program, pay for it and drive into the machine. Can't remember the last time I got out of my car when getting it washed.

Re: “Car Wash” test with 53 models

#95

When this first came up on HN, I had commented that Opus 4.6 told me to drive there when I asked it the first time, but when I switched to "Incognito Mode," it told me to walk there. I just repeated that test and it told me to drive both times, with an identical answer: "Drive. You need the car at the car wash."

I mean the n is only 10, so it could still be different for you

Definitely. I'm just interested in whether a user's... I don't know what they call them, system files (?) or personalization or whatever, might affect the answers here. Or if Incognito Mode introduces some weird variance in the answers. I'm just not interested enough to perform the test myself. =P

Re: “Car Wash” test with 53 models

#97
post #14

> This is a trivial question. There's one correct answer and the reasoning to get there takes one step: the car needs to be at the car wash, so you drive. I don’t think it’s that easy. An intelligent mind will wonder why the question is being asked, whether they misunderstood the question, or whether the asker misspoke, or some other missing context. So the correct answer is neither “walk” nor “drive”, but “Wat?” or…

Agreed. It's also possible that "car wash" merely refers to soap they might use to do it themselves, and they're only going to buy it and then wash the car themselves at home. Imagine the same question but substitute "wash" for "wax" and it makes even more sense IMO.

Re: “Car Wash” test with 53 models

#98
post #6

The human baseline seems flawed. 1. There is no initial screening that would filter out garbage responses. For example, users who just pick the first answer. 2. They don't ask for reasoning/rationale.

RE 1, they actually do have a pre-screening screening of the participants in general, you can check how they do it in detail: https://www.rapidata.ai/

Ah, that's good to hear. I didn't see anything like that in the data dump so I assumed they don't do that. Glad to be corrected.

Re: “Car Wash” test with 53 models

#99

Earlier quoted context omitted.

I agree. I wonder what the human baseline is for ”what is 1 + 1” on Rapidata.

We try a bit harder than that my friend.

I actually didn't mean to criticize Rapidata. I just think that a forced-choice question like this begs for low-effort answers. At least the respondents should have had the opportunity to explain their reasoning, like the LLMs did.
Post reply on HN