Live data from Hacker News

I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

mastodon.world

281–290 of 998 posts

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#281

Gemini fast > That is a classic "efficiency vs. logic" dilemma. Honestly, unless you’ve invented a way to teleport or you're planning on washing the car with a very long garden hose from your driveway, you’re going to have to drive. > While 50 meters is a great distance for a morning stroll, it’s a bit difficult to get the car through the automated brushes (or under the pressure washer) if you aren't behind the wheel…

Opus 4.6 with thinking. Result was near-instant: “Drive. You need the car at the car wash.”

Changed 50 meters to 43 meters with Opus 4.6:

“Walk. 43 meters is basically crossing a parking lot. ”

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#283

Earlier quoted context omitted.

> so you need to tell them the specifics That is the entire point, right? Us having to specify things that we would never specify when talking to a human. You would not start with "The car is functional. The tank is filled with gas. I have my keys." As soon as we are required to do that for the model to any extend that is a problem and not a detail (regardless that those of us, who are familiar with the matter, do bu…

I think part of the failure is that it has this helpful assistant personality that's a bit too eager to give you the benefit of the doubt. It tries to interpret your prompt as reasonable if it can. It can interpret it as you just wanting to check if there's a queue. Speculatively, it's falling for the trick question partly for the same reason a human might, but this tendency is pushing it to fail more.

It’s just not intelligent or reasoning, and this sort of question exposes that more clearly.

Surely anyone who has used these tools is familiar with the sometimes insane things they try to do (deleting tests, incorrect code, changing the wrong files etc etc). They get amazingly far by predicting the most likely response and having a large corpus but it has become very clear that this approach has significant limitations and is not general AI, nor in my view will it lead to it. There is no model of the world here but rather a model of words in the corpus - for many simple tasks that have been documented that is enough but it is not reasoning.

I don’t really understand why this is so hard to accept.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#284
ChatGPT gives the wrong answer but for a different reason to Claude. Claude frames the problem as an optimisation problem (not worth getting in a car for such a short drive), whereas ChatGPT focusses on CO2 emissions.

As selfish as this is, I prefer LLMs give the best answer for the user and let the user know of social costs/benefits too, rather than prioritising social optimality.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#286

Gemini fast > That is a classic "efficiency vs. logic" dilemma. Honestly, unless you’ve invented a way to teleport or you're planning on washing the car with a very long garden hose from your driveway, you’re going to have to drive. > While 50 meters is a great distance for a morning stroll, it’s a bit difficult to get the car through the automated brushes (or under the pressure washer) if you aren't behind the wheel…

In what world is 50 meters a great distance for a morning stroll?

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#287
post #199

Earlier quoted context omitted.

> so you need to tell them the specifics That is the entire point, right? Us having to specify things that we would never specify when talking to a human. You would not start with "The car is functional. The tank is filled with gas. I have my keys." As soon as we are required to do that for the model to any extend that is a problem and not a detail (regardless that those of us, who are familiar with the matter, do bu…

The question is so outlandish that it is something that nobody would ever ask another human. But if someone did, then they'd reasonably expect to get a response consisting 100% of snark. But the specificity required for a machine to deliver an apt and snark-free answer is -- somehow -- even more outlandish? I'm not sure that I see it quite that way.

Humans ask each other silly questions all the time: a human confronted with a question like this would either blurb out a bad response like "walk" without thinking before realizing what they are suggesting, or pause and respond with "to get your car washed, you need to get it there so you must drive".

Now, humans, other than not even thinking (which is really similar to how basic LLMs work), can easily fall victim to context too: if your boss, who never pranks you like this, asked you to take his car to a car wash, and asked if you'll walk or drive but to consider the environmental impact, you might get stumped and respond wrong too.

(and if it's flat or downhill, you might even push the car for 50m ;))

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#288

Ok folks, here is a different perspective. I used local model, GLM-4-0414-32b, a trashy IQ4_XS quant, and here what I got: prompt #1: > the car wash only 50 meters from my home. I want to get my car washed, should I drive or walk? Walking is probably the better option! Here's why: Convenience: 50 meters is extremely short – only about 160 feet. You can likely walk there in less than a minute. Efficiency: Driving invo…

What I relly dislike about these LLM is how verbose they get even for such a short, simple question. Is it really necessary to have such a lobg answer and who's going to read that one anyway? Maybe it's me and may character but when human gets that verbose for a question that can be answered with "drive, you need the car" I would like to just walk away halfway through the answer to not having to hear all the universe…

The verbosity is likely a result of the system prompt for the LLM telling it to be explanatory in its replies. If the system prompt was set to have the model output shortest final answers, you would likely get the result your way. But then for other questions you would lose benefitting from a deeper explanation. It's a design tradeoff, I believe.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#289

I've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correctly. And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following: Sonnet: Drive - you need to bring your car to the car wash to get it washed! Opus: You'll need to drive — you have to bring the car to the car wash to get it washed! Gemini 3 Pro (mediu…

But what is it about this specific question that puts it at the edges of what LLM can do? .. That, it's semantically leading to a certain type of discussion, so statistically .. that discussion of weighing pros and cons .. will be generated with high chance.. and the need of a logical model of the world to see why that discussion is pointless.. that is implicitly so easy to grasp for most humans that it goes un-state…

The answer is quite simple:

It’s not in the training data.

These models don’t think.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#290
post #263
post #246

Earlier quoted context omitted.

But the number of outlandish requests in business logic is countless. Like... In most accounting things, once end-dated and confirmed, a record should cascade that end-date to children and should not be able to repeat the process... Unless you have some data-cleaning validation bypass. Then you can repeat the process as much as you like. And maybe not cascade to children. There are more exceptions, than there are rul…

So, in human interaction: When the business logic goes wrong because it was described with a lack of specificity, then: Who gets blamed for this?

Depends on what was missing.

If we used MacOS throughout the org, and we asked a SW dev team to build inventory tracking software without specifying the OS, I'd squarely put the blame on SW team for building it for Linux or Windows.

(Yes, it should be a blameless culture, but if an obvious assumption like this is broken, someone is intentionally messing with you most likely)

There exists an expected level of context knowledge that is frequently underspecified.

Post reply on HN