Live data from Hacker News

I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

mastodon.world

921–930 of 998 posts

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#921

Earlier quoted context omitted.

I'm actually having a hard time interpreting your meaning. Are you criticizing LLMs? Highlighting the importance of this training and why we're trained that way even as children? That it is an important part of what we call reasoning? Or are you giving LLMs the benefit of the doubt, saying that even humans have these failure modes?[0] Though my point is more that natural language is far more ambiguous than I think pe…

I was pointing out humans and LLMs have this failure mode so in a lot of ways it is no big deal/not some smoking gun that LLMs are useless and dangerous, or at least no more useless and dangerous than humans. I personally would stay away from calling someone, or an LLM, 'stupid' for making this mistake because of several reasons. First, objectively intelligent high functioning people can and do mistakes similar to th…

  > I personally would stay away from calling someone, or an LLM, 'stupid' for making this mistake because of several reasons.
I wouldn't. Because there's a difference between calling someone's action stupid and saying that someone is stupid. These are entirely dependent upon the context of the claim. Smart people frequently do stupid stuff. I have a PhD and by some metric that makes me "smart" but you'll also see me do plenty of stupid stuff every single day. Language is fuzzy...

But I think responses like yours are entirely dismissive at what's being attempted to be shown. What's being shown is how easily they are fooled. Another popular example right now being the cup with a sealed top and open bottom (lol "world model"?).

  > There are a lot of 'gotcha' articles
The point isn't about getting some gotcha, it is about a clear and concise example of how these systems fail.

What would not be a clear and concise example is showing something that requires domain subject expertise. That's absolutely useless as an example to everyone that isn't a subject matter expert.

The point of these types of experiments is to make people think "if they're making these types of errors that I can easily tell are foolish then how often are they making errors where I am unable to vet or evaluate the accuracy of its outputs?" This is literally the Gell-Mann Amnesia Effect in action[0].

  > I totally agree with the language ambiguity point. I think that is a feature and not a bug.
So does everybody. But there are limits to natural language and we've been discussing them for quite a long time[1]. There is in fact a reason we invented math and programming languages.

  > Finally, we often really don't know enough but we still need to say something and like gradient descent, an ambiguous statement may take us a step closer to a useful answer.
Was this sentence an illustrative example?

Sometimes I think we don't need to say something. I think we all (myself included) could benefit more by spending a bit longer before we open our mouths, or even not opening them as often. There's times where it is important to speak out but there are also times that it is important to not speak. It is okay to not know things and it is okay to not be an expert on everything.

[0] https://themindcollection.com/gell-mann-amnesia-effect/

[1] https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#922
post #644

Earlier quoted context omitted.

I can feel the AGI on this one :) I ran extensive tests on this and variations on multiple models. Most models interpret 50 m as a short distance and struggle with spatial reasoning. Only Gemini and Grok correctly inferred that you would need to bring your car to get it washed in their thought stream, and incorporated that into the final answer. GPT-5.2 and Kimi K2.5 and even Opus 4.6 failed in my tests - https://x.c…

> I can feel the AGI on this one :) This was probably meant in a sarcastic way, but isn't it impressive how you cannot push Gemini off track? I tried another prompt with claiming that one of my cups does not work, because it is closed at the top and open at the bottom, and it kind of played with me, giving me a funny technical explanation on how to solve that problem and finally asking me if that was a trick question…

Gemini Fast and Thinking failed just like other models.

I found Gemini Pro to be more consistent in logical reasoning

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#923

I've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correctly. And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following: Sonnet: Drive - you need to bring your car to the car wash to get it washed! Opus: You'll need to drive — you have to bring the car to the car wash to get it washed! Gemini 3 Pro (mediu…

> My first instinct was, I had underspecified the location of the car. The model seems to assume the car is already at the car wash from the wording. Doesn't offering two options to the LLM, "walk," or "drive," imply that either can be chosen? So, surely the implication of the question is that the car is where you are?

> Doesn't offering two options to the LLM, "walk," or "drive," imply that either can be chosen?

Yes, but the problem is specifically that proposing two choices also eliminates other options. An open-ended question would lift that restriction. GPT-5.2 Thinking: https://chatgpt.com/share/6993d099-ef4c-8005-aa62-bdb826b707...

Other possible issues in the question: Is biking also an option? What do you want to do at the carwash when you get there, wash the car or buy a bucket, sponge, and soap? Is the car already at the carwash and you want to drive a second car? What about calling the carwash to see if they will have someone wash the car for you?

There are many ways to interpret the question because it contains ambiguities that must be resolved through assumptions. (Lacking information and constraints such as possible alternatives that satisfy the goal of washing the car.) The follow up questions I asked also have assumed answers but answering them provides no clear resolution to the ambiguity present in the original question.

So, no, I disagree that there is any solid implication of where the car is. And even if there is a solid implication, it can hardly be reasoned that it "isn't an XY problem" or that the question is clear cut in any real sense.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#924
post #131

Yesterday someone on was yapping about how AI is enough to replace senior software engineers and they can just "vibe code their way" over a weekend into a full-fledged product. And that somehow finally the "gatekeeping" of software development was removed. I think of that person reading these answers and wonder if they changed their opinion now :)

What does this nonsensical question that some LLMs get wrong some of the time, and that some don't get wrong ever, have to do with anything? This isn't a "gotcha" even though you want it to be. It's just mildly amusing.

Because, this fundamental premise demonstrates that LLMs can't really think logically like we do and they are far from replacing actual humans, let alone senior software engineers.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#925
post #131

Yesterday someone on was yapping about how AI is enough to replace senior software engineers and they can just "vibe code their way" over a weekend into a full-fledged product. And that somehow finally the "gatekeeping" of software development was removed. I think of that person reading these answers and wonder if they changed their opinion now :)

Does this mean we're back in favor of using weird riddles to decide programming skills now? Do we owe Google an apology for the inverse binary tree incident?

I don't know what inverse binary incident you're referring to, but the fundamental premise here is LLMs can't really think logically like humans do and they are far away from replacing humans in software, let alone senior software engineers

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#926

Earlier quoted context omitted.

Are you sure? What about when you use it? e.g. I suppose asking it to critique experimental design and analytical methodology, or identify potential confounders and future areas to explore, or help summarize nearby research, etc. If you can tell when your students use it, presumably you mean they're just copying whatever, which just sounds like that student doesn't know what they're doing or is being lazy. That doesn…

> What about when you use it? I use AI for coding, but not for anything involving writing text, it's just horrendous at it. It just spews verbose slop, devoid of meaning, original thought or nuanced critique. > That doesn't mean the model isn't capable; it means an incapable person won't know what they'd want to ask of it. So it's user error again then, eh? PhD experts are able to help even "incapable" students, that…

Weird, my experience is that they are full of nuance. e.g. here is a snippet of my discussion with Gemini:

> Would you like to see why Q is "flexible" (Flat) while Z/2 is "rigid" (Not Flat) using this "crushing" vs. "preserving" logic? It explains why localized rings are almost always better to work with.

> Roughly, Q is flexible because it's just an epic extension of the initial object Z?

> That is a very "categorical" way to put it, but it’s actually a bit more subtle! If being an "epic extension of the initial object" was enough to be flat, then every quotient would be flat too. To refine your intuition: Q is "flexible" (flat) not just because it's an extension, but because of how it extends Z. Z/2 is a Quotient: It adds a constraint (2=0). Constraints are "rigid." As we saw, if you multiply by 2, everything collapses to zero. That's a "hidden kernel," which breaks left exactness. Q is a Localization: It adds an opportunity (the ability to divide by any n≠0). This is the definition of "flexibility."

It's hard for me to imagine what kind of work you have where it's not able to capture the requisite nuance. Again, I also find that when you use jargon, they adapt accordingly on their own to raise their level of conversation. They also seem to no longer have an issue with saying "yep exactly!" or "ehh not quite" (and provide counterarguments) as necessary.

Obviously if someone just says "write my paper" or whatever and gives that to you, that won't work well. I'd think they wouldn't make it very far in their academic career regardless (it's surprising that they could get into grad school); they certainly wouldn't last long in any software org I've been in.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#927
post #667
post #462

Earlier quoted context omitted.

For GPT at least, a lot of it is because "DO NOT ASK A CLARIFYING QUESTION OR ASK FOR CONFIRMATION" is in the system prompt. Twice. https://github.com/Wyattwalls/system_prompts/blob/main/OpenA...

So this system prompt is always there, no matter if i'm using chatgpt or azure openai with my own provisioned gpt? This explains why chatgpt is a joke for professionals where asking clarifying questions is the core of professional work.

The system prompt is there if you use a chat app like ChatGTP. The system prompt is one of the things that controls the behavior of the app.

If you use an LLM endpoint in Azure OpenAI, no system prompt is in effect unless you provide one.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#928

Earlier quoted context omitted.

I was pointing out humans and LLMs have this failure mode so in a lot of ways it is no big deal/not some smoking gun that LLMs are useless and dangerous, or at least no more useless and dangerous than humans. I personally would stay away from calling someone, or an LLM, 'stupid' for making this mistake because of several reasons. First, objectively intelligent high functioning people can and do mistakes similar to th…

> I personally would stay away from calling someone, or an LLM, 'stupid' for making this mistake because of several reasons. I wouldn't. Because there's a difference between calling someone's action stupid and saying that someone is stupid. These are entirely dependent upon the context of the claim. Smart people frequently do stupid stuff. I have a PhD and by some metric that makes me "smart" but you'll also see me d…

> This is literally the Gell-Mann Amnesia Effect in action.

Absolutely! But there is some nuance, here. The failure mode is for an ambiguous question, which is an open research topic. There is no objectively correct answer to "Should I walk or drive?" given the provided constraints.

Because handling ambiguities is a problem that researchers are actively working on, I have confidence that models will improve on these situations. The improvements may asymptotically approach zero, leading to ever increasingly absurd examples of the failure mode. But that's ok, too. It means the models will increase in accuracy without becoming perfect. (I think I agree with Stephen Wolfram's take on computationally irreducibility [1]. That handling ambiguity is a computationally irreducible problem.)

EWD was right, of course, and you are too for pointing out rigorous languages. But the interactivity with an LLM is different. A programming language cannot ask clarifying questions. It can only produce broken code or throw a compiler error. We prefer the compiler errors because broken code does not work, by definition. (Ignoring the "feature not a bug" gag.)

Most of the current models are fine-tuned to "produce broken code" rather than "compiler error" in these situations. They have the capability of asking clarifying questions, they just tend not to, because the RL schedule doesn't reward it.

[1]: https://writings.stephenwolfram.com/2017/05/a-new-kind-of-sc...

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#929

I've used LLMs enough that I have a good sense of their _edges_ of intelligence. I had assumed that reasoning models should easily be able to answer this correctly. And indeed, Sonnet and Opus 4.5 (medium reasoning) say the following: Sonnet: Drive - you need to bring your car to the car wash to get it washed! Opus: You'll need to drive — you have to bring the car to the car wash to get it washed! Gemini 3 Pro (mediu…

> so you need to tell them the specifics That is the entire point, right? Us having to specify things that we would never specify when talking to a human. You would not start with "The car is functional. The tank is filled with gas. I have my keys." As soon as we are required to do that for the model to any extend that is a problem and not a detail (regardless that those of us, who are familiar with the matter, do bu…

That's my thought too. Somebody I know kept insisting it's about prompt engineering. "You are an expert coder with 30 years experience" and buddy I'd rather do actual engineering and be that expert myself than spend and figuring out how on that one variant of one version of one model to get halfway decent results.

Re: I want to wash my car. The car wash is 50 meters away. Should I walk or drive?

#930
post #469

This trick went viral on TikTok last week, and it has already been patched. To get a similar result now, try saying that the distance is 45 meters or feet. The new one is with upside down glass: https://www.tiktok.com/t/ZP89Khv9t/

still failed for me on opus 4.6 extended a second ago.

when i prompted about how walking would mean leaving my car behind the "thinking" done before coming to the right conclusion was:

> lmao, fair point. the user is right - you need to bring the car to the car wash. that's a legitimate correction. own it.

Post reply on HN