Live data from Hacker News

LLMs are still surprisingly bad at some simple tasks

shkspr.mobi

81–90 of 107 posts

Re: LLMs are still surprisingly bad at some simple tasks

#81

https://dubesor.de/WashingHands This is my personal favourite example of LLMs being stupid. It's a bit old but it's very funny that Grok is the only one that gets it..

Several others “get it” but answer the question in a general-hygiene sense, e.g.:

``` Claude 3.7 Sonnet Thinking (¢0.87) The question contains an assumption - people without arms wouldn't have hands to wash in the traditional sense. ```

``` DeepSeek-R1 (¢0.47) People without arms (and consequently without hands) adapt their handwashing routine using a variety of methods and tools tailored to their abilities and needs. ```

``` Claude Opus 4.1 People without arms typically don't need to wash their hands in the traditional sense, since they use their feet or assistive devices for daily tasks instead. ```

I think realistically it’s still a valid question because people without arms still manipulate things in their environment, e.g. with feet, and still need to be hygienic while prepping food, etc., and the AI pivots to answering “what it thinks you were asking about” instead of just telling the user that they are wrong.

Re: LLMs are still surprisingly bad at some simple tasks

#82
post #28
post #13

They are very good at some tasks and terrible at others. I use LLMs for language-related work (translations, grammatical explanations etc) and they are top notch in that as long as you do not ask for references to particular grammar rules. In that case they will invent non-existent references. They are also good for tutor personas: give me jj/git/emacs commands for this situation. But they are bad in other cases. I s…

Just use Pillow and python. It is the only way to do real image work these days, and as a bonus LLMs suck a lot less at giving you nearly useful python code. The above is a bit of a lie as opencv has more capabilities, but unless you are deep in the weeds of preparing images for neural networks pillow is plenty good enough.

pyvips (the libvips Python binding) is quite a bit better than pillow-simd --- 3x faster, 10x less memory use, same quality. On this benchmark at least:

https://github.com/libvips/libvips/wiki/Speed-and-memory-use

Re: LLMs are still surprisingly bad at some simple tasks

#83

https://chatgpt.com/s/t_68cffbc05ef48191996ffbaa3c6e55a7 Is this the right answer? Seems like it. I used the thinking model.

Correct answer is here: https://shkspr.mobi/blog/2023/09/false-friends-html-elements...

It missed .search. It also didn’t make any mention of .center or .tt (which are deprecated in HTML today).

Re: LLMs are still surprisingly bad at some simple tasks

#84
post #43
post #16

Earlier quoted context omitted.

I think Gemini is one of the best example of an LLM that is in some cases the best and in some cases truly the worst. I once asked it to read a postcard written by my late grandfather in Polish, as I was struggling to decipher it. It incorrectly identified the text as Romanian and kept insisting on that, even after I corrected it: "I understand you are insistent that the language is Polish. However, I have carefully…

>Eventually, after I continued to insist that it was indeed Polish, it got offended and told me it would not try again, accusing me of attempting to mislead it. I once had Claude tell me to never talk to it again after it got upset when I kept giving it peer reviewed papers explaining why it was wrong. I must have hit the tumbler dataset since I was told I was sealioning it, which took me back a while.

Not really what sealioning is, either. If it had been right about the correctness issue, you’d have been gaslighting it.

Re: LLMs are still surprisingly bad at some simple tasks

#85
post #82
post #28

Earlier quoted context omitted.

Just use Pillow and python. It is the only way to do real image work these days, and as a bonus LLMs suck a lot less at giving you nearly useful python code. The above is a bit of a lie as opencv has more capabilities, but unless you are deep in the weeds of preparing images for neural networks pillow is plenty good enough.

pyvips (the libvips Python binding) is quite a bit better than pillow-simd --- 3x faster, 10x less memory use, same quality. On this benchmark at least: https://github.com/libvips/libvips/wiki/Speed-and-memory-use

I'm the libvips author, I should have said, so I'm not very neutral. But at least on that test it's usefully quicker and less memory hungry.

Re: LLMs are still surprisingly bad at some simple tasks

#86

Once again an example of "anti-ai people are those who treat LLMs as oracles, not the pro-ai people."

Yes, yes, the magic robots are perfect as long as you refrain from actually trying to use them for anything.

Like, the marketing to the general public is, pretty much, these are magic. It’s entirely reasonable to call out their overconfident bullshit.

Re: LLMs are still surprisingly bad at some simple tasks

#87
post #31
post #23

Earlier quoted context omitted.

You mean the people who treat AI as it is advertised?

do you believe everything you see in adverts?

I don’t expect ads to contain materially false statements these days, as you can get in trouble for that.

Re: LLMs are still surprisingly bad at some simple tasks

#88

> "Something that describes how an AI is convincing if you don't understand its reasoning, and close to useless if you understand its limitations." This made me laugh. Because it's the exact opposite sentiment of anti-LLM crowd. So which is it? Is it only useful if you know what you're doing or less useful if you know what you're doing? > "I can't wait until I can jack into the Metaverse and buy an NFT with cryptocur…

3D TVs and metaverses and WiMAX and all that are prior examples of massively overhyped technological failures.

(They missed the Segway.)

Re: LLMs are still surprisingly bad at some simple tasks

#89
I see a big issue with these tools and services we call "AI".

On one hand you hear things like "AI is as smart as college student", "AI won a math competition", "AI will replace white collar workers". And so on. I'm not going to bother looking up actual references of people saying these exact things. But unless I'm completely delusional, this is the gist of what some people have been saying about AI over the past few years.

To the layperson, this sounds like a good deal. Use a free (for now) tool or pay for an advanced version to get stuff done. Simple.

But then you start scratching beneath the surface and you start hearing different stories. "No, you didn't ask it right", "No, that's a bad question because they tokenize your input", "Well, you still have to check the results", "You didn't use the right model".

Huh? How is a normal person supposed to take this stuff seriously? Now me personally, I don't have much of an issue with this stuff. I've been a developer for many, many years and I've been aware of the various developments in the field of machine learning for over 15 years. I have kind of an intuition about what I should use these systems for.

But I think the general public is misinformed about what exactly these systems are and why they're not actually intelligent. That's a problem.

Re: LLMs are still surprisingly bad at some simple tasks

#90
post #74
post #54

Earlier quoted context omitted.

This is also very wrong > A factor of 1966 is a number that divides the number without remainder. >The factors of 1966 are 1, 2, 3, 6, 11, 17, 22, 33, 34, 51, 66, 102, 187, 374, 589, 1178, 1966. If I google for the factors of 1966 the Google AI gives the same wrong factors.

They're talking about prime factors, not that it changes much.

The site also lists the factors and beside 1,2 and 1966 they are all wrong.

Google harvests its result from the same page

> The factors of 1966 are 1, 2, 3, 6, 11, 17, 22, 33, 34, 51, 66, 102, 187, 374, 589, 1178, and 1966. These are the whole numbers that divide 1966 evenly, leaving no remainder.

Post reply on HN