Once again an example of "anti-ai people are those who treat LLMs as oracles, not the pro-ai people."
You mean the people who treat AI as it is advertised?
LLMs are still surprisingly bad at some simple tasks
31–40 of 107 posts
Re: LLMs are still surprisingly bad at some simple tasks
#32The point from the end of the post that AI produces output that sounds correct is exactly what I try to emphasize to friends and family when explaining appropriate uses of LLMs. AI is great at tasks where sounding correct is the essence of the task (for example "change the style of this text"). Not so great when details matter and sounding correct isn't enough, which is what the author here seems to have rediscovered…
> When LLMs say something true, it’s a coincidence of the training data that the statement of fact is also a likely sequence of words; Do you know what a "coincidence" actually is? The definition you're using is wrong. It's not a coincidence that I train a model on healthcare regulations and it answers a question about healthcare regulations correctly. None of that is coincidental. If I trained it on healthcare regul…
If you train a model on only healthcare regulations it wont answer questions about healthcare regulation, it will produce text that looks like healthcare regulations.
Re: LLMs are still surprisingly bad at some simple tasks
#33Re: LLMs are still surprisingly bad at some simple tasks
#34Is this the right answer? Seems like it. I used the thinking model.
Re: LLMs are still surprisingly bad at some simple tasks
#35Re: LLMs are still surprisingly bad at some simple tasks
#36> “To stave off some obvious comments: > yoUr'E PRoMPTiNg IT WRoNg! > Am I though?” Yes. You’re complaining that Gemini “shits the bed”, despite using 2.5 Flash (not Pro), without search or reasoning. It’s a fact that some models are smarter than others. This is a task that requires reasoning so the article is hard to take seriously when the author uses a model optimised for speed (not intelligence), and doesn’t even…
OP here. I literally opened up Gemini and used the defaults. If the defaults are shit, maybe don't offer them as the default? Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?" Either way, disappointing.
> Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?"
That's literally what ChatGPT did for me[0], which is consistent from what they shared at the last keynote (quick-low reasoning answer per default first, with reasoning/search only if explicitly prompted or as a follow-up). It did miss one match tough, as it somehow didn't parse the `` element from the MDN docs.
[0]: https://chatgpt.com/share/68cffb5c-fd14-8005-b175-ab77d1bf58...
Re: LLMs are still surprisingly bad at some simple tasks
#37Earlier quoted context omitted.
OP here. I literally opened up Gemini and used the defaults. If the defaults are shit, maybe don't offer them as the default? Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?" Either way, disappointing.
You are pointing out a maturity issue, not a capability problem. It's clear to everyone that LLM products are immature, but saying they are incapable is misleading
Re: LLMs are still surprisingly bad at some simple tasks
#38https://chatgpt.com/share/68cffaab-4c14-8006-89a2-1818172e4d... Tried on ChatGPT, seems fine.
Re: LLMs are still surprisingly bad at some simple tasks
#39> “To stave off some obvious comments: > yoUr'E PRoMPTiNg IT WRoNg! > Am I though?” Yes. You’re complaining that Gemini “shits the bed”, despite using 2.5 Flash (not Pro), without search or reasoning. It’s a fact that some models are smarter than others. This is a task that requires reasoning so the article is hard to take seriously when the author uses a model optimised for speed (not intelligence), and doesn’t even…
OP here. I literally opened up Gemini and used the defaults. If the defaults are shit, maybe don't offer them as the default? Or, if LLMs are so smart, why doesn't it say "Hmmm, would you like to use a different model for this?" Either way, disappointing.
Re: LLMs are still surprisingly bad at some simple tasks
#40They are very good at some tasks and terrible at others. I use LLMs for language-related work (translations, grammatical explanations etc) and they are top notch in that as long as you do not ask for references to particular grammar rules. In that case they will invent non-existent references. They are also good for tutor personas: give me jj/git/emacs commands for this situation. But they are bad in other cases. I s…
I think Gemini is one of the best example of an LLM that is in some cases the best and in some cases truly the worst. I once asked it to read a postcard written by my late grandfather in Polish, as I was struggling to decipher it. It incorrectly identified the text as Romanian and kept insisting on that, even after I corrected it: "I understand you are insistent that the language is Polish. However, I have carefully…