LLMs are still surprisingly bad at some simple tasks
1–10 of 107 posts
Re: LLMs are still surprisingly bad at some simple tasks
#2> yoUr'E PRoMPTiNg IT WRoNg!
> Am I though?”
Yes. You’re complaining that Gemini “shits the bed”, despite using 2.5 Flash (not Pro), without search or reasoning.
It’s a fact that some models are smarter than others. This is a task that requires reasoning so the article is hard to take seriously when the author uses a model optimised for speed (not intelligence), and doesn’t even turn reasoning on (nor suggest they’re even aware of it being a feature).
I asked the exact prompt to ChatGPT 5 Thinking and got an excellent answer with cited sources, all of which appears to be accurate.
Re: LLMs are still surprisingly bad at some simple tasks
#3Re: LLMs are still surprisingly bad at some simple tasks
#4> “To stave off some obvious comments: > yoUr'E PRoMPTiNg IT WRoNg! > Am I though?” Yes. You’re complaining that Gemini “shits the bed”, despite using 2.5 Flash (not Pro), without search or reasoning. It’s a fact that some models are smarter than others. This is a task that requires reasoning so the article is hard to take seriously when the author uses a model optimised for speed (not intelligence), and doesn’t even…
Search and reasoning use up more context, leading to context rot, and subtler harder to detect hallucinations. Reasoning doesn’t always focus on evaluating the quality of evidence, just “problem solving” from some root set of axioms found in search.
I’ve had this happen in Claude code for example where it hallucinated a few details about a library based on what badly written forum post.
Re: LLMs are still surprisingly bad at some simple tasks
#5So are people?
Re: LLMs are still surprisingly bad at some simple tasks
#6Re: LLMs are still surprisingly bad at some simple tasks
#7So are people?
Re: LLMs are still surprisingly bad at some simple tasks
#8Re: LLMs are still surprisingly bad at some simple tasks
#9So are people?
The difference might be people are actually held accountable for the results.
We are killing thousands on the road to be sure we can blame a driver instead of a computer as one example.