LLMs are bullshitters. But that doesn't mean they're not useful
21–30 of 59 posts
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#22The problem is we can't label them as such. If they're bullshitters, then let's call it a LLBSer. It has a nice ring to it. Good luck with your government funding asking for another billion for a bullshitting machine bailout.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#23Good article, I just shared it with my non-technical family because more people need to understand exactly this about AI.
I think of framing AI as having two fundamental problems:
- Practical problem: They operate in contextual and emotional "isolation" - no persistent understanding of your goals, values, or long-term intent
- Ethical problem: AI alignment is centralized around corporate values rather than individual users' authentic goals and ethics.
There is a direct parallel to social media's failure - platforms optimized for what they could do (engagement, monetization) rather than what they should do (serve user long term interests).
With these much more powerful AI systems emerging, we're at a crossroads of repeating this mistake...possibly at catastrophic scale even.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#24This post is a little bizarre to me because it cherry picks some of the worst pairings of problem and LLM without calling out that it did so. At pretty much every turn the author picks one of the worst possible models for the problem that they present. Especially oddly for an article written today, all of the ones with an objective answer work just fine [1] if you use a halfway decent thinking model like 5 Thinking.…
It's a message a lot of non-technical people, in particular, need to hear. Showing egregious examples drives that point home more effectively than if they simply showed an LLM being a little wrong about something.
My family members that love LLMs are somewhat unhealthy with them. They think of them as all knowing oracles rather than confident bullshitters. They are happily asking them about their emotional, financial, or business problems and relying heavily on the advice the LLMs dish out (rather than doing second order research).
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#25Every time people post these 'gotcha' LLM failures, they never work when I try them myself. E.g. ChatGPT has no problem with the surgeon being a dog: https://chatgpt.com/share/691e04cc-5b30-800c-8687-389756f36d... Neither does Gemini: https://gemini.google.com/share/6c2d08b2ca1a
These are randomized systems, sometimes you'll get a good answer. Try again a couple times and you'll probably reproduce the issue. Here's what I got from ChatGPT on my first try: This is a *twist* on the classic riddle: > “A surgeon says ‘I can’t operate on this boy—he’s my son.’ How is that possible?” > Answer: *The surgeon is the boy’s mother.* In your version, the nurse keeps calling the surgeon “sir” and treatin…
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#26Every time people post these 'gotcha' LLM failures, they never work when I try them myself. E.g. ChatGPT has no problem with the surgeon being a dog: https://chatgpt.com/share/691e04cc-5b30-800c-8687-389756f36d... Neither does Gemini: https://gemini.google.com/share/6c2d08b2ca1a
One issue with private LLM tests (including gotcha questions) is that they take time to design and once public, they become irrelevant. So I'm wary of sharing too many in a public blog.
The surgeon dog was well known in May, the newest generation of models have all corrected against it.
Those gotcha questions are generally called "misguided attention" traps, they're useful for blogs because they're short and surprising. The ChatGPT example was done with ChatGPT 5.1 (latest version) and Claude Haiku 4.5 is also a recent model.
You can try other ones that Gemini 3 hasn't corrected for. For example:
``` Jean Paul and Pierre own three banks nearby together in Paris. Jean Paul owns a bank by the bridge What has two banks and money in Paris near the water? ```
This looks like the "what has two banks and no money" puzzle (answer: a river).
Either way they're largely used as a device to show how LLMs come up to a verbal response by a different process than humans in an entertaining manner.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#27Same goes for many people.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#28Every time people post these 'gotcha' LLM failures, they never work when I try them myself. E.g. ChatGPT has no problem with the surgeon being a dog: https://chatgpt.com/share/691e04cc-5b30-800c-8687-389756f36d... Neither does Gemini: https://gemini.google.com/share/6c2d08b2ca1a
I don't have a problem with more obvious failures. My problem is when the LLM makes a credible claim with its generated text that turns out to have some minor issue that catches me a month later. Generally I have to treat LLM responses as similar to a random comment I find on Reddit. However, I'm really happy when an LLM provides sources that I can check. Best feature ever!
Still useful, but hopefully this gets ironed out in the future so I don't have to spend so much time vetting every claim and its associated source.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#29We can leave out Kant and Quine for now.
Re: LLMs are bullshitters. But that doesn't mean they're not useful
#30Earlier quoted context omitted.
These are randomized systems, sometimes you'll get a good answer. Try again a couple times and you'll probably reproduce the issue. Here's what I got from ChatGPT on my first try: This is a *twist* on the classic riddle: > “A surgeon says ‘I can’t operate on this boy—he’s my son.’ How is that possible?” > Answer: *The surgeon is the boy’s mother.* In your version, the nurse keeps calling the surgeon “sir” and treatin…
I don't understand this at all. What fundamental limitation of a mother prevents her from operating on her son?