"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.
Isn't misleading the correct option here then?
Disagreement among frontier LLMs on real-world fact-checks
171–180 of 377 posts
Re: Disagreement among frontier LLMs on real-world fact-checks
#172Re: Disagreement among frontier LLMs on real-world fact-checks
#173Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
> “Artificial intelligence will cause widespread job loss among software engineers.”
https://lenz.io/c/ai-software-engineers-job-loss-impact-05e4...
this is a statement about the future. who knows? dataset also includes
> Robots will not replace human teachers in schools in the near future.
or
> Papua New Guinea has very few female members of parliament.
what counts as very few?
> “Taurine supplementation supports mood and emotional health in humans.”
why is this labeled as misleading? i'm not even sure when I'm supposed to use the misleading label
> Anaximander was the first scientist in recorded history.
this is a judgement call as the term scientist didn't exist.
the claims that feel actually solidly answerable seem to have much better LLM performance
Re: Disagreement among frontier LLMs on real-world fact-checks
#174Earlier quoted context omitted.
So it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.
It's an omission on my side. Will add in the next version.
Re: Disagreement among frontier LLMs on real-world fact-checks
#175Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
Re: Disagreement among frontier LLMs on real-world fact-checks
#176Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…
The statistic is about commercial production, not number akmonds grown.
Looks safe to say that even majority of almonds are not grown in California.
Re: Disagreement among frontier LLMs on real-world fact-checks
#177Earlier quoted context omitted.
Two of the five models used (Gemini+Search and Sonar Pro) have retrieval capabilities and used search when classifying the claims. The disagreement between them is still quite significant - 42%.
Here are those disagreements: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... One example: Researchers estimate that the average person ingests about 5 grams of plastic per week, which is approximately the weight of a credit card. Gemini retrieval: Misleading Sonar pro: Mostly True
Was the research flagrantly incorrect? Yes. But that does not affect the truth of the statement.
Re: Disagreement among frontier LLMs on real-world fact-checks
#178Earlier quoted context omitted.
What's the correct answer for "During a private Saturday call, Democratic members of the United States House of Representatives from Virginia and Hakeem Jeffries discussed strategies after losing a redistricting case at the Supreme Court of Virginia, including trying to flip two or three Republican-held seats under the existing map."? You can only say True, False, Mostly True or Misleading. (And you're not allowed to…
Search was enabled for 2 of the 5 models -- Gemini and Sonar Pro. The disagreement between them is still high - different verdict on 42% of the claims. Fully agree, that some of those claims are hard to classify for a human as well -- the real-world messiness...
Other burning questions: What methodology was used to choose the question set? Why not allow explanations? How many passes were done for each LLM?
Re: Disagreement among frontier LLMs on real-world fact-checks
#179I'm not being snarky here. Without something to compare to the 67% number tells us nothing. And it's known that many humans disagree with human fact checkers too (see: any election around the world.)
Re: Disagreement among frontier LLMs on real-world fact-checks
#180Earlier quoted context omitted.
The examples seem intentionally diverse, but I haven't seen one that I would be surprised for someone to post about in the format of "ChatGPT/Gemini/Claude/Qwen/... says:" So the examples are good, I think. The rest is philosophy. The links you posted only show a frozen loading spinner for me (iOS Safari). (I looked at the csv in Numbers instead)
Weird, I'm loading them in Mobile Safari myself.
After a couple of seconds, the result does appear.
Happened to be just within my threshold for considering it broken, because the URL bar was "finished", and the spinner doesn't spin, but the last point is probably caused by my a11y settings (prefer no animations and no autoplay).