Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

171–180 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#171
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

Isn't misleading the correct option here then?

No, "misleading" is a statement that is used because it suggests something else. It's a curious category because, differently from true and false, it's not about the statement itself but rather the intention behind its usage or the way it might be understood. It's frankly more of a political judgement than a matter of facts.

Re: Disagreement among frontier LLMs on real-world fact-checks

#172
No human baseline to compare it to. Without that you are missing an important check on the task being poorly constructed. More importantly there is an implied reference thats missing. The implication is that people would have done better, or that perfect agreement is possible.

Re: Disagreement among frontier LLMs on real-world fact-checks

#173
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

yeah i really don't like the corpus of statements and it makes me doubt lenz. consider

> “Artificial intelligence will cause widespread job loss among software engineers.”

https://lenz.io/c/ai-software-engineers-job-loss-impact-05e4...

this is a statement about the future. who knows? dataset also includes

> Robots will not replace human teachers in schools in the near future.

or

> Papua New Guinea has very few female members of parliament.

what counts as very few?

> “Taurine supplementation supports mood and emotional health in humans.”

why is this labeled as misleading? i'm not even sure when I'm supposed to use the misleading label

> Anaximander was the first scientist in recorded history.

this is a judgement call as the term scientist didn't exist.

the claims that feel actually solidly answerable seem to have much better LLM performance

Re: Disagreement among frontier LLMs on real-world fact-checks

#174
post #36

Earlier quoted context omitted.

So it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.

It's an omission on my side. Will add in the next version.

[dead]

Re: Disagreement among frontier LLMs on real-world fact-checks

#175
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

This is not how people use LLMs. If you ask one of these questions you’d get a longer answer, often grounded on the internet. I speculate that conditional on a smart human operator interpreting the results, such interpretations across vendors converge more often than this report makes it seem.

Re: Disagreement among frontier LLMs on real-world fact-checks

#176
post #77
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…

The 51% is US, the question was about California.

The statistic is about commercial production, not number akmonds grown.

Looks safe to say that even majority of almonds are not grown in California.

Re: Disagreement among frontier LLMs on real-world fact-checks

#177
post #118
post #116

Earlier quoted context omitted.

Two of the five models used (Gemini+Search and Sonar Pro) have retrieval capabilities and used search when classifying the claims. The disagreement between them is still quite significant - 42%.

Here are those disagreements: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... One example: Researchers estimate that the average person ingests about 5 grams of plastic per week, which is approximately the weight of a credit card. Gemini retrieval: Misleading Sonar pro: Mostly True

Internally the statement is perfectly true: some researchers did estimate this, and the credit card is a fair proxy for a 5g mass.

Was the research flagrantly incorrect? Yes. But that does not affect the truth of the statement.

Re: Disagreement among frontier LLMs on real-world fact-checks

#178
post #123
post #93

Earlier quoted context omitted.

What's the correct answer for "During a private Saturday call, Democratic members of the United States House of Representatives from Virginia and Hakeem Jeffries discussed strategies after losing a redistricting case at the Supreme Court of Virginia, including trying to flip two or three Republican-held seats under the existing map."? You can only say True, False, Mostly True or Misleading. (And you're not allowed to…

Search was enabled for 2 of the 5 models -- Gemini and Sonar Pro. The disagreement between them is still high - different verdict on 42% of the claims. Fully agree, that some of those claims are hard to classify for a human as well -- the real-world messiness...

Why was it enabled for only 2 of the 5?

Other burning questions: What methodology was used to choose the question set? Why not allow explanations? How many passes were done for each LLM?

Re: Disagreement among frontier LLMs on real-world fact-checks

#179
And how many claims human experts disagree on in the exact same setting?

I'm not being snarky here. Without something to compare to the 67% number tells us nothing. And it's known that many humans disagree with human fact checkers too (see: any election around the world.)

Re: Disagreement among frontier LLMs on real-world fact-checks

#180
post #167

Earlier quoted context omitted.

The examples seem intentionally diverse, but I haven't seen one that I would be surprised for someone to post about in the format of "ChatGPT/Gemini/Claude/Qwen/... says:" So the examples are good, I think. The rest is philosophy. The links you posted only show a frozen loading spinner for me (iOS Safari). (I looked at the csv in Numbers instead)

Weird, I'm loading them in Mobile Safari myself.

Sorry, I didn't wait quite long enough after the last output line appeared.

After a couple of seconds, the result does appear.

Happened to be just within my threshold for considering it broken, because the URL bar was "finished", and the spinner doesn't spin, but the last point is probably caused by my a11y settings (prefer no animations and no autoplay).

Post reply on HN