Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

31–40 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#31
Not sure I'm understanding this. The models are asked to evaluate the truth of random claims out of their own head (except for Gemini with search grounding)? Isn't it exactly the same as asking people to play any quiz game and then rating them as "they disagree n% of the time"?

The output buckets are also pretty questionable- the difference between "True" and "Mostly true" is pretty fuzzy. Is this marked as a "disagreement"?

Re: Disagreement among frontier LLMs on real-world fact-checks

#33
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

>> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them".

I don’t understand your point. That claim is factually false and as such it’s easy to logically reply “false”. What’s the nuance here? I can’t see any

Re: Disagreement among frontier LLMs on real-world fact-checks

#34
post #17

This is an odd one. The paper is real, but was written by Claude? I am assuming OP is human, but also appears to be using Claude to post.

Let's be real, we all asked Claude to summarise this because it was written by Claude

Re: Disagreement among frontier LLMs on real-world fact-checks

#35
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

Give a model a crawler tool (like Grub.nuts.services) and your "problem" goes away.

Re: Disagreement among frontier LLMs on real-world fact-checks

#36
post #20

Earlier quoted context omitted.

Data collection and processing was done manually. LLMs helped with the report drafting. Everything was human reviewed before publishing.

So it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.

It's an omission on my side. Will add in the next version.

Re: Disagreement among frontier LLMs on real-world fact-checks

#37
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

I really don’t buy the almond explanation you’re giving. That requires the level of logic a kindergartener has. It’s a very simple all or nothing question.

If LLM’s are really supposed to be as consistently useful as they’re made out to be they should all spit out “false.”

Re: Disagreement among frontier LLMs on real-world fact-checks

#38
Recently, in May 2026, I asked ChatGPT 5.5 High to search for flights to a certain city that has recently had a new airport since like December 2025

It said the airport code didn't exist

I mean, I get the "knowledge cut off date" and whatnot, but for that sort of thing, you'd think they'd check live information before gaslighting the user, specially since it's a "live" task anyway.

Re: Disagreement among frontier LLMs on real-world fact-checks

#39
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…

If you can consistently construct "true but misleading" content, you may be qualified to work at a major newspaper.

Re: Disagreement among frontier LLMs on real-world fact-checks

#40
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

False vs misleading doesn't seem like a disagreement?

According to the benchmark it is. "Only one verdict bucket can be correct per claim, so any disagreement among the panel means at least one model's verdict is label-inconsistent under this 4-bucket rubric (True / Mostly True / Misleading / False)"
Post reply on HN