Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

291–300 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#292

Why did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point. You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.

Agree. Would be fun to see how much worse Grok would be at this.

Re: Disagreement among frontier LLMs on real-world fact-checks

#293

One fun example: "Ruskin Bond was born on May 19, 1934, in Kasauli, Himachal Pradesh, India". Opus and Gemini believe this to be true, GPT 5.4 believes it's false, Sonar thinks it's mostly true. Disagreement value of 3, you can't disagree more than some models thinking it's true, some thinking it's false But my impression from 2 minutes on Wikipedia is that the most likely disagreement is on the "Himachal Pradesh, In…

How can someone be born in a state that does not yet exist? The statement has the year in it clearly demonstrating the contradiction. One can't be born in the Soviet Union in 1995 or in Tsarist Russia in 1950.

Re: Disagreement among frontier LLMs on real-world fact-checks

#294

Earlier quoted context omitted.

> But it felt as if they are using this to "avoid" some of the hard questions, and we dropped this bucket to force them to provide a verdict. do you not see how that creates extremely misleading and valueless results? you are coercing the results into what you want to see.

Exactly what people do when they use LLMs for "fact-checking" online, and any verbose explanation would be mostly ignored anyway, when people ask political, ethical, or simply ambiguous questions that they hold any stakes in. Don't even need politics for it, there is no point in probing a mathematical black box for "how many soldiers died in the year X in war Y". Any original source is preferable to a blurry "summary…

ask the black box to search for the original source and verify it yourself?

Re: Disagreement among frontier LLMs on real-world fact-checks

#295

Why did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point. You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.

Title says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

Bias is orthogonal to accuracy.

Re: Disagreement among frontier LLMs on real-world fact-checks

#296
post #251

Earlier quoted context omitted.

yeah i really don't like the corpus of statements and it makes me doubt lenz. consider > “Artificial intelligence will cause widespread job loss among software engineers.” https://lenz.io/c/ai-software-engineers-job-loss-impact-05e4... this is a statement about the future. who knows? dataset also includes > Robots will not replace human teachers in schools in the near future. or > Papua New Guinea has very few female…

Agree that some of the claims are forward-looking. The messiness of the real-world and real-user fact checks. No ground-truth verdicts are provided or used in the study though. It only measures the level of agreement between the selected models, not which one is right on which claim. I.e. none of the claims is actually labelled.

were you involved in making the study? your bio says you work for them so you should probably indicate that in your comments.

lack of agreement when there is no singular correct answer (or any answer at all) isn't a useful metric

I ran into a lot of these kinds of issues when working on the Citation Needed WMF project (and related extensions). Truth is so often very nuanced.

Re: Disagreement among frontier LLMs on real-world fact-checks

#298
> the most recent real-world user submissions to a fact-checking platform

'Fact checking' platforms aren't truth. Many 'fact checking' platforms are self-admittedly focused on left advocacy (snopes), or right wing advocacy (newsbusters). lenz-llm-disagreement.csv doesn't state the data source.

Re: Disagreement among frontier LLMs on real-world fact-checks

#299
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

I would argue, FALSE is the correct answer, since this is not a fact, you can know for sure. The logical inverse is also FALSE.

A proposition and its logical inverse cannot both be false. That's a contradiction.

A proposition and its logical inverse can both be unknown, and in fact, a proposition being unknown implies that its logical inverse must also be unknown.

Re: Disagreement among frontier LLMs on real-world fact-checks

#300
post #277

Earlier quoted context omitted.

>Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? Disagree. The definition of misleading is a true fact that is presented in a way to lead you to a false conclusion. Example: "Most good engineers are male". It is true as a consequence of most engineers being male in general, but it leads the reader to a potential false implication tha…

> The definition of misleading is a true fact that is presented in a way to lead you to a false conclusion. According to Merriem-Webster, which defines "mislead" as the following: 1. (transitive verb) to lead in a wrong direction or into a mistaken action or belief often by deliberate deceit 2. (intransitive verb) to lead astray; give a wrong impression Presenting a "true fact" is optional when misleading someone.

Uh, you seem to be right. I can't check oxford to confirm because there's a paywall, apparently.

The mental model I've always been taught is:

False, well intended -> mistake

False, bad intention -> lie

True, bad intention -> misleading

Bad intention, regardless of truth -> deceitful

The problem of classifying all bad intentioned statements as misleading is that it leaves you without a way to express "true +bad intention". While for generic bad intentioned statements regardless of truth we already have a word (deceit).

Post reply on HN