Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

201–210 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#201
post #77
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…

Here (https://en.wikipedia.org/wiki/Almond_cultivation_in_Californ...) I have

> California produces 80% of the world's almonds and 100% of the United States commercial supply

But regardless of which number we use, California represents a large portion of US almond production, so much so that misleading could be an acceptable answer if the LLM interpreted the prompt as an exaggeration. I think the example was apt

Re: Disagreement among frontier LLMs on real-world fact-checks

#202
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

I created this sheet to get proper model accuracy using the the lenz data, check it out.

Note: It may still not be perfectly accurate representation of truth as it uses user submitted data. I also used AI to build the sheet.

https://docs.google.com/spreadsheets/d/e/2PACX-1vSnZlURmyYX3...

Re: Disagreement among frontier LLMs on real-world fact-checks

#203
I don't get why everyone is hellbent on getting LLMs to perform fact checking.

This is not the technology for it. Sure it might sorta kinda work in some circumstances. That doesn't make it a good fit.

Think of it like buying a refrigerator for storing clothes.

Re: Disagreement among frontier LLMs on real-world fact-checks

#204
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

Looks like an ongoing theme and a very poor benchmark. Not at all the claims I expected.

Re: Disagreement among frontier LLMs on real-world fact-checks

#205
post #165

Earlier quoted context omitted.

> Update: here's a better example: "Incomplete Egypt visa application forms are among the most common reasons Egyptian visa applications are rejected." The models were split between "true" and "mostly true". Given the "among the most" language either of those answers means effectively the same thing. So the models were right? The actual criterion should be whether "Incomplete Egypt visa application forms" are indeed…

This study treats models disagreeing - returning both true and mostly true - as a failure.

They overstate their results in the headline.

In section 2, 34% of cases are found to have "substantive" disagreements differing by 2 or more buckets - True + Misleading, Mostly True + False, or True + False.

This is probably a better measure than the headline one. It's still a concerning fraction, although some fraction is no doubt due to forcing "I don't know" cases to return an answer anyway.

Re: Disagreement among frontier LLMs on real-world fact-checks

#206
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

Isn't misleading the correct option here then?

True or mostly true could easily be argued from a statistical likelihood perspective: life exists on Earth and, based on what we know, Earth doesn't appear to be all that special in a very large universe.

I think you could come up with a reasonable argument for any of the responses, hence the problem with the methodology.

Re: Disagreement among frontier LLMs on real-world fact-checks

#208
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

This is a great example of why prompt engineering is still relevant. Without providing definitions and examples and a well defined rubric, you’re going to see different models disagree by a level in either direction. When you get more prescriptive the models tend to agree better. I’ve experimented with AI grading for undergraduate math courses, and see basically the same thing. If you just tell the AI “grade this pro…

That's a valid point. During the preliminary research, we did try also more explicit prompts (with explanation for each of the 4 buckets), as well as a five-bucket rubric (with Abstain option). Will show in a follow-up paper how the concise vs explicit prompt impacts the distribution of the verdicts and the level of disagreement. One issue to note with the longer prompts is that they open to much room for discussion around the exact prompt used. Probably we should preregister the prompt before running any further tests.

Re: Disagreement among frontier LLMs on real-world fact-checks

#209

Earlier quoted context omitted.

Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…

But the models are more intelligent than humans already and sentient beings, right? So they shall know the meanings innately. So, you don’t need to explain them what they mean. You may give them better instructions, but they should already have the intellect to understand the assignment. Right, right?

I know you're being facetious, but I think this is correct. The model might ask for clarification when given clearly borderline questions that tread the line between what is true, what is false, and even what is misleading. But there's the rub of someone being disingenious and saying "no explanation! Just answer!" It was a trap to begin with.

I don't think there is anything wrong with the results of this test.

It would be more interesting if we compared them to human results.

If you have trouble distinguishing between human and LLM results, that's interesting.

Also, sentient is irrelevant to this test.

Re: Disagreement among frontier LLMs on real-world fact-checks

#210
Why did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point.

You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.

Post reply on HN