Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

211–220 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#211
post #165

Earlier quoted context omitted.

> Update: here's a better example: "Incomplete Egypt visa application forms are among the most common reasons Egyptian visa applications are rejected." The models were split between "true" and "mostly true". Given the "among the most" language either of those answers means effectively the same thing. So the models were right? The actual criterion should be whether "Incomplete Egypt visa application forms" are indeed…

This study treats models disagreeing - returning both true and mostly true - as a failure.

Agree with @pjdesno, that the 34% substantive or polar disagreement might be a better headline number. Or even the 21% polar disagreement (at least one model True, and at least one model False), which is still high for many real-world applications.

Re: Disagreement among frontier LLMs on real-world fact-checks

#212
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

Misleading is not analogous with True or False.

Depending on the question, True or False can be objectively right/wrong. Misleading is going to be a judgement call.

This is the inherent problem with "fact checking." It's hard to be completely objective. Even when the question has an objective answer, simply choosing where to look and what facts to verify is itself a bias. Looking at this instead of that, or looking at this but not also this other thing that adds context, etc.

Frankly i think disagreeing often is the expected outcome. Fact checking is jsut kinda bullshit. It's spin dressed up as objectivity. I hope people remember that "fact checking" is a relatively modern thing.

Re: Disagreement among frontier LLMs on real-world fact-checks

#213

Dissent and consensus among frontier models is a good thing. Just like on a team of high performers, there are a million ways to skin a grape. In my research, I've found that models perform better when they operate as a collective system with reputation, incentives, and accountability instead of isolated oracles answering alone. Agreement, dissent, and correctness should all carry rewards and consequences. Just like…

Funny timing. I've been working on a prediction market orchestration that runs Claude and a few others over Polymarket/Kalshi. The models are NOT unanimous. At all, really. I spent about a month convinced that I could just run all five and take majority vote. Eventually I pivoted to a chaining approach where I benchmark areas each model excels, and settled on more like a graph-like architecture where outputs get split and verified by another, then reconstructed, and re-verified at each stage. Has actually been working out pretty well so far, 2 months in consistent profit, but I'm not a millionaire yet.

Re: Disagreement among frontier LLMs on real-world fact-checks

#214
post #20

Earlier quoted context omitted.

Data collection and processing was done manually. LLMs helped with the report drafting. Everything was human reviewed before publishing.

So it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.

I think you mean fallible.

It's also a bit weird to "disclose use of LLMs". It rubs me wrong, the same way parents breathlessly talking about "screen time" rubbed me wrong: it's too general, and with such a broad brush, it's going to sweep up a bunch of perfectly fine usage with a bunch of dubious usage. On the flip side, if folks do start disclosing all the time, it's going to turn into a Prop 65 warnings in CA, where everything says it has lead in it, so folks pretty much ignore it and move on.

If the report's conclusions and reasoning lean on LLMs, or if the data processing itself was done with LLMs, that would be interesting, and I wouldn't treat it as some sort of disclosure, but rather discuss it under methodology. Using LLMs to polish the language a bit after writing an initial draft with key findings? Much less interesting.

I realize this is now a religious issue, and some folks are allergic to anything that touched an LLM. I just don't think that perspective is going to end up having a good shelf life.

Re: Disagreement among frontier LLMs on real-world fact-checks

#215
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

I feel like the prompting could be tweaked to improve response.

Models often have a reasoning/thinking/research mode that is triggered by asking slightly differently.

Still though, Gemini can be a little weak on this front default but can be aligned to behave better.

Re: Disagreement among frontier LLMs on real-world fact-checks

#216

And how many claims human experts disagree on in the exact same setting? I'm not being snarky here. Without something to compare to the 67% number tells us nothing. And it's known that many humans disagree with human fact checkers too (see: any election around the world.)

Agree. Human experts also struggle agreeing on this type of claims. The inter-annotator agreement on the verdicts on the AVeriTeC corpus across 50 organizations is κ=0.619 - substantial but well short of perfect.

Re: Disagreement among frontier LLMs on real-world fact-checks

#217

Why did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point. You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.

Title says “Frontier” which would exclude Grok.

Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

Re: Disagreement among frontier LLMs on real-world fact-checks

#218
post #99
post #91

Earlier quoted context omitted.

Better options would have been "True", "False", "Unknown" (which opinions would fall under too). That also includes an interesting assessment of how well LLMs can identify missing information. My guess is they would be a very low number of "unknown" and a much higher level of agreement (assuming equal representation). Unless the RLHF techniques have gotten better at getting an LLM to say "I don't know", which I doubt…

Tried initially with a fifth bucket, Abstain. It was actually heavily used by some of the models. But it felt as if they are using this to "avoid" some of the hard questions, and we dropped this bucket to force them to provide a verdict.

[dead]

Re: Disagreement among frontier LLMs on real-world fact-checks

#219
post #77

Earlier quoted context omitted.

> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…

Here ( https://en.wikipedia.org/wiki/Almond_cultivation_in_Californ... ) I have > California produces 80% of the world's almonds and 100% of the United States commercial supply But regardless of which number we use, California represents a large portion of US almond production, so much so that misleading could be an acceptable answer if the LLM interpreted the prompt as an exaggeration. I think the example was apt

"All almonds are grown in the U.S. state of California." implies "No almonds are grown outside the U.S. state of California."

You find one almond tree outside of California that grows almonds, where such almonds are grown intentionally, and the claim is false.

Re: Disagreement among frontier LLMs on real-world fact-checks

#220
> No Abstain option is offered (a forced choice keeps the comparison symmetric across models).

Well that's your problem right there: They removed any confidence indicator and forced a choice.

For example:

Statement: Individuals who prefer music with less positive emotional content tend to have higher intelligence.

Gemini: That statement is supported by recent psychological research, though with some important scientific caveats regarding how strong that link actually is.

How should the agent classify this? True? Mostly true? Misleading? False?

Post reply on HN