[flagged]
Disagreement among frontier LLMs on real-world fact-checks
361–370 of 377 posts
Re: Disagreement among frontier LLMs on real-world fact-checks
#362Earlier quoted context omitted.
"False" isn't correct in strict boolean terms either, since that implies that the inverse is true. Claiming "there is extraterrestrial life in the universe" is false is logically equivalent to claiming that "no extraterrestrial life exists anywhere in the universe" is true. Both statements would have to be interpreted as "false" under your criteria, as neither has any evidence to substantiate it. That leads us to a l…
If we strictly follow logic, then nobody and nothing can claim that anything is true or false. We just stick these labels to things which seems to have high enough probability. The problem is that “high enough” is very-very-very different for different people, topics, and even time.
Sure they can. They may or may not be correct, but that's a matter of empirical validation entirely distinct from the logic flow itself. Whether or not a conclusion is logically implied by its hypotheses has nothing to do with whether the input hypotheses are themselves true. Logic is just the reasoning process.
> We just stick these labels to things which seems to have high enough probability. The problem is that “high enough” is very-very-very different for different people, topics, and even time.
That's true, but outside the scope of logic, and is entirely a matter of semantics.
Re: Disagreement among frontier LLMs on real-world fact-checks
#363Earlier quoted context omitted.
Two of the models used have retrieval capabilities and have access to newer information through search. The other three are parametric.
The title mention "fact-checks", but "fact checking" is a process in which facts are checked against sources, not one where you are given a random fact and have to tell if it's true or false from your own memory. That's what is normally called a quiz game. So a more honest title for this research would be "Models answer differently to quiz questions".
from their future
It's as if you asked a human in 2026 to "fact check" something from 2027.
You're going to get an educated guess at best, a coin flip at worst.
Re: Disagreement among frontier LLMs on real-world fact-checks
#364Earlier quoted context omitted.
Your progression is basically the exact same progression as things like Wikipedia, and web search in ggeneral. So, I guess we dont need to hypothesis. Just look around and see how its played out. How many people take the first result on Google as gospel when looking things up?
Google search and Wikipedia both started out being fairly reliable to their source of truth. Google pretty much guaranteed that their top results were relevant to the search query. And wikipedia had an army of people making sure everything was backed up by the references. Crucially, neither claimed to be an arbiter of truth.
Re: Disagreement among frontier LLMs on real-world fact-checks
#365Earlier quoted context omitted.
Google search and Wikipedia both started out being fairly reliable to their source of truth. Google pretty much guaranteed that their top results were relevant to the search query. And wikipedia had an army of people making sure everything was backed up by the references. Crucially, neither claimed to be an arbiter of truth.
I don't recall any LLM provider claiming to be a arbiter of truth.
Re: Disagreement among frontier LLMs on real-world fact-checks
#366mostly true and misleading counting as separate buckets while also being somewhat orthogonal conceptually is also stupid.
a better output format might be true|false|unknowable, confidence where confidence is 0..1
at least then you can compare agreement among models as a distance measurement and not a moronic bucket agreement
the conclusion is actually obvious: llms are good enough for most of this work and it is definitely cheaper, so you should use llms for fact checking at least as a first pass
Re: Disagreement among frontier LLMs on real-world fact-checks
#367Earlier quoted context omitted.
If we strictly follow logic, then nobody and nothing can claim that anything is true or false. We just stick these labels to things which seems to have high enough probability. The problem is that “high enough” is very-very-very different for different people, topics, and even time.
> If we strictly follow logic, then nobody and nothing can claim that anything is true or false. Sure they can. They may or may not be correct, but that's a matter of empirical validation entirely distinct from the logic flow itself. Whether or not a conclusion is logically implied by its hypotheses has nothing to do with whether the input hypotheses are themselves true. Logic is just the reasoning process. > We just…
I replied like that because I think you applied logic too strictly already.
Re: Disagreement among frontier LLMs on real-world fact-checks
#368Re: Disagreement among frontier LLMs on real-world fact-checks
#369Re: Disagreement among frontier LLMs on real-world fact-checks
#370Earlier quoted context omitted.
Assuming this isn’t a satire reply: https://www.pnas.org/doi/10.1073/pnas.2603294123 Hope this helps!
The article compares in particular Grokipedia to Wikipedia, and it states: > Similarity measures across the two platforms reveal a bimodal structure: many Grokipedia articles closely resemble their Wikipedia counterparts, while a considerable subset diverges. Political bias differences emerge primarily within the divergent subset, where Grokipedia shows a relative rightward shift in the ideological orientation of fre…