Earlier quoted context omitted.
Many of the rows in that spreadsheet reference "current events", which models aren't expected to do much better at than a human making an educated guess! They all have cutoff dates either last year or early this year and know nothing about what happened in "April 2026". This is doubly problematic because you evaluated earlier models like Gemini Pro 3 instead of 3.1, GPT 5.4 instead of 5.5, etc... Given that it's only…
Two of the models used have retrieval capabilities and have access to newer information through search. The other three are parametric.
Disagreement among frontier LLMs on real-world fact-checks
61–70 of 377 posts
Re: Disagreement among frontier LLMs on real-world fact-checks
#62"None of these claims is older than February 15, 2026" All of the models they tested were trained on data from before February 15th ... being asked specific questions about things that happened after they were trained.
Re: Disagreement among frontier LLMs on real-world fact-checks
#63(Brought to you by) Lenz...? a crummy commercial...? ...son of a bitch
Re: Disagreement among frontier LLMs on real-world fact-checks
#64That's better than all agreeing on the wrong answer, however.
Re: Disagreement among frontier LLMs on real-world fact-checks
#65Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
Re: Disagreement among frontier LLMs on real-world fact-checks
#66Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
I expect the models are inferring quite a bit from the short prompt, and with structured outputs it would be quite easy to have them give the one word response in one field and explain why in another
Re: Disagreement among frontier LLMs on real-world fact-checks
#67Re: Disagreement among frontier LLMs on real-world fact-checks
#68Take just one random example: `Hostels in Kota, Rajasthan commonly use caged ceiling fans as a preventive measure against student suicides`
While `Hostels in Kota, Rajasthan commonly use caged ceiling fans` may be a verifiable facts (though I doubt if there are any statistics for verification but let's say there are), `a preventive measure against student suicides` is a claim that no one can prove that. It can just a believe at most.
Arh. Did Biden stole Thump 2nd term? Truth or fact or claim?
Re: Disagreement among frontier LLMs on real-world fact-checks
#69Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…