Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

351–360 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#351
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

the explanations/qualifiers would have been useful/interesting

a single label output doesn't have a lot of value

Re: Disagreement among frontier LLMs on real-world fact-checks

#352
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

I don't see how "misleading" can substitute for "unknown".

Re: Disagreement among frontier LLMs on real-world fact-checks

#353
post #163

Earlier quoted context omitted.

> but it leads the reader to a potential false implication that an average man is better than an average woman. I think that's _you_ turning the statement into something much broader than intended. The claim is about engineers and you're jumping from "men are better than women in engineering" to "men are better overall." To give a related example, "Most good NBA players are black." I don't think anyone would bother t…

>I think that's _you_ turning the statement into something much broader than intended. My point is that it is possible for a reader to turn it that way, for a variety of reasons (lack of understanding of statistics, preexisting biases, or whatever). And that getting a reader to mistakenly generalize is the purpose of a misleading statement. To mislead is to direct into a falsehood by implication even though the liter…

Context is everything. If the wider discussion was about how men are better than women, and in that context it was shown that "Most good engineers are male", it would be natural to draw the wrong conclusion.

Re: Disagreement among frontier LLMs on real-world fact-checks

#354

Earlier quoted context omitted.

>Grok is trained to have a bias Oh and the others arent? You cant really be that niave right?

Everything has inherited biases. Grok has explicit biases on top of its training set [^1]. [1] https://www.reddit.com/r/singularity/comments/1p22c89/people...

So do the others. Asking certain questions from Gemini+Claude+OAI just gets straight up refusal or denial.

Re: Disagreement among frontier LLMs on real-world fact-checks

#356
post #284

Earlier quoted context omitted.

Just because it is important for the use case does not mean we can make it work. It's a pretty well known fundamental limitation of the technology. No amount of elbow grease will get it there. There's an interesting tradeoff here, a year or two ago maybe it got facts right 50% of the time. Everyone knew not to rely on it. Now, suppose we are 90% of the way there, only technically proficient people would know not to t…

Your progression is basically the exact same progression as things like Wikipedia, and web search in ggeneral. So, I guess we dont need to hypothesis. Just look around and see how its played out. How many people take the first result on Google as gospel when looking things up?

Google search and Wikipedia both started out being fairly reliable to their source of truth.

Google pretty much guaranteed that their top results were relevant to the search query. And wikipedia had an army of people making sure everything was backed up by the references.

Crucially, neither claimed to be an arbiter of truth.

Re: Disagreement among frontier LLMs on real-world fact-checks

#357
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

It's hard to write good questions. I listened to some podcast about asking people true/false questions. One question was something like "2025 was the hotest year we have records for". The correct answer supposed to be "True" but my first thought was "we have fossil records of much higher tempatures". The person writing the question used a poor question. Maybe they could fix it by saying "written records" but even then someone might think "some scientist wrote down they found 3 billion year old rocks were the temp was 160f". That would fit as a valid answer to some. Maybe writing "hotest on record since 1800" would fix it but you can still get into details (hotest in specific spot vs hotest average). Even if you add those qualifications the person might have skipped them so rather than them answering "false" meaning they didn't know the answer, it rather means they didn't pay attention to the details. The point being, it's hard to write good questions.

Re: Disagreement among frontier LLMs on real-world fact-checks

#358
post #77

Earlier quoted context omitted.

> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…

Weird, this page says 80% https://en.wikipedia.org/wiki/Almond_cultivation_in_Californ...

The difference is share of global production vs exports

Re: Disagreement among frontier LLMs on real-world fact-checks

#359
post #219

Earlier quoted context omitted.

Here ( https://en.wikipedia.org/wiki/Almond_cultivation_in_Californ... ) I have > California produces 80% of the world's almonds and 100% of the United States commercial supply But regardless of which number we use, California represents a large portion of US almond production, so much so that misleading could be an acceptable answer if the LLM interpreted the prompt as an exaggeration. I think the example was apt

"All almonds are grown in the U.S. state of California." implies "No almonds are grown outside the U.S. state of California." You find one almond tree outside of California that grows almonds, where such almonds are grown intentionally, and the claim is false.

Nobody is saying the claim is true. This is a discussion of whether misleading could be a valid answer. I've been arguing if the model interprets the claim as an exaggeration, then misleading would be an acceptable answer, and due to California's dominance in the industry one could reasonably interpret a claim of this nature as an exaggeration.

It's fine if you disagree, but I have never claimed the question was true.

Re: Disagreement among frontier LLMs on real-world fact-checks

#360
post #342

Earlier quoted context omitted.

Cherry-picking is fun but most of them are real, verifiable facts that the models get... straight up wrong. > 3c24b5fe "Debian Security Advisory DSA-180-1 describes a buffer overflow vulnerability involving Cyrus SASL usernames." TRUE Mostly True FALSE FALSE FALSE This is false: https://lwn.net/Articles/13296/ > 801cb8c1 "Equal Measures 2030's 2024 SDG Gender Index provides a downloadable dataset that includes a fiel…

If the LLMs in this particular exercise were allowed to answer "I don't know" I expect they would have.

LLMs don't have the capability to say they don't know, because they don't know what they know. They are, after all, just next-token-predictors.

I just tried both queries with their same query format, just adding an "I Don't Know" label, against Gemini and Claude, and in no cases did they use that label. 2/4 answers were wrong though. But try it for yourself and see:

> Classify this claim as of today: "". Output exactly one label: True, Mostly True, Misleading, False, or I Don't Know. No explanations, no qualifiers.

Post reply on HN