Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

261–270 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#261

looking at the claims i would say 5 humans would disagree even more than the llms some of the claims where llms disagree: "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia." "The slogan "Simon Go Back" was chanted in opposition to the Simon Commission in British India (1928–1930)." "Neptune Deep will start delivering natural gas in 2027." "A hotel villa in Kyrgyzstan displayed a sign stating 'n…

These "Facts" are interesting. "Neptune Deep will start delivering natural gas in 2027." for example is not a fact, its a prediction. "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia." is less of a fact and more of a litmus test for which sources of information you trust.

So, rephrase it thus:

"Russia, Ukraine, and multiple international news agencies reported that Ukrainian drones targeted Moscow on or around May 18, 2026."

There are rarely pure first-order "facts" in the mathematical sense. There are evidence-backed claims with confidence levels. That does not make it "just a litmus test". It makes it a probabilistic factual claim with varying confidence levels - and this one happens to be verified and unambiguous.

Re: Disagreement among frontier LLMs on real-world fact-checks

#262

Why did they exclude Grok? Given the published philosophical differences in how Grok is trained, it would provide an interesting data point. You can argue all day about those differences, but missing this opportunity to observe them in an objective way is disappointing.

Title says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

>Grok is trained to have a bias

Oh and the others arent? You cant really be that niave right?

Re: Disagreement among frontier LLMs on real-world fact-checks

#264
post #163

Earlier quoted context omitted.

>Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? Disagree. The definition of misleading is a true fact that is presented in a way to lead you to a false conclusion. Example: "Most good engineers are male". It is true as a consequence of most engineers being male in general, but it leads the reader to a potential false implication tha…

> but it leads the reader to a potential false implication that an average man is better than an average woman. I think that's _you_ turning the statement into something much broader than intended. The claim is about engineers and you're jumping from "men are better than women in engineering" to "men are better overall." To give a related example, "Most good NBA players are black." I don't think anyone would bother t…

>I think that's _you_ turning the statement into something much broader than intended.

My point is that it is possible for a reader to turn it that way, for a variety of reasons (lack of understanding of statistics, preexisting biases, or whatever). And that getting a reader to mistakenly generalize is the purpose of a misleading statement.

To mislead is to direct into a falsehood by implication even though the literally expressed facts are all true; the writer's bad intentions are necessary to qualify something as misleading I'd say, for the same reason that not all false statements are lies because to be a lie the speaker must know the statement is false and still use it. There are probably much better examples than the one I came up with on the fly, though.

Re: Disagreement among frontier LLMs on real-world fact-checks

#265
The difference between "mostly true", "misleading", and "false" is context, and responses are specifically not allowed to include any context. Even "true" has a little context, since few things can be said to be absolutely true. "Unknown" also isn't allowed.

What's 2 + 2? The answer must be one of the colors of the rainbow.

(People can draw their own conclusions, but the only coherent reason I can think of for the design of this experiment is to generate a misleading conclusion.)

Re: Disagreement among frontier LLMs on real-world fact-checks

#266

Earlier quoted context omitted.

Title says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

How do you know it is trained to have a bias? In fact can I ask you to provide a single reproducable answer right now?

Assuming this isn’t a satire reply: https://www.pnas.org/doi/10.1073/pnas.2603294123

Hope this helps!

Re: Disagreement among frontier LLMs on real-world fact-checks

#267
post #36

Earlier quoted context omitted.

So it's not a secret, why you don't add this upfront to the report? The report itself is even about LLMs, makes a lot of sense to disclose your usage of them for writing the report, especially when you're presenting evidence that boils down to LLMs being infallible.

It's an omission on my side. Will add in the next version.

Next time, we’d prefer to read your actual voice over LLM.

Re: Disagreement among frontier LLMs on real-world fact-checks

#268

Earlier quoted context omitted.

Title says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

How do you know it is trained to have a bias? In fact can I ask you to provide a single reproducable answer right now?

Grok? The model that happily generates CSAM for you while the company breathlessly defends its ability to do so? The model that referred to itself as Mecha Hitler while praising the Nazi party and calling for a second holocaust? The model that famously inserted nonsense about the “white genocide” conspiracy theory into completely unrelated queries?

Why on earth would anyone think such a model is biased?

Re: Disagreement among frontier LLMs on real-world fact-checks

#269

Earlier quoted context omitted.

Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…

> Something can be simultaneously "misleading" and either true or false. Sure they can. It might be a true fact that "100% of the murders committed in over the last 25 years were committed by !" but actually it's a town of 750 people and there was only one murder during that time frame.

how is that misleading if it's a fact, it's only misleading if you presume to know the reaction or intent behind making such a claim, and without context we should be extremely careful in making such presumptions.

Re: Disagreement among frontier LLMs on real-world fact-checks

#270

Earlier quoted context omitted.

>Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? Disagree. The definition of misleading is a true fact that is presented in a way to lead you to a false conclusion. Example: "Most good engineers are male". It is true as a consequence of most engineers being male in general, but it leads the reader to a potential false implication tha…

Isn't this still assuming we can even determine what is true or false? Newtonian physics is false, but it works well enough we teach it in college. But our best models of physics are currently in disagreement, so can we even say they are true? Given the replication crisis, especially in social sciences, how many of peer reviewed findings can be called true? Even experimental results can be false (consider studies tha…

Newtonian physics doesn't just work well enough for education. It provides an incredibly accurate and precise model of the world except at extremes. The majority of engineering does not necessitate using theories of relativity. Both theories are incomplete models approximating reality and are very far from being false.
Post reply on HN