Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

301–310 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#301
post #169

Earlier quoted context omitted.

> "On May 18, 2026, Ukraine carried out a drone attack on Moscow, Russia" I actually don't know which way you came down on that one? I think strictly it's false but "mostly true" would be justifiable? (as in, to say it's false would be misleading if it lead the reader to assume there was no attack around that time). https://www.washingtonpost.com/world/2026/05/17/ukrainian-dr... It seems it happened Saturday 16th ove…

It's impossible to answer if you don't have a search tool, and three out of the five tested models didn't have a search tool.

It's impossible to answer unless you have a *100% complete search tool*.

No sytem can know everything. It doesn't matter how many tools you give it. It's always wrong to force binary True / False without shades of "I don't know"

Re: Disagreement among frontier LLMs on real-world fact-checks

#302

Earlier quoted context omitted.

How do you know it is trained to have a bias? In fact can I ask you to provide a single reproducable answer right now?

Assuming this isn’t a satire reply: https://www.pnas.org/doi/10.1073/pnas.2603294123 Hope this helps!

The article compares in particular Grokipedia to Wikipedia, and it states:

> Similarity measures across the two platforms reveal a bimodal structure: many Grokipedia articles closely resemble their Wikipedia counterparts, while a considerable subset diverges. Political bias differences emerge primarily within the divergent subset, where Grokipedia shows a relative rightward shift in the ideological orientation of frequently cited news media sources, particularly in articles related to religion and history.

Whether this constitutes a gain in bias depends on the base level bias of Wikipedia, as the bias of Grokipedia was measured relative to Wikipedia in this paper. One could plausibly argue that, if Wikipedia has leftward bias, then Grokipedia ended up less biased overall, or more centrally biased.

Re: Disagreement among frontier LLMs on real-world fact-checks

#304
GIGO is an acronym I learned in the 1970s. Things haven't changed much since then.

We live an an era where people have "their own truth", so why not let the AIs have theirs too?

The AI companies have editorial privilege on the content they feed their LLMs, and on the prompts that the users never see. I don't know why they feel a need to interfere when their AI produces something that's politically incorrect. Perhaps it's because they have a fundamental credibility problem with their products...

Re: Disagreement among frontier LLMs on real-world fact-checks

#305
post #103

"Extraterrestrial life exists somewhere in the universe." GPT-5.4: Misleading Opus 4.7: Misleading Gemini 3: FALSE Gemini 3 (Retrieval): FALSE Sonar Pro: FALSE It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options.

I would think ‘false’ is the only correct answer a there’s no evidence to prove the claim, so the claim is safely assumed false. Then again maybe that’s why I’m an atheist, not an agnostic?

True or False: I am wearing a blue shirt.

Re: Disagreement among frontier LLMs on real-world fact-checks

#306

Earlier quoted context omitted.

Title says “Frontier” which would exclude Grok. Grok is trained to have a bias, which a lot of people like, but it’s not meant to be accurate.

>Grok is trained to have a bias Oh and the others arent? You cant really be that niave right?

Everything has inherited biases. Grok has explicit biases on top of its training set [^1].

[1] https://www.reddit.com/r/singularity/comments/1p22c89/people...

Re: Disagreement among frontier LLMs on real-world fact-checks

#307
post #250

Earlier quoted context omitted.

But people use it for that. So what's your point?

It's a marketing failure (or success, depending on how you see it). AI is pretty useful for a great many things, but to really attract more and more investment the current technique seems to be convincing people that AI is useful for everything.

You're probably right, but since Google Search displays an AI-generated answer as the first result, most people end up using this feature more often than they originally intended. It's there now, and it will likely replace traditional search for the general public. Not entirely, but perhaps to a large extent.

Edit: corrected bad spelling with AI XD

Re: Disagreement among frontier LLMs on real-world fact-checks

#308

Earlier quoted context omitted.

>Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? Disagree. The definition of misleading is a true fact that is presented in a way to lead you to a false conclusion. Example: "Most good engineers are male". It is true as a consequence of most engineers being male in general, but it leads the reader to a potential false implication tha…

Isn't this still assuming we can even determine what is true or false? Newtonian physics is false, but it works well enough we teach it in college. But our best models of physics are currently in disagreement, so can we even say they are true? Given the replication crisis, especially in social sciences, how many of peer reviewed findings can be called true? Even experimental results can be false (consider studies tha…

True and False in general communication means based on best available evidence and expertise statement contains no obvious contradictions or falsehoods based on an optimistic parsing of meaning language and intent. Notably this leaves out misleading or missing data because those concerns are separate from truth and falsehood.

E.g. if I say the earth is round we optimistically parse round to include oblate spheroid and rate it true.

If I say that the earth is flat we rate it as false because there is no reasonable interpretation possible other than confusion or malice.

Re: Disagreement among frontier LLMs on real-world fact-checks

#309
post #203

I don't get why everyone is hellbent on getting LLMs to perform fact checking. This is not the technology for it. Sure it might sorta kinda work in some circumstances. That doesn't make it a good fit. Think of it like buying a refrigerator for storing clothes.

Nietzsche might say this is not the fantasy of truth, but of comfort. The Last Man wants a machine to say 'fact wrong' or 'fact right' so the abyss of no ultimate truth can be made small enough to sleep beside.

Re: Disagreement among frontier LLMs on real-world fact-checks

#310

Earlier quoted context omitted.

>Grok is trained to have a bias Oh and the others arent? You cant really be that niave right?

Everything has inherited biases. Grok has explicit biases on top of its training set [^1]. [1] https://www.reddit.com/r/singularity/comments/1p22c89/people...

It’s part of the system prompt. It doesn’t constitute a bias in the model itself.
Post reply on HN