Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

311–320 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#311
post #240

Earlier quoted context omitted.

Good point. Processing the substance of the answer might be too labor-consuming (1,000 claims x 5 models), but "thinking out loud" might improve the quality of the answers indeed. And we can still force/ask them to respond with a clear verdict at the end of their reasoning, as per the chosen rubric.

If you have the model use a tool you can define the schema as a free text rationale field followed by one in the set of possible answers, so everything is nicely formatted as a JSON.

Some models struggle combining JSON schema and web search capabilities.

Re: Disagreement among frontier LLMs on real-world fact-checks

#312
post #193

Earlier quoted context omitted.

> It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options. It's even weirder to suggest that the disagreement is indicative of a problem. If you asked five very knowledgeable humans on this subject to select the correct answer on a multiple-choice questionnaire, they would almost certainly vary significantly more than these 5 LLMs. Not to say that hallu…

What are you talking about, it had the option for nuanced responses, but it chose the more binary responses. It could have chosen no explanations, no qualifiers but instead it showed off LLMs incapability for nuance. These types of experiments prove to me that there is no real "reasoning" happening and "reasoning/thinking" tokens as a concept are mostly there to convince people to use models that consume more tokens…

> What are you talking about, it had the option for nuanced responses

The prompt allowed for exactly four valid outputs and explicitly disallowed explanations and qualifiers.

> Output exactly one label: True, > Mostly True, Misleading, or False. > No explanations, no qualifiers.

How is that a nuanced response?

> These types of experiments prove to me that there is no real "reasoning" happening and "reasoning/thinking"

My suggestion is that five presumably reasoning and thinking humans would also have variation in their responses to the exact same prompt.

Re: Disagreement among frontier LLMs on real-world fact-checks

#313
post #251

Earlier quoted context omitted.

Agree that some of the claims are forward-looking. The messiness of the real-world and real-user fact checks. No ground-truth verdicts are provided or used in the study though. It only measures the level of agreement between the selected models, not which one is right on which claim. I.e. none of the claims is actually labelled.

were you involved in making the study? your bio says you work for them so you should probably indicate that in your comments. lack of agreement when there is no singular correct answer (or any answer at all) isn't a useful metric I ran into a lot of these kinds of issues when working on the Citation Needed WMF project (and related extensions). Truth is so often very nuanced.

They introduced themselves as the study author here: https://news.ycombinator.com/item?id=48307887#48307899

Re: Disagreement among frontier LLMs on real-world fact-checks

#314
It's becoming increasingly clear to me that - at least right now - AI is only useful for 2 things:

1. Coding, with it being more useful the better you are at coding without AI

2. Any expert in their field asking questions about their field, who bother to fact check the output. E.g. "claude pls search these 1000 files and tell me if you find anywhere that they're discussing the settlement" and then the user checks the files/line numbers to make sure that it's correct - basically a turbocharged search that may have false negatives (content existed but I didn't find it) or false positives (content that I classified in a certain way but it was wrong). It takes an expert to tell the latter one in some cases.

Re: Disagreement among frontier LLMs on real-world fact-checks

#315
post #250

Earlier quoted context omitted.

It's a marketing failure (or success, depending on how you see it). AI is pretty useful for a great many things, but to really attract more and more investment the current technique seems to be convincing people that AI is useful for everything.

You're probably right, but since Google Search displays an AI-generated answer as the first result, most people end up using this feature more often than they originally intended. It's there now, and it will likely replace traditional search for the general public. Not entirely, but perhaps to a large extent. Edit: corrected bad spelling with AI XD

Search and fact checking are different problems though.

LLMs are pretty decent at 'search' given the inherent knowledge compression, and some amount of inaccuracy is fine.

Re: Disagreement among frontier LLMs on real-world fact-checks

#316
post #203

I don't get why everyone is hellbent on getting LLMs to perform fact checking. This is not the technology for it. Sure it might sorta kinda work in some circumstances. That doesn't make it a good fit. Think of it like buying a refrigerator for storing clothes.

Nietzsche might say this is not the fantasy of truth, but of comfort. The Last Man wants a machine to say 'fact wrong' or 'fact right' so the abyss of no ultimate truth can be made small enough to sleep beside.

Imagine the dystopian future where your freedom depends on convincing a panel of AI judges that you are innocent.

I assume you'd have access to AI lawyers too, better ones if you can pay for larger/newer models! Meanwhile the judges are N year old models because they are state funded, and they work 'fine'.

Re: Disagreement among frontier LLMs on real-world fact-checks

#317
post #313

Earlier quoted context omitted.

were you involved in making the study? your bio says you work for them so you should probably indicate that in your comments. lack of agreement when there is no singular correct answer (or any answer at all) isn't a useful metric I ran into a lot of these kinds of issues when working on the Citation Needed WMF project (and related extensions). Truth is so often very nuanced.

They introduced themselves as the study author here: https://news.ycombinator.com/item?id=48307887#48307899

ah. I missed that.

Re: Disagreement among frontier LLMs on real-world fact-checks

#318

Earlier quoted context omitted.

Exactly what people do when they use LLMs for "fact-checking" online, and any verbose explanation would be mostly ignored anyway, when people ask political, ethical, or simply ambiguous questions that they hold any stakes in. Don't even need politics for it, there is no point in probing a mathematical black box for "how many soldiers died in the year X in war Y". Any original source is preferable to a blurry "summary…

ask the black box to search for the original source and verify it yourself?

Sure, I like using LLMs in this way, and it often shows that it's very important to verify, because often a claim is "sourced" by what appears to be more of a fuzzy text or semantic match, sometimes even ignoring logical negations.

Especially in niche subjects.

For factual claims, I've fared better with Wikipedia and looking up the sources linked there.

Anyway, as AI text and media generation erodes the credibility of all online sources, these questions about source checking matter less and less: what if the source itself is a long and convincing-sounding text with poor sources?

This problem existed before already, but it boils down to a simple fact:

logic or maths alone cannot derive an authority that verifies claims about the real world other than weighting texts.

The question "what is the current population if Paris" can be answered by LLMs, but basically only by weighting sources, and assigning some credibility to them.

There's no real point in getting some weighted average of sources on this question, but so far, it doesn't hurt either.

Re: Disagreement among frontier LLMs on real-world fact-checks

#319
post #289

Earlier quoted context omitted.

I would think ‘false’ is the only correct answer a there’s no evidence to prove the claim, so the claim is safely assumed false. Then again maybe that’s why I’m an atheist, not an agnostic?

"False" isn't correct in strict boolean terms either, since that implies that the inverse is true. Claiming "there is extraterrestrial life in the universe" is false is logically equivalent to claiming that "no extraterrestrial life exists anywhere in the universe" is true. Both statements would have to be interpreted as "false" under your criteria, as neither has any evidence to substantiate it. That leads us to a l…

If we strictly follow logic, then nobody and nothing can claim that anything is true or false. We just stick these labels to things which seems to have high enough probability. The problem is that “high enough” is very-very-very different for different people, topics, and even time.

Re: Disagreement among frontier LLMs on real-world fact-checks

#320
Totally aside from disagreement between models unbiased by prior input any such experiment may fail to capture the outcomes experienced by real users whose prior text exchanges may substantially change the text recieved.

For instance see the folks who think that they have "awakened" their instance of ChatGPT.

Actual usage may diverge to a greater degree than models

Post reply on HN