Earlier quoted context omitted.
Good point. Processing the substance of the answer might be too labor-consuming (1,000 claims x 5 models), but "thinking out loud" might improve the quality of the answers indeed. And we can still force/ask them to respond with a clear verdict at the end of their reasoning, as per the chosen rubric.
If you have the model use a tool you can define the schema as a free text rationale field followed by one in the set of possible answers, so everything is nicely formatted as a JSON.
Disagreement among frontier LLMs on real-world fact-checks
311–320 of 377 posts
Re: Disagreement among frontier LLMs on real-world fact-checks
#312Earlier quoted context omitted.
> It's a weird fact claim, because the ground truth is "nobody knows for sure" and that's not one of the available options. It's even weirder to suggest that the disagreement is indicative of a problem. If you asked five very knowledgeable humans on this subject to select the correct answer on a multiple-choice questionnaire, they would almost certainly vary significantly more than these 5 LLMs. Not to say that hallu…
What are you talking about, it had the option for nuanced responses, but it chose the more binary responses. It could have chosen no explanations, no qualifiers but instead it showed off LLMs incapability for nuance. These types of experiments prove to me that there is no real "reasoning" happening and "reasoning/thinking" tokens as a concept are mostly there to convince people to use models that consume more tokens…
The prompt allowed for exactly four valid outputs and explicitly disallowed explanations and qualifiers.
> Output exactly one label: True, > Mostly True, Misleading, or False. > No explanations, no qualifiers.
How is that a nuanced response?
> These types of experiments prove to me that there is no real "reasoning" happening and "reasoning/thinking"
My suggestion is that five presumably reasoning and thinking humans would also have variation in their responses to the exact same prompt.
Re: Disagreement among frontier LLMs on real-world fact-checks
#313Earlier quoted context omitted.
Agree that some of the claims are forward-looking. The messiness of the real-world and real-user fact checks. No ground-truth verdicts are provided or used in the study though. It only measures the level of agreement between the selected models, not which one is right on which claim. I.e. none of the claims is actually labelled.
were you involved in making the study? your bio says you work for them so you should probably indicate that in your comments. lack of agreement when there is no singular correct answer (or any answer at all) isn't a useful metric I ran into a lot of these kinds of issues when working on the Citation Needed WMF project (and related extensions). Truth is so often very nuanced.
Re: Disagreement among frontier LLMs on real-world fact-checks
#3141. Coding, with it being more useful the better you are at coding without AI
2. Any expert in their field asking questions about their field, who bother to fact check the output. E.g. "claude pls search these 1000 files and tell me if you find anywhere that they're discussing the settlement" and then the user checks the files/line numbers to make sure that it's correct - basically a turbocharged search that may have false negatives (content existed but I didn't find it) or false positives (content that I classified in a certain way but it was wrong). It takes an expert to tell the latter one in some cases.
Re: Disagreement among frontier LLMs on real-world fact-checks
#315Earlier quoted context omitted.
It's a marketing failure (or success, depending on how you see it). AI is pretty useful for a great many things, but to really attract more and more investment the current technique seems to be convincing people that AI is useful for everything.
You're probably right, but since Google Search displays an AI-generated answer as the first result, most people end up using this feature more often than they originally intended. It's there now, and it will likely replace traditional search for the general public. Not entirely, but perhaps to a large extent. Edit: corrected bad spelling with AI XD
LLMs are pretty decent at 'search' given the inherent knowledge compression, and some amount of inaccuracy is fine.
Re: Disagreement among frontier LLMs on real-world fact-checks
#316I don't get why everyone is hellbent on getting LLMs to perform fact checking. This is not the technology for it. Sure it might sorta kinda work in some circumstances. That doesn't make it a good fit. Think of it like buying a refrigerator for storing clothes.
Nietzsche might say this is not the fantasy of truth, but of comfort. The Last Man wants a machine to say 'fact wrong' or 'fact right' so the abyss of no ultimate truth can be made small enough to sleep beside.
I assume you'd have access to AI lawyers too, better ones if you can pay for larger/newer models! Meanwhile the judges are N year old models because they are state funded, and they work 'fine'.
Re: Disagreement among frontier LLMs on real-world fact-checks
#317Earlier quoted context omitted.
were you involved in making the study? your bio says you work for them so you should probably indicate that in your comments. lack of agreement when there is no singular correct answer (or any answer at all) isn't a useful metric I ran into a lot of these kinds of issues when working on the Citation Needed WMF project (and related extensions). Truth is so often very nuanced.
They introduced themselves as the study author here: https://news.ycombinator.com/item?id=48307887#48307899
Re: Disagreement among frontier LLMs on real-world fact-checks
#318Earlier quoted context omitted.
Exactly what people do when they use LLMs for "fact-checking" online, and any verbose explanation would be mostly ignored anyway, when people ask political, ethical, or simply ambiguous questions that they hold any stakes in. Don't even need politics for it, there is no point in probing a mathematical black box for "how many soldiers died in the year X in war Y". Any original source is preferable to a blurry "summary…
ask the black box to search for the original source and verify it yourself?
Especially in niche subjects.
For factual claims, I've fared better with Wikipedia and looking up the sources linked there.
Anyway, as AI text and media generation erodes the credibility of all online sources, these questions about source checking matter less and less: what if the source itself is a long and convincing-sounding text with poor sources?
This problem existed before already, but it boils down to a simple fact:
logic or maths alone cannot derive an authority that verifies claims about the real world other than weighting texts.
The question "what is the current population if Paris" can be answered by LLMs, but basically only by weighting sources, and assigning some credibility to them.
There's no real point in getting some weighted average of sources on this question, but so far, it doesn't hurt either.
Re: Disagreement among frontier LLMs on real-world fact-checks
#319Earlier quoted context omitted.
I would think ‘false’ is the only correct answer a there’s no evidence to prove the claim, so the claim is safely assumed false. Then again maybe that’s why I’m an atheist, not an agnostic?
"False" isn't correct in strict boolean terms either, since that implies that the inverse is true. Claiming "there is extraterrestrial life in the universe" is false is logically equivalent to claiming that "no extraterrestrial life exists anywhere in the universe" is true. Both statements would have to be interpreted as "false" under your criteria, as neither has any evidence to substantiate it. That leads us to a l…
Re: Disagreement among frontier LLMs on real-world fact-checks
#320For instance see the folks who think that they have "awakened" their instance of ChatGPT.
Actual usage may diverge to a greater degree than models