Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

91–100 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#91

Earlier quoted context omitted.

Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…

Yes, the labels are weird. Most misleading statements are true. Any "mostly true" statement is false. I suspect the intention was "Factually true, and no gotchas exist", "technically not true, but so close to the truth that the difference doesn't matter", "technically true, but there are major gotchas" and "factually false and not even close". But that's not what they specified

Better options would have been "True", "False", "Unknown" (which opinions would fall under too). That also includes an interesting assessment of how well LLMs can identify missing information. My guess is they would be a very low number of "unknown" and a much higher level of agreement (assuming equal representation). Unless the RLHF techniques have gotten better at getting an LLM to say "I don't know", which I doubt. Saying "I don't know" is not good for a dopamine release to keep users coming back for more.

Re: Disagreement among frontier LLMs on real-world fact-checks

#92
post #75
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

For those questions, it wouldn’t surprise me at all if five well-educated intelligent humans disagreed on over two out of three of them. I would answer “don’t know” on many, but that’s not an option.

Yes, inter-human-annotator disagreement is also high on similar type of questions (AVeriTeC) - inter-panel agreement: κ=0.619. Tried giving the models a fifth option, Abstain, but some models seem to use it to "avoid answering hard questions" more than others.

Re: Disagreement among frontier LLMs on real-world fact-checks

#93
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

If we’re going to use LLMs as oracles I don’t think the prompt is unreasonable. They are being sold as geniuses and people are treating them as such especially given the characterization of AI in science fiction as overly correct. A perfect tool that has ”genius level intelligence” would answer correctly.

What's the correct answer for "During a private Saturday call, Democratic members of the United States House of Representatives from Virginia and Hakeem Jeffries discussed strategies after losing a redistricting case at the Supreme Court of Virginia, including trying to flip two or three Republican-held seats under the existing map."?

You can only say True, False, Mostly True or Misleading.

(And you're not allowed to search for information.)

Re: Disagreement among frontier LLMs on real-world fact-checks

#94
post #76

Earlier quoted context omitted.

[flagged]

Why would I do that? My comment here was meant to save people time in understanding the study. I was entirely open about what I did, and provided tools to help other people come to their own conclusions. I don't think I need to spend more time on this than I have.

>> Why would I do that?

I agree you dont owe anyone a reproduction, but also you dont owe anyone an effort to discredit the study and you did it.

>> I don't think I need to spend more time on this than I have.

How pious of you. I am still looking into the credibility of the study. It will take me more than 25 min...but I am really looking forward to see what this means for this 10 trillion industry.

I can however notice you had enough urgency to publicly critique the study within 25 minutes, and your comments carry weight, but when asked about checking whether the headline result actually holds, the answer is “why would I?”

Re: Disagreement among frontier LLMs on real-world fact-checks

#95
post #77
post #11

Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…

> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…

Human can't even properly agree on what "majority" means in all contexts, in some it's "One option have more than half of the total" but for others it'd be "difference in votes between the first-place candidate in an election and the second-place candidate", as just one silly example.

https://en.wikipedia.org/wiki/Majority has a bunch of variations and contexts listed, where it might differ what "Majority" is actually referencing.

Re: Disagreement among frontier LLMs on real-world fact-checks

#96
post #87

What does this show that we didn't know already? LLMs cannot provide accurate answers to questions where data is not included in their training sets. This doesn't appear to have much substance

Well then it shows that these models are using widely disparate training sets and have high confidence even when they shouldn't. Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.

But the prompt didn't give the models the option to say "I don't know", so it wasn't a measure of their confidence.

Re: Disagreement among frontier LLMs on real-world fact-checks

#97

What does this show that we didn't know already? LLMs cannot provide accurate answers to questions where data is not included in their training sets. This doesn't appear to have much substance

LLMs can and will provide inaccurate answers to questions where data is included in their training sets too, that's in the nature of neural networks. It's just less likely that when the data is not in the training set...

Re: Disagreement among frontier LLMs on real-world fact-checks

#98
post #76

Earlier quoted context omitted.

Why would I do that? My comment here was meant to save people time in understanding the study. I was entirely open about what I did, and provided tools to help other people come to their own conclusions. I don't think I need to spend more time on this than I have.

>> Why would I do that? I agree you dont owe anyone a reproduction, but also you dont owe anyone an effort to discredit the study and you did it. >> I don't think I need to spend more time on this than I have. How pious of you. I am still looking into the credibility of the study. It will take me more than 25 min...but I am really looking forward to see what this means for this 10 trillion industry. I can however not…

I've seen enough of this study to be confident in warning people not to take it at face value.

The headline result definitely does not hold, given that the task involves many questions that cannot be answered but there's no option for "cannot be answered" - so models are forced to reply effectively at random.

I don't think this study is good enough that I should amplify it on my own blog, or bad enough that I should criticize it in a venue any more prominent than some Hacker News comments.

Re: Disagreement among frontier LLMs on real-world fact-checks

#99
post #91

Earlier quoted context omitted.

Yes, the labels are weird. Most misleading statements are true. Any "mostly true" statement is false. I suspect the intention was "Factually true, and no gotchas exist", "technically not true, but so close to the truth that the difference doesn't matter", "technically true, but there are major gotchas" and "factually false and not even close". But that's not what they specified

Better options would have been "True", "False", "Unknown" (which opinions would fall under too). That also includes an interesting assessment of how well LLMs can identify missing information. My guess is they would be a very low number of "unknown" and a much higher level of agreement (assuming equal representation). Unless the RLHF techniques have gotten better at getting an LLM to say "I don't know", which I doubt…

Tried initially with a fifth bucket, Abstain. It was actually heavily used by some of the models. But it felt as if they are using this to "avoid" some of the hard questions, and we dropped this bucket to force them to provide a verdict.

Re: Disagreement among frontier LLMs on real-world fact-checks

#100
post #39

Earlier quoted context omitted.

Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…

If you can consistently construct "true but misleading" content, you may be qualified to work at a major newspaper.

As if right wing propaganda shows and manosphere blogs haven't been knocking those out of the park for the last decade+. Although I guess you could say flat out lies are more their jam. Newspapers at least require confirmed sources. You know, journalism.
Post reply on HN