Earlier quoted context omitted.
Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…
Yes, the labels are weird. Most misleading statements are true. Any "mostly true" statement is false. I suspect the intention was "Factually true, and no gotchas exist", "technically not true, but so close to the truth that the difference doesn't matter", "technically true, but there are major gotchas" and "factually false and not even close". But that's not what they specified
Disagreement among frontier LLMs on real-world fact-checks
91–100 of 377 posts
Re: Disagreement among frontier LLMs on real-world fact-checks
#92Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
For those questions, it wouldn’t surprise me at all if five well-educated intelligent humans disagreed on over two out of three of them. I would answer “don’t know” on many, but that’s not an option.
Re: Disagreement among frontier LLMs on real-world fact-checks
#93Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
If we’re going to use LLMs as oracles I don’t think the prompt is unreasonable. They are being sold as geniuses and people are treating them as such especially given the characterization of AI in science fiction as overly correct. A perfect tool that has ”genius level intelligence” would answer correctly.
You can only say True, False, Mostly True or Misleading.
(And you're not allowed to search for information.)
Re: Disagreement among frontier LLMs on real-world fact-checks
#94Earlier quoted context omitted.
[flagged]
Why would I do that? My comment here was meant to save people time in understanding the study. I was entirely open about what I did, and provided tools to help other people come to their own conclusions. I don't think I need to spend more time on this than I have.
I agree you dont owe anyone a reproduction, but also you dont owe anyone an effort to discredit the study and you did it.
>> I don't think I need to spend more time on this than I have.
How pious of you. I am still looking into the credibility of the study. It will take me more than 25 min...but I am really looking forward to see what this means for this 10 trillion industry.
I can however notice you had enough urgency to publicly critique the study within 25 minutes, and your comments carry weight, but when asked about checking whether the headline result actually holds, the answer is “why would I?”
Re: Disagreement among frontier LLMs on real-world fact-checks
#95Here's the prompt they used: Classify this claim as of : " " Output exactly one label: True, Mostly True, Misleading, or False. No explanations, no qualifiers. The claims look like this: https://lenz.io/research/llm-disagreement/data.csv I put that in Datasette Lite to make it easier to explore. Here's an example of a disagreement: https://lite.datasette.io/?csv=https%3A%2F%2Fstatic.simonwil... The claim was "All alm…
> The almond thing is false, but I'd argue that "misleading" might be defensible if you were to accompany it with "the majority of almonds are grown in California, but not all of them". The "majority" in this case meaning about 51%, according to Wikipedia[1]? How could 51% ever be considered to be close to "all", such that "misleading" would be a valid answer? Am I missing something? [1]: https://en.wikipedia.org/wik…
https://en.wikipedia.org/wiki/Majority has a bunch of variations and contexts listed, where it might differ what "Majority" is actually referencing.
Re: Disagreement among frontier LLMs on real-world fact-checks
#96What does this show that we didn't know already? LLMs cannot provide accurate answers to questions where data is not included in their training sets. This doesn't appear to have much substance
Well then it shows that these models are using widely disparate training sets and have high confidence even when they shouldn't. Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.
Re: Disagreement among frontier LLMs on real-world fact-checks
#97What does this show that we didn't know already? LLMs cannot provide accurate answers to questions where data is not included in their training sets. This doesn't appear to have much substance
Re: Disagreement among frontier LLMs on real-world fact-checks
#98Earlier quoted context omitted.
Why would I do that? My comment here was meant to save people time in understanding the study. I was entirely open about what I did, and provided tools to help other people come to their own conclusions. I don't think I need to spend more time on this than I have.
>> Why would I do that? I agree you dont owe anyone a reproduction, but also you dont owe anyone an effort to discredit the study and you did it. >> I don't think I need to spend more time on this than I have. How pious of you. I am still looking into the credibility of the study. It will take me more than 25 min...but I am really looking forward to see what this means for this 10 trillion industry. I can however not…
The headline result definitely does not hold, given that the task involves many questions that cannot be answered but there's no option for "cannot be answered" - so models are forced to reply effectively at random.
I don't think this study is good enough that I should amplify it on my own blog, or bad enough that I should criticize it in a venue any more prominent than some Hacker News comments.
Re: Disagreement among frontier LLMs on real-world fact-checks
#99Earlier quoted context omitted.
Yes, the labels are weird. Most misleading statements are true. Any "mostly true" statement is false. I suspect the intention was "Factually true, and no gotchas exist", "technically not true, but so close to the truth that the difference doesn't matter", "technically true, but there are major gotchas" and "factually false and not even close". But that's not what they specified
Better options would have been "True", "False", "Unknown" (which opinions would fall under too). That also includes an interesting assessment of how well LLMs can identify missing information. My guess is they would be a very low number of "unknown" and a much higher level of agreement (assuming equal representation). Unless the RLHF techniques have gotten better at getting an LLM to say "I don't know", which I doubt…
Re: Disagreement among frontier LLMs on real-world fact-checks
#100Earlier quoted context omitted.
Without providing definitions of "True / Mostly True / Misleading / False" to each rater, I rate the article's claim that "Only one verdict bucket can be correct per claim" as false . Something can be simultaneously "misleading" and either true or false. Which category should something go in if it's "mostly false"? How much can something be wrong before it goes from "mostly true" to "false" (objectively, both have so…
If you can consistently construct "true but misleading" content, you may be qualified to work at a major newspaper.