Live data from Hacker News

Disagreement among frontier LLMs on real-world fact-checks

lenz.io

371–377 of 377 posts

Re: Disagreement among frontier LLMs on real-world fact-checks

#371

Earlier quoted context omitted.

Assuming this isn’t a satire reply: https://www.pnas.org/doi/10.1073/pnas.2603294123 Hope this helps!

This doesn’t show grok as a model has bias but only that the product that uses grok has bias. Even the referenced papers to show models can have bias don’t show anything about grok. Overall you have given me zero evidence that grok model itself has some political bias. FWIW I don’t mind bias but I haven’t seen evidence of it.

No serious person should use Grok for anything remotely important. Musk openly changes it when it gives responses that anger his followers

Re: Disagreement among frontier LLMs on real-world fact-checks

#372
post #219

Earlier quoted context omitted.

"All almonds are grown in the U.S. state of California." implies "No almonds are grown outside the U.S. state of California." You find one almond tree outside of California that grows almonds, where such almonds are grown intentionally, and the claim is false.

Nobody is saying the claim is true. This is a discussion of whether misleading could be a valid answer. I've been arguing if the model interprets the claim as an exaggeration, then misleading would be an acceptable answer, and due to California's dominance in the industry one could reasonably interpret a claim of this nature as an exaggeration. It's fine if you disagree, but I have never claimed the question was true…

A statement that is obviously false cannot reasonably be “misleading”.

Re: Disagreement among frontier LLMs on real-world fact-checks

#373
post #96
post #87

Earlier quoted context omitted.

Well then it shows that these models are using widely disparate training sets and have high confidence even when they shouldn't. Questions like "is mouthwash effective" presumably has one solid data source -- medical journals.

But the prompt didn't give the models the option to say "I don't know", so it wasn't a measure of their confidence.

I mean that's true but I don't think that's realistically what's going on when one model gives an unqualified "Yes" and the other gives an unqualified "no."

You can argue the study isn't as case-closed-decisive as we'd ideally like, but it's certainly evidence. It's probably hard to design a better study.

Re: Disagreement among frontier LLMs on real-world fact-checks

#374
post #372

Earlier quoted context omitted.

Nobody is saying the claim is true. This is a discussion of whether misleading could be a valid answer. I've been arguing if the model interprets the claim as an exaggeration, then misleading would be an acceptable answer, and due to California's dominance in the industry one could reasonably interpret a claim of this nature as an exaggeration. It's fine if you disagree, but I have never claimed the question was true…

A statement that is obviously false cannot reasonably be “misleading”.

An exaggeration can. If I said "the C language was a million times faster than python" that would be an exaggeration. It would both be obviously false (most things are only trivially faster) and misleading.

If the LLM interpreted the original statement as an exaggeration, then misleading could be an acceptable answer to a false statement.

Re: Disagreement among frontier LLMs on real-world fact-checks

#375

Earlier quoted context omitted.

how is that misleading if it's a fact, it's only misleading if you presume to know the reaction or intent behind making such a claim, and without context we should be extremely careful in making such presumptions.

It's misleading because a single murder in this case is not statistically significant, but phrasing it using probabilistic terminology (i.e. percentages) obscures that fact and implies that you have enough data for the probabilistic language to be relevant. Choosing to use percentages when there is a countable or small amount of data is typically misleading, even though it is "technically" true. In fact, a misleading…

>obscures ... implies

See: Presumption

Re: Disagreement among frontier LLMs on real-world fact-checks

#377
post #367
post #362

Earlier quoted context omitted.

> If we strictly follow logic, then nobody and nothing can claim that anything is true or false. Sure they can. They may or may not be correct, but that's a matter of empirical validation entirely distinct from the logic flow itself. Whether or not a conclusion is logically implied by its hypotheses has nothing to do with whether the input hypotheses are themselves true. Logic is just the reasoning process. > We just…

You fucked up causality. My sentence has a clear causality order, which you ignored. You are right if we ignore that, and you are right that people ignore logic many times. For a good reason. I replied like that because I think you applied logic too strictly already.

> You fucked up causality. My sentence has a clear causality order, which you ignored.

I'm afraid I'm not seeing anything related to causality in your previous comment. On top of that, any discussion of causality would also itself be an empirical question, and so would be upstream of drawing logical conclusions out of already accepted hypotheses.

> I replied like that because I think you applied logic too strictly already.

Boolean logic is consists of on deterministic evaluation of binary values, so there is no concept of "too strictly" applicable to it. It's essentially math, wherein there's a specific correct answer determined by rules that are always applied the same way.

Post reply on HN