Live data from Hacker News

When LLM judges agree, should we believe them?

amazon.science

31–40 of 43 posts

Re: When LLM judges agree, should we believe them?

#31

Earlier quoted context omitted.

This was the first time I'd heard that gotcha question. I just threw it at Opus 5: None — a bass in water is a fish, and fish are notoriously bad at music. The instrument version plays four strings as standard (five and six-string basses exist for players who want to go lower or higher), and it prefers to stay dry. Seems like a pretty good answer to me!

Indeed, giving a definite answer to an ambiguous nonsense question is still incorrect. A fish can play with as many strings as it finds, but only one when on a hook. Yet this too is an incorrect answer, as it again ignores the ambiguity in the phrasing. =3

This feels like an xkcd 169 situation, to be honest.

Re: When LLM judges agree, should we believe them?

#32

Earlier quoted context omitted.

Indeed, giving a definite answer to an ambiguous nonsense question is still incorrect. A fish can play with as many strings as it finds, but only one when on a hook. Yet this too is an incorrect answer, as it again ignores the ambiguity in the phrasing. =3

This feels like an xkcd 169 situation, to be honest.

It was actually a trivial allusion to a rather old poetic parable, and highlights a foundational flaw in LLM inference model statistical salience.

If a LLM based chat bot does ever answer it correctly, than you know with a fair degree of certainty it was content moderators stepping into the chat. Have a wonderful day. =3

https://en.wikisource.org/wiki/The_Poems_of_John_Godfrey_Sax...

https://en.wikipedia.org/wiki/Blind_men_and_an_elephant

Re: When LLM judges agree, should we believe them?

#33
post #12

Earlier quoted context omitted.

I think you should have read the article first, at minimum the subheader > Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.

This doesn't even make logical sense. Actual judges have highly correlated outputs. This would literally be actively sculpting the range of opinions you want to see. It's like the idea of political districting that thinks that the aim should be to balance each district between "the two" political parties. You're not doing anything but institutionalizing two political parties and constant conflict. You're setting the…

> the aim should be to balance each district between "the two" political parties.

I've never heard anyone explicitly advocating for gerrymandering in favor of conflict/variance before? Is that a real thing?

Re: When LLM judges agree, should we believe them?

#34
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

No modern LLM can tell how many eyes the magic card Pit Imp has. They all say 2. This is across all reasoning levels and paid Gemini, Claude, and GPT (Sol)

Drawing a line red to split up the image then has them answer correctly.

Their failure modes are highly correlated.

Re: When LLM judges agree, should we believe them?

#35

Earlier quoted context omitted.

This feels like an xkcd 169 situation, to be honest.

It was actually a trivial allusion to a rather old poetic parable, and highlights a foundational flaw in LLM inference model statistical salience. If a LLM based chat bot does ever answer it correctly, than you know with a fair degree of certainty it was content moderators stepping into the chat. Have a wonderful day. =3 https://en.wikisource.org/wiki/The_Poems_of_John_Godfrey_Sax... https://en.wikipedia.org/wiki/Bli…

OK, I still have no idea what you actually think the "correct" answer is.

Re: When LLM judges agree, should we believe them?

#36

Earlier quoted context omitted.

It was actually a trivial allusion to a rather old poetic parable, and highlights a foundational flaw in LLM inference model statistical salience. If a LLM based chat bot does ever answer it correctly, than you know with a fair degree of certainty it was content moderators stepping into the chat. Have a wonderful day. =3 https://en.wikisource.org/wiki/The_Poems_of_John_Godfrey_Sax... https://en.wikipedia.org/wiki/Bli…

OK, I still have no idea what you actually think the "correct" answer is.

I think they want the correct answer to be "your question doesn't really make sense, so I'm not going to answer it". (But I also think Opus's answer is better.)

Re: When LLM judges agree, should we believe them?

#37

Earlier quoted context omitted.

It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.

They have already shown algorithmic discrimination in predicting recidivism for brown people, as they are nonsensically overrepresented in the statistical data of US prison populations. Folks should sue in a class-action lawsuit, any legal firm worth their beautiful walnut desks would seriously be happy take on that constitutionally backed mission. =3

> brown people, as they are nonsensically overrepresented in the statistical data of US prison populations

"Nonsensical?" They commit violent crimes, they get prosecuted for said violent crimes, and are serving prison sentences for those crimes. The algorithm picks up on this trend using the same logic that insurance actuaries use, which has also been largely neutered by critical theory.

What even is the argument here-- they're all innocent? Cops are ignoring piles of dead white people and their white murderers to only go patrol brown neighborhoods? We both know neither claim is true. The usual complaint is that cops avoid their neighborhoods and/or are lazy in investigating the crimes they report. The idea of overpolicing has always been a Marxist double-bind...nonsensical, I daresay.

Re: When LLM judges agree, should we believe them?

#38
post #23
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

> and by having a second one (with a different context) check almost entirely eliminates the problem. You solved one of the largest problems with current LLMs. How is it possible that nobody tried that before? Because they do. There are already LLMs checking outputs of other LLMs, the bullshit answers that you see are the results of failures on that checks. If you remove all checks LLMs will create hallucinations eve…

I'm sorry you're so upset, but it's true. Try it for yourself - get your LLM to hallucinate something, and then paste that text into another chat window and ask it to verify the result for you.

Re: When LLM judges agree, should we believe them?

#39
post #34
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

No modern LLM can tell how many eyes the magic card Pit Imp has. They all say 2. This is across all reasoning levels and paid Gemini, Claude, and GPT (Sol) Drawing a line red to split up the image then has them answer correctly. Their failure modes are highly correlated.

Yes - those failures (like strawberry) are. And those failures are very rare, which is why you had to reach for the Pit Imp MTG card, which I had to Google to understand your point.

Re: When LLM judges agree, should we believe them?

#40

Earlier quoted context omitted.

OK, I still have no idea what you actually think the "correct" answer is.

I think they want the correct answer to be "your question doesn't really make sense, so I'm not going to answer it". (But I also think Opus's answer is better.)

All the answers I've seen so far are assuredly not inaccurate (the elephants trunk is like a snake), but never fully correct (an elephant is not a snake).

While the LLM spits out each ambiguous context search result, it never answers the actual query without a human cheaters help. =3

https://en.wikipedia.org/wiki/Pareidolia

Post reply on HN