When LLM judges agree, should we believe them?
11–20 of 44 posts
Re: When LLM judges agree, should we believe them?
#12Without reading the article (doesn't matter if it's pro or contra): no, of course not. It shouldn't even be a debatable question.
> Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
Re: When LLM judges agree, should we believe them?
#13I believe it depends on the LLM itself. Like what model as each model has diff weights and diff data trained onn
Re: When LLM judges agree, should we believe them?
#14Re: When LLM judges agree, should we believe them?
#15They would all agree raspberry has two Rs
Re: When LLM judges agree, should we believe them?
#16For info: the advisor(s) available are higher end models. For example: you use sonnet, the available advisors are opus and fable. If you use Haiku, the advisor are sonnet, opus and fable.
Re: When LLM judges agree, should we believe them?
#17While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.
Re: When LLM judges agree, should we believe them?
#18While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.
It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.
Folks should sue in a class-action lawsuit, any legal firm worth their beautiful walnut desks would seriously be happy take on that constitutionally backed mission. =3
Re: When LLM judges agree, should we believe them?
#19Without reading the article (doesn't matter if it's pro or contra): no, of course not. It shouldn't even be a debatable question.
I think you should have read the article first, at minimum the subheader > Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
Because there's no "discounting of opinions". They are running a separate LLM to "score" opinions. And the result is still "no" regardless of "lineages" or "sources".
And the end of the article leads me to believe that the entire article and approach is LLM-induced garbage:
--- start quote ---
When LLM judges agree, we should ask why. Sometimes agreement is independent evidence. Sometimes it is a shared blind spot. A good aggregation method should be able to tell the difference.
--- end quote ---