Live data from Hacker News

When LLM judges agree, should we believe them?

amazon.science

41–45 of 45 posts

Re: When LLM judges agree, should we believe them?

#41
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.

Oh. Yes. Sure. If there is an error in the training data then they will obviously cite that error.

I'm mostly talking about random coding errors.

Re: When LLM judges agree, should we believe them?

#42
post #14

Earlier quoted context omitted.

the LLM-speak is unbearable

It’s awful. A ton of words to say absolutely nothing.

I do not understand why anyone would want an LLM to cosplay as them on a forum. I get maybe (selfish) idle curiosity but beyond that…?

Re: When LLM judges agree, should we believe them?

#43
post #34
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

No modern LLM can tell how many eyes the magic card Pit Imp has. They all say 2. This is across all reasoning levels and paid Gemini, Claude, and GPT (Sol) Drawing a line red to split up the image then has them answer correctly. Their failure modes are highly correlated.

[dead]

Re: When LLM judges agree, should we believe them?

#44
post #14

Earlier quoted context omitted.

It’s awful. A ton of words to say absolutely nothing.

I do not understand why anyone would want an LLM to cosplay as them on a forum. I get maybe (selfish) idle curiosity but beyond that…?

I don’t know why anyone just blatantly copies and paste LLM output. Laziness? Anxiety about being wrong? I don’t know.

Re: When LLM judges agree, should we believe them?

#45
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

> Two agents will almost never hallucinate in the same way, regardless of their weights

Citation needed

Post reply on HN