Live data from Hacker News

When LLM judges agree, should we believe them?

amazon.science

41–47 of 47 posts

Re: When LLM judges agree, should we believe them?

#41
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.

Oh. Yes. Sure. If there is an error in the training data then they will obviously cite that error.

I'm mostly talking about random coding errors.

Re: When LLM judges agree, should we believe them?

#42
post #14

Earlier quoted context omitted.

the LLM-speak is unbearable

It’s awful. A ton of words to say absolutely nothing.

I do not understand why anyone would want an LLM to cosplay as them on a forum. I get maybe (selfish) idle curiosity but beyond that…?

Re: When LLM judges agree, should we believe them?

#43
post #34
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

No modern LLM can tell how many eyes the magic card Pit Imp has. They all say 2. This is across all reasoning levels and paid Gemini, Claude, and GPT (Sol) Drawing a line red to split up the image then has them answer correctly. Their failure modes are highly correlated.

[dead]

Re: When LLM judges agree, should we believe them?

#44
post #14

Earlier quoted context omitted.

It’s awful. A ton of words to say absolutely nothing.

I do not understand why anyone would want an LLM to cosplay as them on a forum. I get maybe (selfish) idle curiosity but beyond that…?

I don’t know why anyone just blatantly copies and paste LLM output. Laziness? Anxiety about being wrong? I don’t know.

Re: When LLM judges agree, should we believe them?

#45
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

> Two agents will almost never hallucinate in the same way, regardless of their weights

Citation needed

Re: When LLM judges agree, should we believe them?

#46
post #10

While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

By design, even the same LLM, when asked the same question multiple times, will almost never hallucinate in the same way. By that (flawed) logic, you could have the same LLM check itself.

Re: When LLM judges agree, should we believe them?

#47
Since the baseline LLM architecture is similar with each other, maybe already there are relations between their each opinion. Of course, this is an assumption and the explicit training seems effective in this case. But, I'm curious whether the approach would be effective in other cases (in terms of generalization?)
Post reply on HN