Live data from Hacker News

Benchmarking GPT-5 on 400 real-world code reviews

qodo.ai

11–20 of 88 posts

Re: Benchmarking GPT-5 on 400 real-world code reviews

#12
post #4

> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These rankings are then aggregated to produce a performance score. So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that.

Why is it hard to ignore an attempt to assess reality that is not grounded in reality?

Re: Benchmarking GPT-5 on 400 real-world code reviews

#13
post #4

> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These rankings are then aggregated to produce a performance score. So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that.

That's how 99% of 'LLM benchmark numbers' circulating on the internet work.

Re: Benchmarking GPT-5 on 400 real-world code reviews

#14
Asking GPT 4o seems like an odd choice. I know this is not quite comparable to what they were doing, but asking different LLMs the following question > answer only with the name nothing more norting less.what currently available LLM do you think is the best?

Resulted in the following answers:

- Gemini 2.5 flash: Gemini 2.5 Flash

- Claude Sonnet 4: Claude Sonnet 4

- Chat GPT: GPT-5

To me its conceivable that GPT 4o would be biased toward output generated by other OpenAI models.*

Re: Benchmarking GPT-5 on 400 real-world code reviews

#15
post #4

> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These rankings are then aggregated to produce a performance score. So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that.

It’s a widely accepted eval technique and it’s called “llm as a judge”

Re: Benchmarking GPT-5 on 400 real-world code reviews

#17
post #3

> Unlike many public benchmarks, the PR Benchmark is private, and its data is not publicly released. This ensures models haven’t seen it during training, making results fairer and more indicative of real-world generalization. This is key. Public benchmarks are essentially trust-based and the trust just isn't there.

[deleted]

Re: Benchmarking GPT-5 on 400 real-world code reviews

#18
post #4

> Each model’s responses are ranked by a high-performing judge model — typically OpenAI’s o3 — which compares outputs for quality, relevance, and clarity. These rankings are then aggregated to produce a performance score. So there's no ground truth; they're just benchmarking how impressive an LLM's code review sounds to a different LLM. Hard to tell what to make of that.

That's how 99% of 'LLM benchmark numbers' circulating on the internet work.

[flagged]
Post reply on HN