Live data from Hacker News

Separating signal from noise in coding evaluations

openai.com

11–20 of 108 posts

Re: Separating signal from noise in coding evaluations

#14
post #3

Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.

Unless they have something in the labs that massively departs from their current products, AGI isn't on the table and is purely hype for marketing purposes.

Re: Separating signal from noise in coding evaluations

#18
post #16
post #9

Didn't we all know from the start that all of SWE-Bench was flawed? Even the authors concede the limitations and have long since moved on.

SWE-Bench Pro was created to replace SWE-Bench and fix these problems.

SWE-bench Verified was created to fix the problems of SWE-bench.

Then SWE-Bench Pro was created because SWE-bench Verified had flaws.

Now SWE-Bench Pro is shown to have flaws.

Re: Separating signal from noise in coding evaluations

#19
post #3

Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.

they should be consulting Donald Rumsfeld and make sure they implement the Unknown-Unknowns benchmark, because thats how they get you

Re: Separating signal from noise in coding evaluations

#20
post #3

Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.

This ties into the bias-variance tradeoff (https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff) common with building non-LLM models. The solutions can only be a) figure out how to get LLMs smaller with similar performance so they don't memorize things/game the benchmarks and b) build benchmarks that are indeed comprehensive for all real-world data, which is infeasible.
Post reply on HN