Separating signal from noise in coding evaluations
21–30 of 108 posts
Re: Separating signal from noise in coding evaluations
#22Re: Separating signal from noise in coding evaluations
#23Earlier quoted context omitted.
SWE-Bench Pro was created to replace SWE-Bench and fix these problems.
SWE-bench Verified was created to fix the problems of SWE-bench. Then SWE-Bench Pro was created because SWE-bench Verified had flaws. Now SWE-Bench Pro is shown to have flaws.
Re: Separating signal from noise in coding evaluations
#24What is considered SOTA for SWE benchmarks now?
Re: Separating signal from noise in coding evaluations
#25What is considered SOTA for SWE benchmarks now?
Re: Separating signal from noise in coding evaluations
#26Re: Separating signal from noise in coding evaluations
#27Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.
Re: Separating signal from noise in coding evaluations
#28Earlier quoted context omitted.
SWE-Bench Pro was created to replace SWE-Bench and fix these problems.
SWE-bench Verified was created to fix the problems of SWE-bench. Then SWE-Bench Pro was created because SWE-bench Verified had flaws. Now SWE-Bench Pro is shown to have flaws.
Re: Separating signal from noise in coding evaluations
#29Re: Separating signal from noise in coding evaluations
#30Achieving AGI will be more than just passing all benchmarks, it has to account for the unknown problems too.
This ties into the bias-variance tradeoff ( https://en.wikipedia.org/wiki/Bias%E2%80%93variance_tradeoff ) common with building non-LLM models. The solutions can only be a) figure out how to get LLMs smaller with similar performance so they don't memorize things/game the benchmarks and b) build benchmarks that are indeed comprehensive for all real-world data, which is infeasible.
In one sense, yes, tradeoffs are inescapable as the scope expands to the maximal possible scope. In another sense... it depends on the level of abstraction we're talking about.