Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
1–10 of 46 posts
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#2Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#3Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#4The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#5Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#6The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does
> down to 0C / 32 F (he didn't test colder conditions)
I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#7Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the generated code is semantically identical to the reference solution but even slightly exceeds a length cutoff, it fails? That seems like a genuinely flawed criterion, and the LLM-based graders themselves feel highly unstable.
Honestly, my impression is that LLM advancement is now completely dictated by benchmarks. But looking at recent trends, token prices are skyrocketing while coding capabilities haven't shown any massive improvements beyond a certain model generation. In fact, comparing GPT-6 Astra to 5.6 SOL, SOL often writes better code.
Considering all this, while training across diverse domains can pack a model with various pieces of knowledge, ultimately, it feels impossible for this approach to do something like derive the theory of relativity from medieval knowledge.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#8Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#9I feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the…
Perhaps it would be possible if the classically observed reality turned out to be possible (consistent under some reasonable conditions) only as a consequences of some deeper, sufficiently determined theory; but that would go far beyond our current level of knowledge, at least as far as I can say.
Re: Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
#10The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does
I'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".
It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from Met office site on their Heathrow station.) While the temperature does drop below 0C sometimes, it is not consistently below zero.
In southern England pretty much nobody changes tyres for winter -- you just use the same set all year. Optimising for "2C in the wet" seems about right...