SWE-bench Verified no longer measures frontier coding capabilities
21–30 of 209 posts
Re: SWE-bench Verified no longer measures frontier coding capabilities
#22Re: SWE-bench Verified no longer measures frontier coding capabilities
#23> We audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions, despite our best efforts in improving on this in the initial creation of SWE-bench Verified. Is this saying a quarter* of the questions and answers were wrong, this whole time?! If so, how was this ever, in any way, a va…
[deleted]
Huh, that is very curious and interesting indeed. If that's indeed true, that Anthropic claims that pass rate while OpenAI claims the test cases are flawed and broken, then clearly one of them aren't telling their whole side...
Re: SWE-bench Verified no longer measures frontier coding capabilities
#24I don't understand these websites which force translation to my native language. I mean, it's fine as it's useful for many people, but where is the button for disabling it ? Or why is it enabled by default ? "codage de pointe" sounds so weird and cringe in French.
Does your browser request French via an Accept-Language header perhaps? What really infuriates me is when sites don’t respect that header and give you a translation based on IP location.
Re: SWE-bench Verified no longer measures frontier coding capabilities
#25Re: SWE-bench Verified no longer measures frontier coding capabilities
#26Earlier quoted context omitted.
[deleted]
> Curiously Opus 4.7 claims to have a 87.6% pass rate and Mythos claims to have a 93.9% pass rate... leading to the conclusion that it's actually possible to "solve" the problems that OpenAI claims are incorrect. Huh, that is very curious and interesting indeed. If that's indeed true, that Anthropic claims that pass rate while OpenAI claims the test cases are flawed and broken, then clearly one of them aren't telling…
https://news.ycombinator.com/item?id=47911074
Citation for the claimed pass rates is: https://llm-stats.com/benchmarks/swe-bench-verified
Re: SWE-bench Verified no longer measures frontier coding capabilities
#27> We audited a 27.6% subset of the dataset that models often failed to solve and found that at least 59.4% of the audited problems have flawed test cases that reject functionally correct submissions, despite our best efforts in improving on this in the initial creation of SWE-bench Verified. Is this saying a quarter* of the questions and answers were wrong, this whole time?! If so, how was this ever, in any way, a va…
The answer is “it works because ML wants to work.” It’s surprising how far you can get with something flawed. It’s also why such huge breakthroughs are possible by noting flaws others haven’t.
Re: SWE-bench Verified no longer measures frontier coding capabilities
#28Curiously Opus 4.7 claims to have a 87.6% pass rate and Mythos claims to have a 93.9% pass rate... leading to the conclusion that it's actually possible to "solve" the problems that OpenAI claims are incorrect.
Re: SWE-bench Verified no longer measures frontier coding capabilities
#29Its pretty clear that any benchmark that comes out will be outdated and exist within the training data with short measure. There will always be an incentive to optimize specifically for these benchmarks even if just for marketing material. Sure there is a training cutoff, but its usually only 3-6 months off of the public release dates. The problem with coding benchmarks then becomes creating novel benchmarks that are…
Re: SWE-bench Verified no longer measures frontier coding capabilities
#30this statement alone seems to invalidate the SWE-bench tests