Looking at the benchmark, https://www.swebench.com/, about half of scored submissions score under 1/3 correct? So they're either not cheating, or not cheating effectively?
Some critical issues with the SWE-bench dataset
21–30 of 121 posts
Re: Some critical issues with the SWE-bench dataset
#22> When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. This matches my intuition about the coding performance of these models a lot better. I don't think any current coding benchmark accurately measures coding performance.
Re: Some critical issues with the SWE-bench dataset
#23I am shocked— shocked —when a vendor cheats in order to increase their benchmark scores. I always tell my customers to ignore benchmarks and compare outcomes with their own workloads. Benchmarks are almost completely useless in the real world.
Re: Some critical issues with the SWE-bench dataset
#24OAI, xAI, Antropic, Google all score incredibly well, then you go to try and write code and its just okay.
They claim it can do PHD level reasoning, but here I am not trusting it on basic computational thinking.
Re: Some critical issues with the SWE-bench dataset
#25> When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. This matches my intuition about the coding performance of these models a lot better. I don't think any current coding benchmark accurately measures coding performance.
o3-mini and gpt-4o are so piss poor in agent coding compared to claude that you don't even need a benchmark
Re: Some critical issues with the SWE-bench dataset
#26> When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. This matches my intuition about the coding performance of these models a lot better. I don't think any current coding benchmark accurately measures coding performance.
Re: Some critical issues with the SWE-bench dataset
#27> 32.67% of the successful patches involve cheating as the solutions were directly provided in the issue report or the comments. Looking at the benchmark, https://www.swebench.com/ , about half of scored submissions score under 1/3 correct? So they're either not cheating, or not cheating effectively?
Re: Some critical issues with the SWE-bench dataset
#28> 32.67% of the successful patches involve cheating as the solutions were directly provided in the issue report or the comments. Looking at the benchmark, https://www.swebench.com/ , about half of scored submissions score under 1/3 correct? So they're either not cheating, or not cheating effectively?
Re: Some critical issues with the SWE-bench dataset
#29> When we filtered out these problematic issues, the resolution rate of SWE-Agent+GPT-4 dropped from 12.47% to 3.97%. This matches my intuition about the coding performance of these models a lot better. I don't think any current coding benchmark accurately measures coding performance.
o3-mini and gpt-4o are so piss poor in agent coding compared to claude that you don't even need a benchmark
Re: Some critical issues with the SWE-bench dataset
#30I would argue almost every popular benchmark quoted by the big LLM companies is tainted. OAI, xAI, Antropic, Google all score incredibly well, then you go to try and write code and its just okay . They claim it can do PHD level reasoning, but here I am not trusting it on basic computational thinking.