Top model scores may be skewed by Git history leaks in SWE-bench
1–10 of 167 posts
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#2Re: Top model scores may be skewed by Git history leaks in SWE-bench
#3Re: Top model scores may be skewed by Git history leaks in SWE-bench
#4Very interested to see the updated results. This could really shake up the leaderboard.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#5Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
I don't get it, who is so opposed to doing the bare minimum of manual work and check what these models are doing? At least back in the day grad students doing an easy meta-paper understood it meant doing some repetitive manual work. Now we got benchmarks by hype vendors who think they can use the thing they are benchmarking to .. mark the bench.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#6Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
So kinda neat to see this paper!
[0]https://github.blog/news-insights/octoverse/octoverse-2024/#...
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#7Like, seriously, how come all these agents are beating Claude Code? In practice, they are shitty and not even close. Yes. I tried them.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#8Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
So the "Verified" part of "SWE Bench Verified" means.. not "Verified" at all. I don't get it, who is so opposed to doing the bare minimum of manual work and check what these models are doing? At least back in the day grad students doing an easy meta-paper understood it meant doing some repetitive manual work. Now we got benchmarks by hype vendors who think they can use the thing they are benchmarking to .. mark the b…
Seems on-brand for an LLM-related thing to claim that it has verified something without actually checking.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#9Re: Top model scores may be skewed by Git history leaks in SWE-bench
#10Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3