Top model scores may be skewed by Git history leaks in SWE-bench
31–40 of 167 posts
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#32Re: Top model scores may be skewed by Git history leaks in SWE-bench
#33[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
Edit: That said, I’m willing to believe based on the information in the thread that this most likely only affects a tiny fraction of runs.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#34[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#35Earlier quoted context omitted.
I was going to argue "LLM's need code samples to-do well on languages and if we are honest C# is a language mostly held in private repo's" but Github's 2024 report[0] says its the 5th most used language (I'm to lazy to check if this report includes private repo's but I'll assume it doesn't). So kinda neat to see this paper! [0] https://github.blog/news-insights/octoverse/octoverse-2024/#...
5th most used language based on private repos that the group making the report has the exclusive direct access to seeing I don't see that contradicting your assumption
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#36That the answers have been available to them in the environment, and they’re still not hitting 100% on this benchmark is a damning indictment of SOTA model performance.
It really isn't. Do you expect SOTA models to answer any answered question on the internet with 100% accuracy? Congrats you just compressed the whole internet (at least a few zettabytes) into a model (a few TB at most?).
The test environment contains the answers to the questions.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#37How can we ever perform this sort of faux-neutral agentic evaluation in an environment where we want agents to have access to the sum total of knowledge (which will necessarily include being able to learn about the evaluation being conducted and its expectations)?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#38I'm not surprised. People really thought the models just kept getting better and better?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#39We relatively quickly identified that the testing set are taken directly from the training set, but the claim has been advertised already so they were more difficult to retract... if it were at all, I left shortly after.
The incentives are not aligned with accurate reporting.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#40Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
I was going to argue "LLM's need code samples to-do well on languages and if we are honest C# is a language mostly held in private repo's" but Github's 2024 report[0] says its the 5th most used language (I'm to lazy to check if this report includes private repo's but I'll assume it doesn't). So kinda neat to see this paper! [0] https://github.blog/news-insights/octoverse/octoverse-2024/#...