Earlier quoted context omitted.
Claude code was severely degraded the last few weeks, very simple terminal prompts were failing for me that it never had problems with.
Follow the money. Or how much comes from your pocket vs. VC and big tech speculators.
Top model scores may be skewed by Git history leaks in SWE-bench
91–100 of 167 posts
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#92[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#93[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#94Earlier quoted context omitted.
I could've phrased it better.
You could rewrite it a 1000 times, if the underlying idea is the same, suggesting something you don't know it's true, the outcome would be the same. Or did you mean something else? What was your intention with the message?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#95Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
So the "Verified" part of "SWE Bench Verified" means.. not "Verified" at all. I don't get it, who is so opposed to doing the bare minimum of manual work and check what these models are doing? At least back in the day grad students doing an easy meta-paper understood it meant doing some repetitive manual work. Now we got benchmarks by hype vendors who think they can use the thing they are benchmarking to .. mark the b…
We've all read & analyzed a large number of agent trajectories. This loophole seems to be something that popped up with the more recent models and we simply weren't aware of it.
As discussed in the github issue, there's a fix in the new version of the SWE-bench containers (currently being rolled out) that makes sure that the relevant commits aren't available.
Part of what makes SWE-bench a very interesting benchmark is the enormous action space that agents that compete on it can take. However that also means that there's unexpected things happening when models get better. We're currently working on making all agent runs easily browsable on a website (rather than having to download our AWS buckets) to get even more eyes on the trajectories. Thanks to everyone who uncovered this loophole.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#96[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
Obviously having something available during test time is more valuable than buried somewhere in the pretraining mixture. But in pretraining it happens presumably with high probability (why wouldn't coding models pretrain on the entire github), while in test time it apparently happened only very occasionally?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#97hah the model should get extra credit for discovering this! > Now I understand the situation perfectly! The issue described in the problem statement is a real bug that was already identified and fixed in later versions of pytest. Since we're working with pytest 5.2.4, we need to apply the same fix. https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa...
Edit: I misunderstood what was being tested; the test is correct.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#98Earlier quoted context omitted.
You could rewrite it a 1000 times, if the underlying idea is the same, suggesting something you don't know it's true, the outcome would be the same. Or did you mean something else? What was your intention with the message?
I meant it as a hint for anyone inclined to dig deeper. It's a possibility rather than something we can confidently dismiss.
Thinking out loud also doesn't make defamation acceptable.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#99I created a new benchmark from Java commits that are new in the past 6 months to add some variety: https://brokk.ai/power-ranking
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#100Earlier quoted context omitted.
Is it wrong? Aren't ethics and intelligence two different axes?
Different, but probably not as orthogonal as one might think. E.g. cooperating ethics had been necessary for the further development of human populations intelligence (and culture, technology, material wealth, nutrition etc that lead to further increases in intelligence). So lack of ethics might be a sign of intelligence, but it's also a parasitic intelligence that benefits the individual, and beyond certain level an…
- don't lie too often
- don't kill members of the in group
Seems like these would be required for any group to survive, which makes sense why they are universal. All other rules/ethics seem to be dependent on resource scarcity.