Live data from Hacker News

Top model scores may be skewed by Git history leaks in SWE-bench

github.com

91–100 of 167 posts

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#91
post #75
post #60

Earlier quoted context omitted.

Claude code was severely degraded the last few weeks, very simple terminal prompts were failing for me that it never had problems with.

Follow the money. Or how much comes from your pocket vs. VC and big tech speculators.

They did a big fundraising round right after so it's easy to suspect they were manipulating profitability growth for it.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#92
post #53

[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…

> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…

[Also on the SWE-bench team] Part of the reason why this didn't surface earlier was that it only seems to affect more recent models, maybe the result of reward hacking during posttraining. We're currently working on making trajectories easier to access for everyone through a web tool (rather than having to download things from aws) to get even more eyes on the trajectories. The interface will also include search & LM inspection tools to specifically look for anything that might qualify as cheating.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#93

[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…

SGTM. The transparency is good.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#94

Earlier quoted context omitted.

I could've phrased it better.

You could rewrite it a 1000 times, if the underlying idea is the same, suggesting something you don't know it's true, the outcome would be the same. Or did you mean something else? What was your intention with the message?

I meant it as a hint for anyone inclined to dig deeper. It's a possibility rather than something we can confidently dismiss.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#95
post #5
post #2

Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3

So the "Verified" part of "SWE Bench Verified" means.. not "Verified" at all. I don't get it, who is so opposed to doing the bare minimum of manual work and check what these models are doing? At least back in the day grad students doing an easy meta-paper understood it meant doing some repetitive manual work. Now we got benchmarks by hype vendors who think they can use the thing they are benchmarking to .. mark the b…

[On the SWE-bench team] As someone pointed out SWE-bench Verified is a subset of tasks that were reviewed to be solvable (i.e., have enough context in the task description) as well are scored with unit tests that aren't overly specific to rule out valid solutions.

We've all read & analyzed a large number of agent trajectories. This loophole seems to be something that popped up with the more recent models and we simply weren't aware of it.

As discussed in the github issue, there's a fix in the new version of the SWE-bench containers (currently being rolled out) that makes sure that the relevant commits aren't available.

Part of what makes SWE-bench a very interesting benchmark is the enormous action space that agents that compete on it can take. However that also means that there's unexpected things happening when models get better. We're currently working on making all agent runs easily browsable on a website (rather than having to download our AWS buckets) to get even more eyes on the trajectories. Thanks to everyone who uncovered this loophole.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#96

[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…

Even if this bug never existed, models can still see lookahead commits during pretraining. Do we expect this bug to have a greater impact than the pretraining leakage?

Obviously having something available during test time is more valuable than buried somewhere in the pretraining mixture. But in pretraining it happens presumably with high probability (why wouldn't coding models pretrain on the entire github), while in test time it apparently happened only very occasionally?

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#97
post #78

hah the model should get extra credit for discovering this! > Now I understand the situation perfectly! The issue described in the problem statement is a real bug that was already identified and fixed in later versions of pytest. Since we're working with pytest 5.2.4, we need to apply the same fix. https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa...

Am I to interpret https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa... as it making a test that only asserts false and saying that the test exercises the function in question?

Edit: I misunderstood what was being tested; the test is correct.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#98

Earlier quoted context omitted.

You could rewrite it a 1000 times, if the underlying idea is the same, suggesting something you don't know it's true, the outcome would be the same. Or did you mean something else? What was your intention with the message?

I meant it as a hint for anyone inclined to dig deeper. It's a possibility rather than something we can confidently dismiss.

If it's a possibility and you don't want to dig deeper better to sit out and not comment anything at all, lest you risk defamation.

Thinking out loud also doesn't make defamation acceptable.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#99
swe-bench's bigger problems include (1) labs train on the test and (2) 50% of the tickets are from django; it's not a representative dataset even if all you care about is Python.

I created a new benchmark from Java commits that are new in the past 6 months to add some variety: https://brokk.ai/power-ranking

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#100
post #87
post #74

Earlier quoted context omitted.

Is it wrong? Aren't ethics and intelligence two different axes?

Different, but probably not as orthogonal as one might think. E.g. cooperating ethics had been necessary for the further development of human populations intelligence (and culture, technology, material wealth, nutrition etc that lead to further increases in intelligence). So lack of ethics might be a sign of intelligence, but it's also a parasitic intelligence that benefits the individual, and beyond certain level an…

Aren't there only two rules that all groups follow in the animal kingdom?

- don't lie too often

- don't kill members of the in group

Seems like these would be required for any group to survive, which makes sense why they are universal. All other rules/ethics seem to be dependent on resource scarcity.

Post reply on HN