Live data from Hacker News

Top model scores may be skewed by Git history leaks in SWE-bench

github.com

61–70 of 167 posts

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#61

Earlier quoted context omitted.

The big labs are almost certainly using compiler/repl output for generated code as an oracle for RL. I doubt they have C# in the mix.

Why do you doubt that? It's a widely used language. And there is even an open source C# REPL.

[deleted]

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#62
post #53

[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…

> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…

I'm also on the SWE-bench team. This was simply a classic bug. We had code before that we believed was sufficient to hide / remove future GitHub history and it turns out it was not. We've patched it.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#63
post #53

Earlier quoted context omitted.

> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

We absolutely did not.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#64
post #56

Earlier quoted context omitted.

[flagged]

They really did a "trust me bro" and "do your own research" huh

the strange thing to me is that people would have it any other way. if you don't trust someone, why would you trust them to do the research for you? bit of entitlement if you ask me

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#66

Earlier quoted context omitted.

Ya what he links directly contradicts what he's saying lol

[flagged]

If you are going to represent your team in public, you owe them better than a response like this.

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#68

Fascinating case showing how LLM promoters will happily take "verified" benchmarks at their word. It's easy to publish "$NEWMODEL received an X% bump in SWE-Bench Verified!!!!". Proper research means interrogating the traces, like these researchers did (the Gist shows Claude 4 Sonnet): https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa... Commentary: https://x.com/bwasti/status/1963288443452051582 , https:/…

The best benchmark is the community vibe in the weeks following a release. Claude benchmarks poorly but vibes well. Gemini benchmarks well and vibes well. Grok benchmarks well but vibes poorly. (yes I know you are gushing with anecdotes, the vibes are simply the approximate color of gray born from the countless black and white remarks.)

the vibes are just a collection anecdotes

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#69

Earlier quoted context omitted.

> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y

We absolutely did not.

[deleted]

Re: Top model scores may be skewed by Git history leaks in SWE-bench

#70
post #55

Earlier quoted context omitted.

[flagged]

Unfortunately the bank account trajectories are not public, because unscupulous corporations such FAANG who let thousands of engineers wade through my chat messages on their platforms might not shy away from bribing academics to improve benchmarks of their billion-dollar AI initiatives. It's also a bribe if my sibling gets a job with $500k annual salary. Tech is not immune to it.

You realize that this problem in SWE-Bench was discovered and publicized by people within those FAANG corporations?
Post reply on HN