Earlier quoted context omitted.
The big labs are almost certainly using compiler/repl output for generated code as an oracle for RL. I doubt they have C# in the mix.
Why do you doubt that? It's a widely used language. And there is even an open source C# REPL.
Top model scores may be skewed by Git history leaks in SWE-bench
61–70 of 167 posts
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#62[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#63Earlier quoted context omitted.
> This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing them. You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case. It's like building a chroot and then allowing `cd ..` to break out of it. What other maybe extremely basic edge cases were missed? > This doesn't change the overal…
> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#64Earlier quoted context omitted.
[flagged]
They really did a "trust me bro" and "do your own research" huh
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#65Re: Top model scores may be skewed by Git history leaks in SWE-bench
#66Re: Top model scores may be skewed by Git history leaks in SWE-bench
#67Turns out the test shouldn't have the answers included in it?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#68Fascinating case showing how LLM promoters will happily take "verified" benchmarks at their word. It's easy to publish "$NEWMODEL received an X% bump in SWE-Bench Verified!!!!". Proper research means interrogating the traces, like these researchers did (the Gist shows Claude 4 Sonnet): https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa... Commentary: https://x.com/bwasti/status/1963288443452051582 , https:/…
The best benchmark is the community vibe in the weeks following a release. Claude benchmarks poorly but vibes well. Gemini benchmarks well and vibes well. Grok benchmarks well but vibes poorly. (yes I know you are gushing with anecdotes, the vibes are simply the approximate color of gray born from the countless black and white remarks.)
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#69Earlier quoted context omitted.
> You're all extremely clever and I can't seem to understand how you missed thinking about such a simple edge case [...] I wouldn't be surprised if they left this loophole on purpose to give some (their?) agents extra leverage. Edit #1: I didn't mean to imply bad intent; just thinking out loud. Edit #2: Please, downvote responsibly. I deserve every one. https://www.youtube.com/watch?v=0FHEeG_uq5Y
We absolutely did not.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#70Earlier quoted context omitted.
[flagged]
Unfortunately the bank account trajectories are not public, because unscupulous corporations such FAANG who let thousands of engineers wade through my chat messages on their platforms might not shy away from bribing academics to improve benchmarks of their billion-dollar AI initiatives. It's also a bribe if my sibling gets a job with $500k annual salary. Tech is not immune to it.