Fascinating case showing how LLM promoters will happily take "verified" benchmarks at their word. It's easy to publish "$NEWMODEL received an X% bump in SWE-Bench Verified!!!!". Proper research means interrogating the traces, like these researchers did (the Gist shows Claude 4 Sonnet): https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa... Commentary: https://x.com/bwasti/status/1963288443452051582 , https:/…
Top model scores may be skewed by Git history leaks in SWE-bench
41–50 of 167 posts
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#42In the meawhile, Oracle stock went up 40% in one one day, based on what Wall Street thinks AI might be...in 4 years...Not a bubble at all...
I think Oracle's stock mostly popped due to a delayed reaction with the US GSA contract it secured in July and the revenue guidance probably related to it: https://www.oracle.com/news/announcement/blog/oracle-cloud-c...
Wall Street is currently heavily punishing any company who misses their quarter, including NVIDIA!, after beating on their quarter.
Oracle had a earnings miss in the current quarter!
Their current REALITY is ~$15B quarterly revenue (with cloud infra ~$3B) and only ~$12B in near-term deferred backlog and deferred backlog is NOT revenue. To justify the valuation, this would imply OCI going from ~$18B in FY26 to ~$140B by FY30 that is an insane promise of +$120B in 4 years but back-loaded into the year 3 or year 4. :-))
Capex needs ~$35B next year just to chase GPUs/power and if they miss one quarter the story implodes. The supposed rational, efficient market, is paying near $1T today for back-loaded hopes.
Is completely bubble math. Like anybody, including Oracle AND their Customers, have ANY idea of their Capex in 4 years.
Complete and total bubble.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#43It's honestly ridiculous they left git history lying around during a benchmark, and this benchmark made to ICLR in Jan 2024 and no one has detected this issue until now. I don't really trust any benchmarking or tools or claims from this space when they can make such huge basic errors.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#44Earlier quoted context omitted.
Are you going to rail on humans for making this mistake in the first place?
No because that's the baseline. It's what you do when you have no other choice. Railing against that would be pointless.
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#45[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
The comment you link to says that "we only performed a quick preliminary search" and "We do not have a method for automatically checking existing trajectories." In other words, it can't confirm that the issue only "affected a tiny fraction of existing agents in a tiny fraction of their runs" as you say. Are you saying that you have since separately confirmed this? Edit: That said, I’m willing to believe based on the…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#46I'm not surprised. People really thought the models just kept getting better and better?
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#47Fascinating case showing how LLM promoters will happily take "verified" benchmarks at their word. It's easy to publish "$NEWMODEL received an X% bump in SWE-Bench Verified!!!!". Proper research means interrogating the traces, like these researchers did (the Gist shows Claude 4 Sonnet): https://gist.github.com/jacobkahn/bd77c69d34040a9e9b10d56baa... Commentary: https://x.com/bwasti/status/1963288443452051582 , https:/…
Claude benchmarks poorly but vibes well. Gemini benchmarks well and vibes well. Grok benchmarks well but vibes poorly.
(yes I know you are gushing with anecdotes, the vibes are simply the approximate color of gray born from the countless black and white remarks.)
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#48Earlier quoted context omitted.
The comment you link to says that "we only performed a quick preliminary search" and "We do not have a method for automatically checking existing trajectories." In other words, it can't confirm that the issue only "affected a tiny fraction of existing agents in a tiny fraction of their runs" as you say. Are you saying that you have since separately confirmed this? Edit: That said, I’m willing to believe based on the…
Ya what he links directly contradicts what he's saying lol
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#49[I'm on the SWE-bench team] Multiple people have looked into this, for example right in that thread: https://github.com/SWE-bench/SWE-bench/issues/465#issuecomme... This issue had affected a tiny fraction of existing agents in a tiny fraction of their runs. And we've now issued a fix. This is a natural part of running a benchmark, I'm sure tiny things like this will keep on getting discovered and we'll keep on fixing…
Re: Top model scores may be skewed by Git history leaks in SWE-bench
#50Not “may be”: just look how swe-bench scores drop to single digits once it in C# https://arxiv.org/html/2506.12286v3
So the "Verified" part of "SWE Bench Verified" means.. not "Verified" at all. I don't get it, who is so opposed to doing the bare minimum of manual work and check what these models are doing? At least back in the day grad students doing an easy meta-paper understood it meant doing some repetitive manual work. Now we got benchmarks by hype vendors who think they can use the thing they are benchmarking to .. mark the b…
I doubt any of the AI company employees are encouraged to go looking for cheating