13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
1–10 of 17 posts
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#2They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#3Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#4Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.
I think you just answered your own question.
Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#5https://swe-rebench.com/about
> Potential data contamination: The SWE-bench dataset, comprising a collection of GitHub issues, has been publicly available since the end of 2023. As a result, models released after this date may have seen these exact issues or highly similar data during training. This raises the risk of inflated performance metrics and makes it harder to distinguish genuine generalization from memorization.
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#6What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#7Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#8Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.
[flagged]
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#9Weird
Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS
#10I've had much better experiences with the "budget" tier models than this test suggests I should. I rarely even consider the higher models due to price/use. Perhaps I'm just more vigilant about spec'ing my prompts out before submitting them? Am I just doing too pleb of work? I'm not trying to write custom cuda kernels or hardware integration.