Live data from Hacker News

13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

swe-rebench.com

1–10 of 17 posts

Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

#4
post #3

Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.

I think you just answered your own question.

Edit: In DeepSWE Sol High scores the same as Fable High for ~1/3 of the cost.

Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

#5
https://swe-rebench.com/about

> Potential data contamination: The SWE-bench dataset, comprising a collection of GitHub issues, has been publicly available since the end of 2023. As a result, models released after this date may have seen these exact issues or highly similar data during training. This raises the risk of inflated performance metrics and makes it harder to distinguish genuine generalization from memorization.

Re: 13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

#10
I've had much better experiences with the "budget" tier models than this test suggests I should. I rarely even consider the higher models due to price/use. Perhaps I'm just more vigilant about spec'ing my prompts out before submitting them? Am I just doing too pleb of work? I'm not trying to write custom cuda kernels or hardware integration.
Post reply on HN