Live data from Hacker News

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

withspecific.com

1–10 of 147 posts

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#4
So TL;DR benchmarking in a completely non-reproducible manner ?

"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".

So basically pinky-promise benchmarking ?

I'm not sure I follow the value here ?

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#7

So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ?

If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#9

So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ?

Doesn’t it ultimately have to be this way, to prevent saturation?
Post reply on HN