Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
1–10 of 146 posts
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#2Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#3Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#4"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".
So basically pinky-promise benchmarking ?
I'm not sure I follow the value here ?
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#5Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#6A bit of a meta question: what are the most relevant benchmarks by now?
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#7So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ?
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#8A bit of a meta question: what are the most relevant benchmarks by now?
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#9So TL;DR benchmarking in a completely non-reproducible manner ? "Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company". So basically pinky-promise benchmarking ? I'm not sure I follow the value here ?