The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change. EDIT: I was using gemini cli... it's not a harness issue lol
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
21–30 of 147 posts
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#22The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change. EDIT: I was using gemini cli... it's not a harness issue lol
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#23I think these benchmarks are not that useful, e.g. this suggests Fable is better than Astra, but in practice Astra is waaaaaay faster (like 5x; it's not even close), and also waaaay less annoying to talk to. There's only two or three sane options here - you can easily try them all and pick yourself.
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#24> Each task comes from a private production codebase that we licensed from a real-world company How does that work?
> How does that work? My gut feeling is that any serious real-world company with a proprietary codebase worth looking at would not be handing out the crown jewels to a third party. License or not. I don't doubt somebody licensed their codebase to them, I just have my doubts about who the "who" could be.
At any rate, I'm not sure it matters whose codebase it is. I'd even say that a shitty codebase might make for a better test.
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#25The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change. EDIT: I was using gemini cli... it's not a harness issue lol
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#26Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#27Earlier quoted context omitted.
If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.
> You're giving up transparency for it being harder to game But then if we take that argument to its natural extreme, surely it means people should take the marketing bullshit published in the 100-page system cards published by Anthropic & co as "valuable" too ?
But even so, pretty much yes: companies that actually have reliable and accurate info in their releases get trusted more. It takes time because the default is to disbelieve info from biased sources, but it is possible to trust some of them more than others.
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#28Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#29this is the closest benchmark to my experience using the model harness combo. Astra for as great as it is falls slightly behind Fable 5.1 for me for large feature work (although it comments code much better). in particular, Fable is able to assess priority better than Astra (meaning Astra sometimes does things that aren’t worthwhile while missing things that are clearly important, particularly on possible ballooning…
Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
#30A bit of a meta question: what are the most relevant benchmarks by now?
https://artificialanalysis.ai/evaluations/terminalbench-v4-0