Live data from Hacker News

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

withspecific.com

141–147 of 147 posts

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#141

Earlier quoted context omitted.

hey this seems really interesting - what prompted you to test multiple agents on your codebase?

Curiosity and anxiety. I rebuilt my entire workflow around agents so the unease that the frontiers would change something (access, pricing, availability) and lock me out of that were high. Also why I spent way too much on hardware (at least that can be deducted). Now the whole stack could run in my house and I feel much better about the situation. Once I got the testing going though it is worth it for its own pursuit…

do you have some time to chat? janak@withspecific.com

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#142
post #108
post #65

Earlier quoted context omitted.

I agree somewhat with the way the agents behave but feel the opposite reaction. With Fable, I get exhausted because it's always dumping out paragraphs of text that explain one approach but have some secret gotcha thrown out in the last two sentences. Then I have to pause and consider the caveat and if it matters and it happens every single time Fable responds and that constantly needing to make a decision that could…

I'm the opposite. Every time I've let Sol/Astra be decisive, I ended up with an overengineered mess. I much prefer getting alerted when there's more than 1 approach to the problem and it's discovered mid-implementation. I don't want to do the grunt work of writing code, but I do want to know the architecture and be responsible for the decisions. Fable is also very good at pushing back when I propose something that wi…

> Fable is also very good at pushing back

Oh man! This also is a pet peeve of mine with Fable. I will look at what it's doing and say "Shouldn't it be done this way?" and then it will spend forever arguing with me that it should be done the way it wanted to do it. It seems to get stuck in a certain way of thinking and will insist its way is right until I can really prove it - or just go over to Astra.

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#144
post #35

I'd love to see these: - DeepSeek V4.1 Flash - Kimi K3 - GLM 5.3 (and flash) - hy4-preview - Grok 4.6 All of these can be acessed using a $10/mo OpenCode Go subscription.

Here's the list: 1 Fable 5.1 38.8% 2 GPT-6 Astra 33.8% 3 Gemini 3.8 Flash 31.2% 4 GLM 5.3 28.8% 5 Grok 4.6 23.8% 5 Muse Spark 1.3 23.8% 7 Kimi K3 18.8% 8 GPT-5.6 Sol 16.2% See number 4, 5 and 7.

Thank you sir!

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#146

Earlier quoted context omitted.

Curiosity and anxiety. I rebuilt my entire workflow around agents so the unease that the frontiers would change something (access, pricing, availability) and lock me out of that were high. Also why I spent way too much on hardware (at least that can be deducted). Now the whole stack could run in my house and I feel much better about the situation. Once I got the testing going though it is worth it for its own pursuit…

do you have some time to chat? janak@withspecific.com

[dead]

Re: Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

#147

I think these benchmarks are not that useful, e.g. this suggests Fable is better than Astra, but in practice Astra is waaaaaay faster (like 5x; it's not even close), and also waaaay less annoying to talk to. There's only two or three sane options here - you can easily try them all and pick yourself.

that's fair - for long horizon engineering tasks would speed still matter?

I'd say so. Do you want your results in a week or a day?
Post reply on HN