I’ve been running a bunch of coding agents on benchmarks recently as part of consulting, and this is actually much more impressive than it seems at first glance. 71.2% puts it at 5th, which is 4 points below the leader (four points is a lot) and just over 1% lower than Anthropic’s own submission for Claude Sonnet 4 - the same model these guys are running. But the top rated submissions aren’t running production produc…
This is classic Goodhart's law. "When a measure becomes a target, it ceases to be a good measure" https://en.wikipedia.org/wiki/Goodhart%27s_law
Qodo CLI agent scores 71.2% on SWE-bench Verified
51–60 of 60 posts
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#52Earlier quoted context omitted.
Finally someone mentions Refact, I was in contact with the team, rooting for them really.
Just looked them up. Their pricing is around buying "coins" with no transparency as to what that gets. Hard pass
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#53I do not want to pay API charges or be limited to a fixed number of "credits" per month.
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#54do we know anything about the size of the model? I can't find the answer.
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#55I updated to the latest version last night. Enjoyed seeing the process permission toggle (rwx). Was a refreshing change to keep the security minded folks less in panic with all the agentic coding adoptions :-)
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#56We need some international body to start running these tests… I just can’t trust these numbers any longer. We need a platform for this, something at least we can get some peer reviews
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#57We need some international body to start running these tests… I just can’t trust these numbers any longer. We need a platform for this, something at least we can get some peer reviews
I’m working on this at STAC Research and looking to connect with others interested in helping. Key challenges are ensuring impartiality (and keeping it that way), making benchmarks ungameable, and guaranteeing reproducibility. We’ve done similar work in finance and are now applying the same principles to AI.
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#58Earlier quoted context omitted.
Labeling them as “wrappers” and “niche business” indicates a strong cognitive bias already. Value can be created on both sides of the equation.
How so? They are wrappers, and it is niche.
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#59Earlier quoted context omitted.
I’m working on this at STAC Research and looking to connect with others interested in helping. Key challenges are ensuring impartiality (and keeping it that way), making benchmarks ungameable, and guaranteeing reproducibility. We’ve done similar work in finance and are now applying the same principles to AI.
That sounds amazing, mind telling us a little more?
The approach is to use workloads defined by developers and end users (not providers) that reflect their real-world tasks. E.g. in finance, delivering market snapshots to trading engines. We test full stacks, holding some layers constant so you can isolate the effect of hardware, software, or models. Every run goes through an independent third-party audit to ensure consistent conditions, no cherry-picking of results, and full disclosure of config and tuning, so that the results are reproducible and the comparisons are fair.
In finance, the benchmarks are trusted enough to drive major infrastructure decisions by the leading banks and hedge funds, and in some cases to inform regulatory discussions, e.g. around how the industry handles time synchronization.
Now starting to apply the same principles to the AI benchmarking space. Would love to talk to anyone who wants to be involved?
Re: Qodo CLI agent scores 71.2% on SWE-bench Verified
#60Earlier quoted context omitted.
That sounds amazing, mind telling us a little more?
Sure! STAC Research has been building and running benchmarks in finance for ~18 years. We’ve had to solve many of the same problems I think you’re highlighting here.. e.g. tech & model providers tuning specifically for the benchmark, results that get published but can’t be reproduced outside the provider’s lab, etc. The approach is to use workloads defined by developers and end users (not providers) that reflect thei…
So the business model would be AI foundries contracting you for evaluating their models?
Do you envision some kind of freely accessible platform for consulting the results?