Live data from Hacker News

Qodo CLI agent scores 71.2% on SWE-bench Verified

qodo.ai

51–60 of 60 posts

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#51
post #2

I’ve been running a bunch of coding agents on benchmarks recently as part of consulting, and this is actually much more impressive than it seems at first glance. 71.2% puts it at 5th, which is 4 points below the leader (four points is a lot) and just over 1% lower than Anthropic’s own submission for Claude Sonnet 4 - the same model these guys are running. But the top rated submissions aren’t running production produc…

This is classic Goodhart's law. "When a measure becomes a target, it ceases to be a good measure" https://en.wikipedia.org/wiki/Goodhart%27s_law

A specific setup for the benchmark is just plain cheating, not Goodhart’s law.

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#52

Earlier quoted context omitted.

Finally someone mentions Refact, I was in contact with the team, rooting for them really.

Just looked them up. Their pricing is around buying "coins" with no transparency as to what that gets. Hard pass

You realize that you can self-host their stuff? https://github.com/smallcloudai/refact

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#56
post #18

We need some international body to start running these tests… I just can’t trust these numbers any longer. We need a platform for this, something at least we can get some peer reviews

I’m working on this at STAC Research and looking to connect with others interested in helping. Key challenges are ensuring impartiality (and keeping it that way), making benchmarks ungameable, and guaranteeing reproducibility. We’ve done similar work in finance and are now applying the same principles to AI.

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#57
post #56
post #18

We need some international body to start running these tests… I just can’t trust these numbers any longer. We need a platform for this, something at least we can get some peer reviews

I’m working on this at STAC Research and looking to connect with others interested in helping. Key challenges are ensuring impartiality (and keeping it that way), making benchmarks ungameable, and guaranteeing reproducibility. We’ve done similar work in finance and are now applying the same principles to AI.

That sounds amazing, mind telling us a little more?

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#58
post #47

Earlier quoted context omitted.

Labeling them as “wrappers” and “niche business” indicates a strong cognitive bias already. Value can be created on both sides of the equation.

How so? They are wrappers, and it is niche.

I think those wrappers could create some potentially complex workflow around LLM API, with various trees of decisions, integrations, eval, rankers, ratets, etc, and this is their added value.

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#59
post #57
post #56

Earlier quoted context omitted.

I’m working on this at STAC Research and looking to connect with others interested in helping. Key challenges are ensuring impartiality (and keeping it that way), making benchmarks ungameable, and guaranteeing reproducibility. We’ve done similar work in finance and are now applying the same principles to AI.

That sounds amazing, mind telling us a little more?

Sure! STAC Research has been building and running benchmarks in finance for ~18 years. We’ve had to solve many of the same problems I think you’re highlighting here.. e.g. tech & model providers tuning specifically for the benchmark, results that get published but can’t be reproduced outside the provider’s lab, etc.

The approach is to use workloads defined by developers and end users (not providers) that reflect their real-world tasks. E.g. in finance, delivering market snapshots to trading engines. We test full stacks, holding some layers constant so you can isolate the effect of hardware, software, or models. Every run goes through an independent third-party audit to ensure consistent conditions, no cherry-picking of results, and full disclosure of config and tuning, so that the results are reproducible and the comparisons are fair.

In finance, the benchmarks are trusted enough to drive major infrastructure decisions by the leading banks and hedge funds, and in some cases to inform regulatory discussions, e.g. around how the industry handles time synchronization.

Now starting to apply the same principles to the AI benchmarking space. Would love to talk to anyone who wants to be involved?

Re: Qodo CLI agent scores 71.2% on SWE-bench Verified

#60
post #59
post #57

Earlier quoted context omitted.

That sounds amazing, mind telling us a little more?

Sure! STAC Research has been building and running benchmarks in finance for ~18 years. We’ve had to solve many of the same problems I think you’re highlighting here.. e.g. tech & model providers tuning specifically for the benchmark, results that get published but can’t be reproduced outside the provider’s lab, etc. The approach is to use workloads defined by developers and end users (not providers) that reflect thei…

Thank you, it’s quite brilliant to transfer those skill like this.

So the business model would be AI foundries contracting you for evaluating their models?

Do you envision some kind of freely accessible platform for consulting the results?

Post reply on HN