A real-world benchmark for AI code review
11–20 of 29 posts
Re: A real-world benchmark for AI code review
#12Some feedback for the team, looked at pricing page and saw it more expensive ($30/dev/mo) and highly limiting (20prs per month per user). We have devs putting up that many prs in a single day. With this kind of plan pretty much no way we would even try this product
Re: A real-world benchmark for AI code review
#13Company creates a benchmark. Same company is best in that benchmark. Story as old as time.
Re: A real-world benchmark for AI code review
#14I'm not as cynical as the others here; if there are no popular code review benchmarks why should they not design one? Apparently this is in support of their 2.0 release: https://www.qodo.ai/blog/introducing-qodo-2-0-agentic-code-r... > We believe that code review is not a narrow task; it encompasses many distinct responsibilities that happen at once. [...] > Qodo 2.0 addresses this with a multi-agent expert review ar…
A lot of this stuff is really new, and we will need to find ways to standardize, but it will take time and consensus.
It took 4 years after the release of the automobile to coin the term milage to refer to miles driven per unit of gasoline. We will in due time create the same metrics for AI.
Re: A real-world benchmark for AI code review
#15Still early in development and has a much simpler goal, but I like simple things that work well.
Re: A real-world benchmark for AI code review
#16Merged PRs being considered good code?
Re: A real-world benchmark for AI code review
#17Company creates a benchmark. Same company is best in that benchmark. Story as old as time.
>thing gets optimized
Re: A real-world benchmark for AI code review
#18> Qodo takes a different approach by starting with real, merged PRs Merged PRs being considered good code?
Re: A real-world benchmark for AI code review
#19Cmd+F - "Overfitting"...nothing. Nope, no mention of how they do anything to alleviate overfitting. These benchmarks are getting tiresome.