Live data from Hacker News

The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

intrepidkarthi.com

101–110 of 115 posts

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#101
post #15

Generation got cheap. Verification got expensive. That proves AI is capable of doing one part of the software engineering process. The 16 devs in the study trusted AI to write the code. Once we trust AI to do the verification as well we'll realise the gains we feel we're getting now. Essentially we're intentionally going slower on the second half because the trust is missing. Alternatively, rather than trusting AI to…

> Once we trust AI to do the verification as well we'll realise the gains we feel we're getting now. I built a UAT agent on top of claude-agent-sdk, it uses Playwright and can spin up a preview instance for PRs we open. It uses its knowledge of the code to create a test plan, runs that test plan, and takes screenshots as it does so. On a recent PR I made a change to our MFA implementation and assumed that I would nee…

I built a very similar tool recently mentioned elsewhere in these comments. I think with the current state of LLMs, harnesses, and related tooling, being able to create or setup self-eval tooling is the biggest differentiator between merely using LLMs to write code vs realizing true 10x productivity wins.

I'm curious whether this is something LLMs are eventually going to be good enough at doing, or something the average developer knows when and how to do, or if this is going be something that's too specialized or difficult for most developers and maybe the next generation of developer tooling products. Now that we're several months into Claude Code crossing the threshold of legitimacy and adoption, I've been surprised at how few projects or developers are doing this yet.

To a certain extent now all you need to do is ask Claude Code for browser automation workflows and CUJ tests in your repo, and ye shall receive, but probably something that just uses base playwright. It would be even better if you could ask to install or use a self-eval tool that already did everything you needed it to do and also knew how to specify/setup automations. I'm assuming the level of agency or mental overhead of embarking on a browser automation side quest is beyond what most developers are used to in the course of their regular work, even though it's not really as hard as it sounds now. If so then self-eval tooling could be a very promising new product category to sell to enterprises.

BTW if you have a link to your project I'd be interested in checking it out! $5.69 for a UAT run sounds very high to me based on how many tokens it typically takes for agents to create automations or steer my similar project, but it could be that your test workloads are much more exhaustive or high-dimensionality than mine are. This is what a basic "go to amazon.com and search for a product, then take a screenshot" automation looks like for ours: https://github.com/accretional/chromerpc/blob/main/recipes/s.... And this is our interactive/dynamic remote steering mode: https://github.com/accretional/chromerpc/blob/main/chrome-pr.... I decided to implement against the Chrome Devtools Protocol (one layer under Playwright) and use grpc service reflection to allow agents to dynamically discover/describe the entire chrome devtools api surface. I just started working on a way to gather traces and monitor/manage the automation run internals because I think there's a ton of opportunity in this problem space for orchestration and RL

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#102

Earlier quoted context omitted.

>I just did a ~6 month project in ~2 weeks using a frontier model. Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.

> 'good' models have been available since the start of the year today: https://www.anthropic.com/news/redeploying-fable-5 35 days ago: https://www.anthropic.com/news/claude-opus-4-8 70 days ago: https://openai.com/index/introducing-gpt-5-5/ 77 days ago: https://www.anthropic.com/news/claude-opus-4-7 119 days ago: https://openai.com/index/introducing-gpt-5-4/ 182 days ago: The start of the year

Opus 4.5/4.6 are what many people consider the first 'good' models and it's from last year/start of this year.

But fine, let's say everything before gpt 5.5 was unusable crap. Then there should still be projects that would normally have previously taken ~2 years done in just two months. Where are they?

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#103

Earlier quoted context omitted.

>I just did a ~6 month project in ~2 weeks using a frontier model. Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.

My guess at what's happening is that people are mostly using the tools on low impact or speculative projects. Notice that he said he wouldn't have attempted it without AI. That's been my experience too. I had an idea that I didn't need so I hadn't bothered doing it, but AI made it easier to just have a go. I suspect people aren't using AI as much on their main profit-making projects (which also are going to be bigger…

Sure, but "10x faster, but only applies on small greenfield, throwaway projects" is a major caveat. In fact, there's a good chance this doesn't disprove the original blog post, you could be way faster on small projects but slower on 'real' projects.

>Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.

But I'm not expecting everyone to have built something like that, but surely among millions of users someone should have, especially the people proclaiming insane productivity gains? There are no super impressive open source projects done using AI and all the companies boasting about how all their code is AI written now don't show much improvement either.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#104
post #28
post #21

Earlier quoted context omitted.

But this isn’t just a narrative but a study. Limited but a study

A study from 2025. Might as well be from five years ago.

So where are all those profits from the higher productivity?

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#105

Earlier quoted context omitted.

> 'good' models have been available since the start of the year today: https://www.anthropic.com/news/redeploying-fable-5 35 days ago: https://www.anthropic.com/news/claude-opus-4-8 70 days ago: https://openai.com/index/introducing-gpt-5-5/ 77 days ago: https://www.anthropic.com/news/claude-opus-4-7 119 days ago: https://openai.com/index/introducing-gpt-5-4/ 182 days ago: The start of the year

Opus 4.5/4.6 are what many people consider the first 'good' models and it's from last year/start of this year. But fine, let's say everything before gpt 5.5 was unusable crap. Then there should still be projects that would normally have previously taken ~2 years done in just two months. Where are they?

[deleted]

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#106

2025 is such old news that this just isn't relevant. METR already redid the study at a later date and now finds a likely 18% speedup "For the subset of the original developers who participated in the later study, we now estimate a speedup of -18% with a confidence interval between -38% and +9%" (note their use of - and + here could be slightly confusing but they do mean 18% faster per the post) https://metr.org/blog/…

[flagged]

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#107
post #5

This study was shown to be flawed at the time; METR has retracted it. And it doesn't take into account current frontier models. AI makes you more productive. This is no longer up for debate. The energy you spend arguing last year's talking points is better spent knuckling down and learning the tools.

[dead]

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#108
post #94
post #65

Earlier quoted context omitted.

Their followup study essentially says the followup study itself is possibly broken because developers will now not participate in some of the non-AI tasks and because the study pays less. I would not, at all, suggest that this second study corrects or debunks the first. Instead what it shows (if anything, i.e. if you can even put aside the regrettable choice to change the payment level, which affects applicant recrui…

The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.

Most useful comment in the thread — participant-level selection is exactly what METR's update flags as the reason their new data is weak. Curious which direction the filtering ran: "AI won't help here" and "I don't want to do this one manually" corrupt the estimate in opposite directions.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#109

These studies are meaningless because speedup is heavily dependent on the kind of work you're doing. No doubt that you can do mechanical refactors 100x faster with AI, and also no doubt that using AI will be slower for tasks where it's less about writing code and more about context/world knowledge or building understanding. Averaging across these tasks doesn't make sense because everyone's work consists of a differen…

[flagged]

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#110

One thing I've noticed with generative AI is it's now easier than ever to write more lines of code. Before, a backend guy asked to add an intranet page would make an austere page -bare html with barely any styling or javascript. Today, the same guy given the same task can turn in something with styling, javascript, internationalisation, interactive form validation, progress spinner, minification build stage, linting,…

[flagged]
Post reply on HN