Live data from Hacker News

The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

intrepidkarthi.com

91–100 of 115 posts

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#91
post #51

These studies are meaningless because speedup is heavily dependent on the kind of work you're doing. No doubt that you can do mechanical refactors 100x faster with AI, and also no doubt that using AI will be slower for tasks where it's less about writing code and more about context/world knowledge or building understanding. Averaging across these tasks doesn't make sense because everyone's work consists of a differen…

I have found llms to be utterly useless for frontend (tailwind included). That is, unless you're building a single page app/landing page that is the typical center column with a hero and below that a 3x3 feature grid with those same 3 colors that all the sloppers show off. I'm not a frontend dev, but these statements are starting to get outright disrespectful to those that are. Do you people understand how much "worl…

This isn't my experience at all, LLMs do graphs and more complex excel-like web pages very well. They also do dashboards very well. They even seem to do 3D stuff with three.js like video games pretty well too, although I haven't tried that myself. Maybe they can't do something "great", but they sure can do good enough in most cases, and good enough is already better than most websites.

>I'm not a frontend dev, but these statements are starting to get outright disrespectful to those that are.

Agree here, especially considering that usually "niche scientific codebases" have terrible code so you don't need a super smart model to get a good bost in software engineering.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#92
post #51

Earlier quoted context omitted.

I have found llms to be utterly useless for frontend (tailwind included). That is, unless you're building a single page app/landing page that is the typical center column with a hero and below that a 3x3 feature grid with those same 3 colors that all the sloppers show off. I'm not a frontend dev, but these statements are starting to get outright disrespectful to those that are. Do you people understand how much "worl…

Fair point, I was more trying to make a statement about the amount of training data available, not the "difficulty" of the task. I just used Tailwind as an example because it is so ubiquitous with so much training data for LLMs to learn from, while any niche application doesn't have that.

Training data has stopped being a good predictor of LLM abilities ever since they started doing heavy RL runs. I'm not sure how much corporate dashboards/I can't believe it's not excel stuff were in the training data, I guess not that many considering that stuff is almost always corporate and kept inside companies, and LLMs are still great at it, good enough to make people that used excel and used it well daily for 10+ years stop using it for lots of stuff.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#93
post #79

Earlier quoted context omitted.

>feel Productivity is not a feeling though. Either you show an increased productivity or it doesn't exist

It is however extremely hard to measure accurately with software engineering and every easy measurement immediately draws ire from devs here.

Yes and no. You can measure things:

* features delivered * time-to-market (idea -> production) * code quality metrics * bugs found / bugs fixed * tickets opened / closed (and with what resolution) / month * revenue vs expenses * (if you must) commits / month, code changes / month

But few software developers seem to do that because it's overhead.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#94
post #65

2025 is such old news that this just isn't relevant. METR already redid the study at a later date and now finds a likely 18% speedup "For the subset of the original developers who participated in the later study, we now estimate a speedup of -18% with a confidence interval between -38% and +9%" (note their use of - and + here could be slightly confusing but they do mean 18% faster per the post) https://metr.org/blog/…

Their followup study essentially says the followup study itself is possibly broken because developers will now not participate in some of the non-AI tasks and because the study pays less. I would not, at all, suggest that this second study corrects or debunks the first. Instead what it shows (if anything, i.e. if you can even put aside the regrettable choice to change the payment level, which affects applicant recrui…

The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#95
post #94
post #65

Earlier quoted context omitted.

Their followup study essentially says the followup study itself is possibly broken because developers will now not participate in some of the non-AI tasks and because the study pays less. I would not, at all, suggest that this second study corrects or debunks the first. Instead what it shows (if anything, i.e. if you can even put aside the regrettable choice to change the payment level, which affects applicant recrui…

The original study itself had at least one developer who later revealed that he had filtered out tasks he prefered not to do without AI: https://xcancel.com/ruben_bloom/status/1943536052037390531 -- given the N was 16, and he seems to have been one of the more AI-experienced devs, and we don't know if the other devs did this, the results of the first study itself could be questioned.

I am not at all suggesting the first study is good, or that I believe its conclusions.

(Or that the failure of the second study validates the conclusions of the first.)

I am just saying that people here who think the second study overturned, debunked or corrected the findings of the first are explicitly wrong, because even its authors admit it is a broken study.

It would take a non-broken study to do that, and it may not actually be possible anymore, which is perhaps the most useful finding of the second study.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#96

"19 August 2025" This may as well have been written in the stone ages, when we were banging AI rocks together. I just did a ~6 month project in ~2 weeks using a frontier model. I wouldn't even have attempted this kind work a year ago, with or without the AIs available at the time!

>I just did a ~6 month project in ~2 weeks using a frontier model. Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.

My guess at what's happening is that people are mostly using the tools on low impact or speculative projects. Notice that he said he wouldn't have attempted it without AI.

That's been my experience too. I had an idea that I didn't need so I hadn't bothered doing it, but AI made it easier to just have a go. I suspect people aren't using AI as much on their main profit-making projects (which also are going to be bigger, more complex and not greenfield - which is all harder for AI).

Also give it a chance - as you said "good" models have only been available very recently and you wouldn't expect everyone to start using them instantly.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#97
post #16

Earlier quoted context omitted.

I'm convinced this is what causes people to feel productive with vim

Triggered by both of these comments.. interaction mode dictates a style of thinking. I have to use a mouse, I'm forced to use my eyes, which also means I probably have to use a massive screen. I have to pay attention to some hyperactive Intellisense-like feature, I'm forced to remove my attention from the problem. It's like saying you're convinced people reporting they feel more productive in a mauve-coloured room ar…

I think most people who strongly identify with tools like vim do so out of a sense of identity-building to "be the kind of developer who is good at vim" / embody some kind of aesthetic or in-group signal moreso than an actual desire to be more effective at getting work done.

As long as you don't have some kind of stochastic or >5s impediment taking you out of a state of flow, most developers' productivity is going be vastly more influenced by their knowledge, understanding, and ability to focus on the problem they are working on than the marginal difference in time it requires to perform some navigation or editing task. Which is not to say that vim is bad or that you shouldn't use it, but that it's just a text editor and if you get triggered by someone not liking it or thinking it's more trouble than it's worth, it might be worth taking a step back and thinking about why it's something that triggers an emotional/defensive response, rather than the kind of reaction you'd have to someone liking strawberry more than vanilla.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#98
post #16

Earlier quoted context omitted.

Triggered by both of these comments.. interaction mode dictates a style of thinking. I have to use a mouse, I'm forced to use my eyes, which also means I probably have to use a massive screen. I have to pay attention to some hyperactive Intellisense-like feature, I'm forced to remove my attention from the problem. It's like saying you're convinced people reporting they feel more productive in a mauve-coloured room ar…

>> I'm forced to use my eyes, which also means I probably have to use a massive screen. I have to pay attention to some hyperactive Intellisense-like feature What the hell are you talking about?

I can quite easily (and often do) use a basic editor while staring at the wall. I've yet to use an IDE where there wasn't some idiotic race between keystrokes and whatever random latency language server just told it to insert parens or a newline after you already typed them, assuming the text is even visible on a 13" screen buried in sidebars and "essential" extensions. They're full attention tools which is a completely different mode of work than is otherwise possible.

It's not to say an IDE isn't a useful thing, they just have their place like anything else. I personally find autocompletion useful for a couple of weeks going into a new language or project after which it's very often more a distraction than a productivity enhancer. Same goes with e.g. Git integration. I wouldn't presume to say a Git integration user simply needs to learn Git in much the same way I wouldn't expect someone to tell me that I can't use Git just because I don't use the IDE Git integration. They're just tools

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#99
post #16

Earlier quoted context omitted.

Triggered by both of these comments.. interaction mode dictates a style of thinking. I have to use a mouse, I'm forced to use my eyes, which also means I probably have to use a massive screen. I have to pay attention to some hyperactive Intellisense-like feature, I'm forced to remove my attention from the problem. It's like saying you're convinced people reporting they feel more productive in a mauve-coloured room ar…

I think most people who strongly identify with tools like vim do so out of a sense of identity-building to "be the kind of developer who is good at vim" / embody some kind of aesthetic or in-group signal moreso than an actual desire to be more effective at getting work done. As long as you don't have some kind of stochastic or >5s impediment taking you out of a state of flow, most developers' productivity is going be…

> I think most people who strongly identify with tools like vim do so out of a sense of identity-building to "be the kind of developer who is good at vim" / embody some kind of aesthetic or in-group signal moreso than an actual desire to be more effective at getting work done.

This is the exact same sweeping inferential leap as the original comment. I happen to think people who drive red cars do so only because they want to incite a sense of danger and potency in their road opponents, people who wear boots obviously want to identify with Ukranians on the front lines and any claims it helps with their flat feet are obvious rubbish.

Tooling and language obsession is boring and borderline offensive to anyone who has been around for a few years. Imagine walking into someone's workplace and demanding they replace their well worn chair, would you do it? Imagine insisting someone use vim because their IDE didn't have a natural pipe-through-shell-command function.

Re: The gauge broke: devs felt 20% faster with AI, measured 19% slower (2025)

#100

"19 August 2025" This may as well have been written in the stone ages, when we were banging AI rocks together. I just did a ~6 month project in ~2 weeks using a frontier model. I wouldn't even have attempted this kind work a year ago, with or without the AIs available at the time!

>I just did a ~6 month project in ~2 weeks using a frontier model. Claims like this are hard for me to take seriously because 'good' models have been available since the start of the year. So, if they really 10x one's productivity, then people should be able to have gotten done 5 years worth of work since then, but I've never actually seen anybody show any project like this.

> 'good' models have been available since the start of the year

today: https://www.anthropic.com/news/redeploying-fable-5

35 days ago: https://www.anthropic.com/news/claude-opus-4-8

70 days ago: https://openai.com/index/introducing-gpt-5-5/ 77 days ago: https://www.anthropic.com/news/claude-opus-4-7

119 days ago: https://openai.com/index/introducing-gpt-5-4/

182 days ago: The start of the year

Post reply on HN