Live data from Hacker News

We are changing our developer productivity experiment design

metr.org

31–40 of 62 posts

Re: We are changing our developer productivity experiment design

#31

Earlier quoted context omitted.

> Anthropic claims they solely use agents to code and don't modify any code manually. Have you used CC? It shows. They did not make their fortune off this, and it’s at least lost me a customer because of how sloppy it is. The model is good, and it’s why they have to gate access to it. I’d much rather use a different harness. I do think you’re on to something though. As societal wealth further concentrates among the f…

...uh, I think Claude Code is great, actually. A lot of that is indeed just the strength of the underlying model, but the local client is great too. Plan mode, checkpoints, subagents... I've been using Claude Code for a year now, and I feel like Anthropic has steadily been eliminating pain points. It's certainly a lot better than the Gemini cli!

Functionality-wise, it's great, but it's a buggy mess, and it seems to be getting worse with each release.

Re: We are changing our developer productivity experiment design

#32
post #13

Those developer quotes are tough to read. Rate limits are going to hit like a truck when the labs eventually need to make a profit.

At this point the AI labs would pretty much have to form an illegal price fixing cartel in order to jack the prices up, they've been competing to drive down prices for so long. They'd have to get the Chinese AI labs to go along with that price fixing too.

You don't need collusion, just the VC money drying up. Economic reality will set the base price.

Re: We are changing our developer productivity experiment design

#33
post #15

"I don't want to do this without AI" sounds like we're already well into the brain atrophy stage of this. Now what? (I'd think about it myself but....)

I don’t want to do work around the house without a fully charged battery for my ryobi either. I don’t want to go on a groccery run without my car. Using tools is not brain atrophy

Re: We are changing our developer productivity experiment design

#34
post #31

Earlier quoted context omitted.

...uh, I think Claude Code is great, actually. A lot of that is indeed just the strength of the underlying model, but the local client is great too. Plan mode, checkpoints, subagents... I've been using Claude Code for a year now, and I feel like Anthropic has steadily been eliminating pain points. It's certainly a lot better than the Gemini cli!

Functionality-wise, it's great, but it's a buggy mess, and it seems to be getting worse with each release.

I’m a heavy user for about four months now, and it’s definitely getting better for me. How would you say it’s getting worse?

Re: We are changing our developer productivity experiment design

#35
post #9

Unless this measures the entire SDLC longitudinally (like say, over a year) I'm not interested. I too can tell Claude Code to do things all day every day, but unless we have data on the defect rate it doesn't matter at all.

I really am quite in awe of Claude Code recently, so definitely not a naysayer, but this is a really important point. It’s so easy to create code, but am I shipping that much to prod than I used to? A bit.

Obviously this highly depends on your company and your setup and risk tolerance and what not.

Re: We are changing our developer productivity experiment design

#36
post #21

> When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI. This implies we are systematically missing tasks which have high expected uplift from AI. In fact, one of the developers in the original study later revealed on Twitter that he had already done exactly that during the study, i.e. filtered out tasks he prefered not to do w…

As one of the naysayers who talked a lot about the original study, I enthusiastically endorse any attempt at all to actually measure AI productivity. An increase from 20% slowdown to 20% speedup over the past year seems broadly consistent with my understanding of how things have gone. I think I remain classified as a "naysayer", though, because the "booster" case has gone from "I'm multiple times more productive" to "I never have to look at code my AI agents just handle everything" over the same period.

Re: We are changing our developer productivity experiment design

#37
post #15

"I don't want to do this without AI" sounds like we're already well into the brain atrophy stage of this. Now what? (I'd think about it myself but....)

I'm pretty sure that this was exactly the response to the first generation of devs who insisted on coding with a terminal instead of submitting punch cards like "real programmers".

Hum... People have this reaction in all kinds of situations, but I don't think programmers ever reacted like that to a large change in abstraction.

Re: We are changing our developer productivity experiment design

#39

Those developer quotes are tough to read. Rate limits are going to hit like a truck when the labs eventually need to make a profit.

Keep in mind that they make large profit on inference. Not enough to make up for losses on training but it won’t be a problem for Chinese labs which will just steal their weights.

Re: We are changing our developer productivity experiment design

#40
post #2

Really interesting updates to their 2025 experiment. Repeat devs from the original experiment went from 0-40% slowdown to now -10-40% speedup - and METR estimates this as a 'lower-bound' more devs saying they dont even want to do 50% of their work without AI, even for 50/hr 30-50% of devs decided not to submit certain tasks without AI, missing the tasks with the highest uplift it also seems like there is a skill gap…

There are some people participating in the study who will fire & forget instructions to Claude/Codex running in parallel worktrees, but would really struggle if they were required to work on their project without AI assistance.

So while some study participants probably are seeing an actual speedup because of the discipline with which they manage their codebase's structure & documentation, other study participants are actually getting worse at non-AI coding.

...and METR's study can't tell which is which because METR's study isn't using any sort of codebase quality metrics for grounding.

Post reply on HN