Earlier quoted context omitted.
> Anthropic claims they solely use agents to code and don't modify any code manually. Have you used CC? It shows. They did not make their fortune off this, and it’s at least lost me a customer because of how sloppy it is. The model is good, and it’s why they have to gate access to it. I’d much rather use a different harness. I do think you’re on to something though. As societal wealth further concentrates among the f…
...uh, I think Claude Code is great, actually. A lot of that is indeed just the strength of the underlying model, but the local client is great too. Plan mode, checkpoints, subagents... I've been using Claude Code for a year now, and I feel like Anthropic has steadily been eliminating pain points. It's certainly a lot better than the Gemini cli!
We are changing our developer productivity experiment design
31–40 of 62 posts
Re: We are changing our developer productivity experiment design
#32Those developer quotes are tough to read. Rate limits are going to hit like a truck when the labs eventually need to make a profit.
At this point the AI labs would pretty much have to form an illegal price fixing cartel in order to jack the prices up, they've been competing to drive down prices for so long. They'd have to get the Chinese AI labs to go along with that price fixing too.
Re: We are changing our developer productivity experiment design
#33"I don't want to do this without AI" sounds like we're already well into the brain atrophy stage of this. Now what? (I'd think about it myself but....)
Re: We are changing our developer productivity experiment design
#34Earlier quoted context omitted.
...uh, I think Claude Code is great, actually. A lot of that is indeed just the strength of the underlying model, but the local client is great too. Plan mode, checkpoints, subagents... I've been using Claude Code for a year now, and I feel like Anthropic has steadily been eliminating pain points. It's certainly a lot better than the Gemini cli!
Functionality-wise, it's great, but it's a buggy mess, and it seems to be getting worse with each release.
Re: We are changing our developer productivity experiment design
#35Unless this measures the entire SDLC longitudinally (like say, over a year) I'm not interested. I too can tell Claude Code to do things all day every day, but unless we have data on the defect rate it doesn't matter at all.
Obviously this highly depends on your company and your setup and risk tolerance and what not.
Re: We are changing our developer productivity experiment design
#36> When surveyed, 30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI. This implies we are systematically missing tasks which have high expected uplift from AI. In fact, one of the developers in the original study later revealed on Twitter that he had already done exactly that during the study, i.e. filtered out tasks he prefered not to do w…
Re: We are changing our developer productivity experiment design
#37"I don't want to do this without AI" sounds like we're already well into the brain atrophy stage of this. Now what? (I'd think about it myself but....)
I'm pretty sure that this was exactly the response to the first generation of devs who insisted on coding with a terminal instead of submitting punch cards like "real programmers".
Re: We are changing our developer productivity experiment design
#38Re: We are changing our developer productivity experiment design
#39Those developer quotes are tough to read. Rate limits are going to hit like a truck when the labs eventually need to make a profit.
Re: We are changing our developer productivity experiment design
#40Really interesting updates to their 2025 experiment. Repeat devs from the original experiment went from 0-40% slowdown to now -10-40% speedup - and METR estimates this as a 'lower-bound' more devs saying they dont even want to do 50% of their work without AI, even for 50/hr 30-50% of devs decided not to submit certain tasks without AI, missing the tasks with the highest uplift it also seems like there is a skill gap…
So while some study participants probably are seeing an actual speedup because of the discipline with which they manage their codebase's structure & documentation, other study participants are actually getting worse at non-AI coding.
...and METR's study can't tell which is which because METR's study isn't using any sort of codebase quality metrics for grounding.