Live data from Hacker News

Are LLM merge rates not getting better?

entropicthoughts.com

11–20 of 175 posts

Re: Are LLM merge rates not getting better?

#11

That's an interesting claim, but I don't see it in my own work. They have got better but it's very hard to quantify. I just find myself editing their work much less these days (currently using GPT 5.4).

Without meaning to sound dismissive, because I'm really not intending to, there's also the possibility that you've gotten worse after enough time using them. You're treating yourself as a constant in this, but man cannot walk in the same river twice.

Re: Are LLM merge rates not getting better?

#12
These studies are always really hard to judge the efficacy of. I would say though the most surprising thing to me about LLMs in the past year is how many people got hyped about the Opus 4.5 release. Having used Claude Code at work since it was released I haven't really noticed any step changes in improvement. Maybe that's because I've never tried to use it to one shot things?

Regardless I'm more inclined to believe that 4.5 was the point that people started using it after having given up on copy/pasting output in 2024. If you're going from chat to agentic level of interaction it's going to feel like a leap.

Re: Are LLM merge rates not getting better?

#13
post #7

Given that it is the general consensus that a step function occurred with Opus 4.5/4.6 only 3 months ago - it seems like an insane omission.

This has been the general consensus for about three years now. "Drastic increases in capability have happened the last 3-6 months" have been a constant refrain.

Without any data from the study past September I think its not unreasonable, if you want to make an argument based on evidence.

For me personally, I agree with you, I'm really seeing it as well.

Re: Are LLM merge rates not getting better?

#14
post #7

Given that it is the general consensus that a step function occurred with Opus 4.5/4.6 only 3 months ago - it seems like an insane omission.

There's a consensus that SOMETHING changed with Opus 4.5. It might have been the "merge rates" metric, it might have not.

I'm certainly getting faster and cleaner-looking solutions for certain issues on Opus 4.6 than I was 5 months ago, but I'm not sure about the ability to solve (or even weigh in) the actual hard stuff, i.e. the stuff I'm paid for.

And I'm definitely not sure about the supposed big step between 4.5 and 4.6. I'm literally not seeing any.

Re: Are LLM merge rates not getting better?

#15

These studies are always really hard to judge the efficacy of. I would say though the most surprising thing to me about LLMs in the past year is how many people got hyped about the Opus 4.5 release. Having used Claude Code at work since it was released I haven't really noticed any step changes in improvement. Maybe that's because I've never tried to use it to one shot things? Regardless I'm more inclined to believe t…

Nah, pre 4.5 it was not comfortable to use agentic coding.

Re: Are LLM merge rates not getting better?

#16
I agree completely. I haven't noticed much improvement in coding ability in the last year. I'm using frontier models.

What's been the game changer are tools like Claude Code. Automatic agentic tool loops purpose built for coding. This is what I have seen as the impetus for mainstream adoption rather than noticeable improvements in ability.

Re: Are LLM merge rates not getting better?

#19
post #7

Given that it is the general consensus that a step function occurred with Opus 4.5/4.6 only 3 months ago - it seems like an insane omission.

This has been the general consensus for about three years now. "Drastic increases in capability have happened the last 3-6 months" have been a constant refrain. Without any data from the study past September I think its not unreasonable, if you want to make an argument based on evidence. For me personally, I agree with you, I'm really seeing it as well.

> "Drastic increases in capability have happened the last 3-6 months" have been a constant refrain.

well, yeah. because that's been the experience for many people.

3 years ago, trying to use ChatGPT 3.5 for coding tasks was more of a gimmick than anything else, and was basically useless for helping me with my job.

today, agentic Opus 4.6 provides more value to me than probably 2 more human engineers on my team would

Re: Are LLM merge rates not getting better?

#20
post #8

> This means llms have not improved in their programming abilities for over a year. Isn’t that wild? Why is nobody talking about this? Because it's not true. They have improved tremendously in the last year, but it looks like they've hit a wall in the last 3 months. Still seeing some improvements but mostly in skills and token use optimization.

> but mostly in skills and token use optimization.

I have heard rumors that token use optimization has been a recent focus to try to tidy up the financials of these companies before they IPO. take that with a grain of salt though

Post reply on HN