Live data from Hacker News

Are LLM merge rates not getting better?

entropicthoughts.com

91–100 of 175 posts

Re: Are LLM merge rates not getting better?

#91
post #72

Earlier quoted context omitted.

The issue with llm’s is trust. I don’t see that ever going away. Humans have learned to trust other humans over a large time scale with rules in place to control behaviour.

That's a big problem with very specific manifestations. My startup helps customers handle regulatory compliance, also by forwarding complex questions to a pool of consultants. We've compared now more than a hundred replies to that of GPT Pro, and the quality is roughly the same. Sometimes a little worse, sometimes a little better. Always more detailed. Never unacceptable. But how to convince our customers that we hav…

Yup exactly.

Being able to hold someone liable for a F up has been how we have been able to function as a society and get to where we are today.

Re: Are LLM merge rates not getting better?

#92
post #80

Earlier quoted context omitted.

What the post is describing is just ANOVA. If removing a category improves the overall fit then fitting the two terms independently has the same optimal solution (with the two independent terms found to be identical). MSE never increases when adding a category. This is why you have to reach to things that penalize adding parameters to models when running model comparisons.

No, the post is doing cross-validation to test predictive power directly. The error will not decompose as neatly then.

Why would they do that and where do you see evidence they did?

Re: Are LLM merge rates not getting better?

#93
post #72

Earlier quoted context omitted.

The issue with llm’s is trust. I don’t see that ever going away. Humans have learned to trust other humans over a large time scale with rules in place to control behaviour.

That's a big problem with very specific manifestations. My startup helps customers handle regulatory compliance, also by forwarding complex questions to a pool of consultants. We've compared now more than a hundred replies to that of GPT Pro, and the quality is roughly the same. Sometimes a little worse, sometimes a little better. Always more detailed. Never unacceptable. But how to convince our customers that we hav…

LLM plus human should be better than either standalone. You won’t be able to make as much money scaling out, though.

You won’t be able to scale out and make as much money though. But surely you’re not only concerned about profit, right? What’s the point of life if you’re just trying to get rich.

Re: Are LLM merge rates not getting better?

#94
post #85

Earlier quoted context omitted.

Steet? Do you mean street? They're smarter in the same way a search engine is smarter.

Yes, "street". Typing from my phone, sorry. And search engines are narrow tools that can only output copies of its dataset. An LLM is capable of surprisingly novel output, even if the exact level of creativity is heavily debated.

Remixes aren't novel.

Re: Are LLM merge rates not getting better?

#95
post #80

Earlier quoted context omitted.

No, the post is doing cross-validation to test predictive power directly. The error will not decompose as neatly then.

Why would they do that and where do you see evidence they did?

Because it's a direct way to measure predictive power, and it says so: "We’ll use leave-one-out cross-validation"

Re: Are LLM merge rates not getting better?

#96
post #4

Interesting article, although with so few data points and such a specific time slice it is difficult to draw serious conclusions about the "improvement" of LLM models. It's notably lacking newer models (4.5 Opus, 4.6 Sonnet) and models from Gemini. LLMs appear to naturally progress in short leaps followed by longer plateaus, as breakthroughs are developed such as chain-of-thought, mixture-of-experts, sub-agents, etc.

[flagged]

Re: Are LLM merge rates not getting better?

#97

I am pretty convinced that for most types of day to day work, any perceived improvements from the latest Claude models for example were total placebo. In blind tests and with normal tasks, people would probably have no idea if they're using Opus 4.5 or 4.6.

I'd agree with you on 4.5 to 4.6, but going from gpt-5 or 4.0 to 4.5 was night and day.

GPT5 added the router, which was def a downgrade. 4.5 was probably the best non-COT model humanity has made. But too expensive to run.

Re: Are LLM merge rates not getting better?

#99
post #69

I don't find this very compelling. If you look at the actual graph they are referencing but never showing [1] there is a clear improvement from Sonnet 3.7 -> Opus 4.0 -> Sonnet 4.5. This is just hidden in their graph because they are only looking at the number of PRs that are mergable with no human feedback whatsoever (a high standard even for humans). And even if we were to agree that that's a reasonable standard, G…

Yes, I think this is basically an instance of the "emergent abilities mirage." https://arxiv.org/abs/2304.15004 If you measure completion rate on a task where a single mistake can cause a failure, you won't see noticeable improvements on that metric until all potential sources of error are close to being eliminated, and then if they do get eliminated it causes a sudden large jump in performance. That's fine if you ju…

That's how the public perceive it though.

It's useless and never gets better until it suddenly, unexpecty got good enough.

Re: Are LLM merge rates not getting better?

#100

There is a decent case for this thesis to hold true especially if we look at the shift in training regimes and benchmarking over the last 1-2 years. Frontier labs don't seem to really push pure size/capability anymore, it's an all in focus on agentic AI which is mainly complex post-training regimes. There are good reasons why they don't or can't do simple param upscaling anymore, but still, it makes me bearish on AGI…

> In practice this still doesn't mean 50 % of white collar can't be automated though.

Let me ask you this, though: if we wanted to, what percentage of white collar jobs could have been automated or eliminated prior to LLMs?

Meta has nearly 80k employees to basically run two websites and three mobile apps. There were 18k people working at LinkedIn! Many big tech companies are massive job programs with some product on the side. Administrative business partners, program managers, tech writers, "stewards", "champions", "advocates", 10-layer-deep reporting chains... engineers writing cafe menu apps and pet programming languages... a team working on in-house typefaces... the list goes on.

I can see AI producing shifts in the industry by reducing demand for meaningful work, but I doubt the outcome here is mass unemployment. There's an endless supply of bs jobs as long as the money is flowing.

Post reply on HN