Live data from Hacker News

On the slow death of scaling

papers.ssrn.com

21–30 of 34 posts

Re: On the slow death of scaling

#21

>the acceptance that there are emergent properties which appear out of nowhere is another way of saying our scaling laws don’t actually equip us to know what is coming. Is this actually accepted? Ever since [0], I thought people recognized that they don't appear out of nowhere. [0] https://arxiv.org/pdf/2304.15004

"Appear out of nowhere" looks like a straw-man. Anyway, there are newer papers. For example "Emergent Abilities in Large Language Models: A Survey"[0]

[0] https://arxiv.org/abs/2503.05788

Re: On the slow death of scaling

#22

>the acceptance that there are emergent properties which appear out of nowhere is another way of saying our scaling laws don’t actually equip us to know what is coming. Is this actually accepted? Ever since [0], I thought people recognized that they don't appear out of nowhere. [0] https://arxiv.org/pdf/2304.15004

> I thought people recognized that they don't appear out of nowhere.

I don't think that paper is widely accepted. Have you seen the authors of that paper, or anyone else, use it to successfully predict (rather than postdict) anything?

Re: On the slow death of scaling

#23
post #3

Earlier quoted context omitted.

If anyone believes they're close to a generalization end-game wrt to AI capabilities, it makes no sense to do anything that could impact their advantage, by enabling others to compete. Collaboration makes sense on timeframes that don't imply zero-sum games. Board games like the Settlers of Catan are good examples of the behavior— concretely the start of the game when everyone trades vs the end of the game when if you…

> Collaboration makes sense on timeframes that don't imply zero-sum games. People are fooling themselves if they think AGI will be zero sum. Even if only one group somehow miraculously develops it, there will immediately be fast followers. And, the more likely scenario is more than one group would independently pull it off - if it's even possible.

[deleted]

Re: On the slow death of scaling

#25

>the acceptance that there are emergent properties which appear out of nowhere is another way of saying our scaling laws don’t actually equip us to know what is coming. Is this actually accepted? Ever since [0], I thought people recognized that they don't appear out of nowhere. [0] https://arxiv.org/pdf/2304.15004

"Appear out of nowhere" looks like a straw-man. Anyway, there are newer papers. For example "Emergent Abilities in Large Language Models: A Survey"[0] [0] https://arxiv.org/abs/2503.05788

I was struck by this in the abstract: The scaling of these models, accomplished by increasing the number of parameters and the magnitude of the training datasets, has been linked to various so-called emergent abilities that were previously unobserved. These emergent abilities, ranging from advanced reasoning and in-context learning to coding and problem-solving...

In my experience with agent assisted coding, how well it works seems very closely tied to the quantity and quality of training material. It also has some identifiable qualities like verifiability that make it a particularly good target for an LLM. I would not call that surprising or emergent.

Re: On the slow death of scaling

#26
> A pervasive belief in scaling has resulted in a massive windfall in capital for industry labs and fundamentally reshaped the culture of conducting science in our field.

People spend money on this because it works. It seems odd to call observable reality a "pervasive belief".

> Academia has been marginalized from meaningfully participating in AI progress and industry labs have stopped publishing.

Firstly, I still see news items about new models that are supposed to do more with less. If these are neither from academia nor industry, where are they coming from?

Secondly, "has been marginalized"? Really? Nobody's going to be uninterested in getting better results with less compute spend, attempts have just had limited effectiveness.

.

> However, it is unclear why we need so many additional weights. What is particularly puzzling is that we also observe that we can get rid of most of these weights after we reach the end of training with minimal loss

I thought the extra weights were because training takes advantage of high-dimensional bullshit to make the math tractable. And that there's some identifiable point where you have "enough" and more doesn't help.

I hadn't heard that anyone had a workable way to remove the extra ones after training, so that's cool.

.

.

The impression I had is that there's a somewhat-fuzzy "correct" number of weights and amount of training for any given architecture and data set / information content. And that when you reach that point is when you stop getting effort-free results by throwing hardware at the problem.

Re: On the slow death of scaling

#27
post #3

Earlier quoted context omitted.

If anyone believes they're close to a generalization end-game wrt to AI capabilities, it makes no sense to do anything that could impact their advantage, by enabling others to compete. Collaboration makes sense on timeframes that don't imply zero-sum games. Board games like the Settlers of Catan are good examples of the behavior— concretely the start of the game when everyone trades vs the end of the game when if you…

> Collaboration makes sense on timeframes that don't imply zero-sum games. People are fooling themselves if they think AGI will be zero sum. Even if only one group somehow miraculously develops it, there will immediately be fast followers. And, the more likely scenario is more than one group would independently pull it off - if it's even possible.

>if it's even possible.

Why do people keep repeating this. The only way artificial intelligence is impossible is if intelligence is impossible. And we're here so that pretty much removes that impediment.

Re: On the slow death of scaling

#30
Scaling works, the problem is that it is practically impossible to scale much. There is only so much energy, text data, GPUs, etc. The folly of scaling is that we are living in a finite world. The huge investments in AI for the past few years are probably hitting the practical limits of scaling for now.

I also feel like most insiders were fully aware of this fact, but it was a neat sales pitch.

Post reply on HN