I really wish more people skeptical of AI capabilities would read about scaling laws -- Lilian is always so marvelous at giving a deep overview of the technical side but the whole point of this is: there are scaling laws, and they hold and continue to hold. This is such a huge basis for the predictions about AI capabilities for the past like 5 years.
And sitting right next to the data and compute factors in every cross entropy loss equation is the entropy of the language, which is just a fixed constant. There’s such a hard cap on cross entropy loss training and I never hear it come up!
Scaling Laws, Carefully
21–22 of 22 posts
Re: Scaling Laws, Carefully
#22I really wish more people skeptical of AI capabilities would read about scaling laws -- Lilian is always so marvelous at giving a deep overview of the technical side but the whole point of this is: there are scaling laws, and they hold and continue to hold. This is such a huge basis for the predictions about AI capabilities for the past like 5 years.
Why should the skeptics be reading it? The scaling laws show diminishing returns on more training data and larger models. From the Kaplan scaling laws paper: > We have observed consistent scalings of language model log-likelihood loss with non-embedding parameter count N, dataset size D, and optimized training computation Cmin, as encapsulated in Equations (1.5) and (1.6). Conversely, we find very weak dependence on…
This is why, because this completely misses the point. Every power law has diminishing marginal returns, that is by construction. The point is that they diminish predictably and smoothly across seven orders of magnitude of compute. There is no plateau anywhere in the observed range, and in fact larger models more sample efficient than expected.
You can't just cherry pick one sentence containing the phrase "diminishing returns" and ignore the actual paper. The whole reason the labs spent the GDP of a small country on GPUs is that the curve is boring and reliable. No one just was like "oops we missed that line in the Kaplan paper, shit".
Lots of arguments as to why things may slow down or hit fundamental limits requiring a paradigm shift. But at several points in the past 5 years many major players have assumed this to be the case and have been continually proven wrong. Not to mention the plethora of socioeconomic horrors coming out of this Pandora's box, which I believe to be a large colocated B300 compute cluster somewhere in the midwest.
A point that would have been consistent with Kaplan was them finding alpha ~ 0.05 on compute means halving the reducible loss is a million times more expensive. But this was invalidated by Chinchilla who found it to be ~0.36 (100x more compute, not 1million). And loss has a non-linear relationship with capabilities, but roughly halving pretraining loss is ~doubling the METR task horizon. We also haven't factored in something incredibly important which is: algorithmic / architectural / training pipeline gains. These are substantial too!