Earlier quoted context omitted.
> We cannot add more compute to a given compute budget C without increasing data D to maintain the relationship. > We must either (1) discover new architectures with different scaling laws, and/or (2) compute new synthetic data that can contribute to learning (akin to dreams). Of course we can, this is a non issue. See e.g. AlphaZero [0] that's 8 years old at this point, and any modern RL training using synthetic dat…
AlphaZero trained itself through chess games that it played with itself. Chess positions have something very close to an objective truth about the evaluation, the rules are clear and bounded. Winning is measurable. How do you achieve this for a language model? Yes, distillation is a thing but that is more about compression and filtering. Distillation does not produce new data in the same way that chess games produce…
The Bitter Lesson Is Misunderstood
21–30 of 259 posts
Re: The Bitter Lesson Is Misunderstood
#22Earlier quoted context omitted.
Very true. Living cells are ~4-5 orders of magnitude more functional-information-dense than the most advanced chips, and there is a lot more living mass than advanced chips. But the networking potential of digital compute is a fundamentally different paradigm than living systems. The human brain is constrained in size by the width of the female pelvis. So while it's expensive, we can trade scope-constrained robustnes…
> The human brain is constrained in size by the width of the female pelvis. Well, it _was_ until recently.
Re: The Bitter Lesson Is Misunderstood
#23The problem I am facing in my domain is that all of the data is human generated and riddled with human errors. I am not talking about typos in phone numbers, but rather fundamental errors in critical thinking, reasoning, semantic and pragmatic oversights, etc. all in long-form unstructured text. It's very much an LLM-domain problem, but converging on the existing data is like trying to converge on noise. The opportun…
This is my current drum I bang on when an uninformed stakeholder tries shoving LLMs blindly down everyone’s throats: it’s the data, stupid . Current data aggregates outside of industries wholly dependent on it (so anyone not in web advertising, GIS, or intelligence) are garbage , riddled with errors and in awful structures that are opaque to LLMs. For your AI strategy to have any chance of success, your data has to b…
Re: The Bitter Lesson Is Misunderstood
#24Re: The Bitter Lesson Is Misunderstood
#25Re: The Bitter Lesson Is Misunderstood
#26Earlier quoted context omitted.
> We cannot add more compute to a given compute budget C without increasing data D to maintain the relationship. > We must either (1) discover new architectures with different scaling laws, and/or (2) compute new synthetic data that can contribute to learning (akin to dreams). Of course we can, this is a non issue. See e.g. AlphaZero [0] that's 8 years old at this point, and any modern RL training using synthetic dat…
AlphaZero trained itself through chess games that it played with itself. Chess positions have something very close to an objective truth about the evaluation, the rules are clear and bounded. Winning is measurable. How do you achieve this for a language model? Yes, distillation is a thing but that is more about compression and filtering. Distillation does not produce new data in the same way that chess games produce…
But generally the idea is that it's, you need some notion of reward, verifiers etc.
Works really well for maths, algorithms, amd many things actually.
See also this very short essay/introduction: https://www.jasonwei.net/blog/asymmetry-of-verification-and-...
That's why we have IMO gold level models now, and I'm pretty confident we'll have superhuman mathematics, algorithmic etc models before long.
Now domains which are very hard to verify - think e.g. theoretical physics etc - that's another story.
Re: The Bitter Lesson Is Misunderstood
#27Re: The Bitter Lesson Is Misunderstood
#28No, the paths forward are: better design, training, feeding in more video, audio, and general data from the outside world. The web is just a small part of our experience. What about apps, webcam streams, radio from all over the world in its many forms, OTA TV, interacting with streaming content via remote, playing every video game, playing board games with humans, feeds and data from robots LLMs control, watching everyone via their phones and computers, car cameras, security footage and CCTV, live weather and atmospheric data, cable television, stereoscopic data, ViewMaster reels, realtime electrical input from various types of brains while interacting with their attached creatures, touch and smell, understanding birth, growth, disease, death, and all facets of life as an observer, observing those as a subject, expanding to other worlds, solar systems, galaxies, etc., affecting time and space, search and communication with a universal creator, and finally understanding birth and death of the universe.
Re: The Bitter Lesson Is Misunderstood
#29I don't understand why we need more data for training. Assuming we've already digitized every book, magazine, research paper, newspaper, and other forms of media, why do we need this "second internet?" Legal issues aside, don't we already have the totality of human knowledge available to us for training?
Re: The Bitter Lesson Is Misunderstood
#30> The path forward: data alchemists (high-variance, 300% lottery ticket) or model architects (20-30% steady gains) No, the paths forward are: better design, training, feeding in more video, audio, and general data from the outside world. The web is just a small part of our experience. What about apps, webcam streams, radio from all over the world in its many forms, OTA TV, interacting with streaming content via remot…