Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

191–200 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#191

> There is no second internet I don't know about that. LLMs have been trained mostly on text. If you add photos, audio and videos, and later even 3D games, or 3D videos, you get massively more data than the old plain text. Maybe by many orders of magnitude. And this is certainly that can improve cognition in general. Getting to AGI without audio and video, and 3D perception seems like a non-starter. And even if we th…

Also, even if we lacked the data to proceed with Chinchilla-optimal scaling that wouldn't be the same as being unable to proceed with scaling, it would just require larger models and more flops than we would prefer.

Re: The Bitter Lesson Is Misunderstood

#192

Earlier quoted context omitted.

We think this because ten years ago we were all having our minds blown by DeepMind's game playing achievements and videos of dancing robots and thought this meant blue collar work would be solved imminently. But most of these solutions were more crude than they let on, and you wouldn't really know unless you were working in AI already. Watch John Carmack's recent talk at Upper Bound if you want him to see him destroy…

Thank you for this update. I vividly remember a few years ago the excitement of John Carmack announcing he was retreating into his cave to do some deep work on AGI, pushing the boundaries of the current AI research. I truly appreciate Carmack's intellectual honesty now at announcing "yeah, no, LLMs are not the way to go to recreate anything remotely close to human intelligence.". In fact, and I quote him, "we do not…

I don't think that quote from Carmack represents some deeply considered conclusion. He started off his efforts with embodiment. He either never considered LLMs a path towards AGI, or thought he didn't personally have anything to contribute to LLMs (he talked about it early on in his journey but I don't remember the specifics). He didn't spend a year investigating LLMs and then decide that they weren't the path to AGI. The point is that he has no special insight regarding LLMs relationship to AGI and its misleading to imply that his current effort towards building AGI that eschew LLMs is an expert opinion.

Re: The Bitter Lesson Is Misunderstood

#193

Earlier quoted context omitted.

> they've trained on Somewhat apples and oranges given billions of years of evolution behind that human. GPT-5 started off as a blank slate.

[flagged]

Neural precursor cells literally move themselves from where they first differentiate to their final location to ensure specific neural structures and information dynamics in the developed brain. It's not declarative memory, but its a memory of the neural architecture etched out over evolutionary time.

Re: The Bitter Lesson Is Misunderstood

#194
post #183

Earlier quoted context omitted.

This comparison is absolute nonsense. "How could a telescope see saturn, human eyes have billions of years of evolution behind them, and we only made telescopes a few hundred years ago, so they should be much weaker than eyes" "How can StockFish play chess better than a human, the human brain has had billions of years of evolution" Evolution is random, slow, and does not mean we arrive at even a local optima.

They're not saying that LLMs should be better than smart teenagers; they're saying that smart teenagers can solve some problems without needing massive amounts of data, so apparently those problems are technically solvable without those amounts of data.

Yes. It is astonishing that LLMs can solve problems that only a handful of very smart teenagers can solve, but LLMs do it by consuming a million times as much content as those teenagers. Running out of data is not a reason for despair.

Also consider that during training LLMs spend much less time on processing, say, TAOCP (Knuth), or SICP (Abelson, Sussman, and Sussman), or Probability Theory (Jaynes) than on the entirety that is r/Frugal.

20 thick books turn a smart teenager into a graduate with a MSc. That's what, 10 million tokens?

When we read difficult, important texts, we reflect on them, make exercises, discuss them, etc. We don't know how to make an LLM do that in a way that improves it. Yet.

Re: The Bitter Lesson Is Misunderstood

#195

Earlier quoted context omitted.

Yes this is new synthetic data which did not exist before. I encourage you to read the link.

I think we're talking past each other, I'll try once more. Suppose you train an LLM on a very small corpus of data, such as all the content of the library of congress. Then you have that LLM author new works. Then you train a new LLM on the original corpus plus this new material. Do you really think you've addressed the core issue in the SP? Can more parameters be meaningfully trained even if you add more GPU? To me,…

When it comes to logical reasoning, the difficulty isn't about having enough new information, but about ensuring the LLMs capture the right information. The problem LLMs have with learning logical reasoning from standard training is that they learn spurious relationships between the context and the next token, undermining its ability to learn fully general logical reasoning. Synthetic data helps because spurious associations are undermined by the randomness inherent in the synthetic data, forcing the model to find the right generic reasoning steps.

Re: The Bitter Lesson Is Misunderstood

#196

Earlier quoted context omitted.

Yes this is new synthetic data which did not exist before. I encourage you to read the link.

I think we're talking past each other, I'll try once more. Suppose you train an LLM on a very small corpus of data, such as all the content of the library of congress. Then you have that LLM author new works. Then you train a new LLM on the original corpus plus this new material. Do you really think you've addressed the core issue in the SP? Can more parameters be meaningfully trained even if you add more GPU? To me,…

Yes if you have some way to verify the quality of the new works and you only include the high quality works in the new LLM's training set.

Re: The Bitter Lesson Is Misunderstood

#197

Earlier quoted context omitted.

AlphaZero trained itself through chess games that it played with itself. Chess positions have something very close to an objective truth about the evaluation, the rules are clear and bounded. Winning is measurable. How do you achieve this for a language model? Yes, distillation is a thing but that is more about compression and filtering. Distillation does not produce new data in the same way that chess games produce…

Simple, you just need to turn language into a game. You make models talk to each other, create puzzles for each other's to solve, ask each other to make cases and evaluate how well they were made. Will some of it look like ramblings of pre-scientific philosophers? (or modern ones because philosophy never progressed after science left it in the dust) Sure! But human culture was once there too. And we pulled ourselves…

>And we pulled ourselves out of this nonsense by the bootstraps.

Human progress was promoted by having to interact with a physical world that anchored our ramblings and gave us a reward function for coherence and cooperation. LLMs would need some analogous anchoring for it to progress beyond incoherent babble.

Re: The Bitter Lesson Is Misunderstood

#198

Earlier quoted context omitted.

"We have to learn the bitter lesson that building in how we think we think does not work in the long run. The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough…

What if the meta bitter lesson is that data scaling is just a more extreme form of the human-centric approach of building knowledge into agents? After all, we're telling the model what to say, think and how to behave. A true general method wouldn't rely on humans at all! Human data would be worthless beyond bootstrapping!

Another meta bitter lesson: we don't understand ourselves well enough to define and build something that thinks as we do.

Re: The Bitter Lesson Is Misunderstood

#199

I really enjoyed reading this article as I found its content extremely insightful, but I fear I must whine for far too long about something entirely minor. As someone that didn't go to expensive maths club, the way people who did, talk about maths is disgraceful imho. Consider the equasion in this article: (C ~ 6 N⋅D) I can look up the symbol for "roughly equals", that was super cool and is a great part of curiousity…

> I just don't understand what's wrong with 6⋅N⋅D or 6ND

I think most people that read this would be confused and try to find out why there are some undisclosed vectorial operations applied to what looks like scalar numbers.

And yeah, mathematical notation is ugly and confusing. But the fix is not as simple as you think it is.

Apropos CQRS, it's a marketing name. It's hard to understand on purpose. Actual CS-made names tend to be easier.

Re: The Bitter Lesson Is Misunderstood

#200

Earlier quoted context omitted.

The problem is not the robot loading the diswasher, it is the dishwasher. The dishwasher (and general kitchen electronics) industry has not innovated in a long time. My prediction is a new player will come in who vertically integrates these currently disjoint industries and product. The tableware used should be compatible with the dishwasher, the packaging of my groceries should be compatible with the cooking system.…

This exists already in the form of "ready meals" a.k.a. TV dinners. Fast food shops are already substantially mechanised; huge effots have been made to robotize cooking, but people are still cheaper to hire. It's still nowhere near the quality of home-cooked food.

Yes, there are a lot of garbage microwave food offerings, especially popular with the US population. As a European I'm talking about quality food made with an automated process and end-to-end automation, including ingredient procurement and cleanup.

Not in competition with trash food but with proper food and local ingredients.

Post reply on HN