Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

201–210 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#202
Humans require a _lot_ less training data to become, for instance, fluent in English. If a given AI algorithm needs to be trained on the entire Internet to accomplish the same, then it seems safe to assume that the data has not really been "mined out".

Generating more training data from the same original data should not be fundamentally problematic in that sense.

Re: The Bitter Lesson Is Misunderstood

#203

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Maybe human brains are constantly generating (and training on) massive amounts of synthetic data and that is how they get so smart?

Re: The Bitter Lesson Is Misunderstood

#204
post #68

Earlier quoted context omitted.

I assume they use N⋅D rather than ND to make it explicit these are 2 different variables. That's not necessary for 6N because variable names don't start with a number by convention.

Its good we all learned this convention. Thanks for teaching to it to me though. To clarify, if it read: C ~ X N⋅D you'd be as confused as me? Its because its a number it has special implied mechanics where we can skip operators because its "obvious".

Well no actually it'd still clear to me that they mean the the multiplication of 3 different variables X, N, and D.

I don't think of it as eliding obvious operators. Rather in mathematics juxtaposition is used as an operator to represent multiplication. You would never elide an addition operator.

So X next to D still means multiplication as long as you can tell that X and D are separate entities.

I would wonder why they switched conventions in the middle of an expression though.

Re: The Bitter Lesson Is Misunderstood

#205

Humans require a _lot_ less training data to become, for instance, fluent in English. If a given AI algorithm needs to be trained on the entire Internet to accomplish the same, then it seems safe to assume that the data has not really been "mined out". Generating more training data from the same original data should not be fundamentally problematic in that sense.

It only seems that way because much of the data that humans use is not in a format that computers would understand. A toddler learning to talk is engaging their full body.

Re: The Bitter Lesson Is Misunderstood

#206

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Maybe human brains are constantly generating (and training on) massive amounts of synthetic data and that is how they get so smart?

This sentence really struck me in a particular way. Very interesting. It does seem like thoughts/stream of consciousness is just your brain generating random tokens to itself and learning from it lol.

Re: The Bitter Lesson Is Misunderstood

#207

Earlier quoted context omitted.

> So it makes me wonder, is embodiment (advanced robotics) 1000x harder than LLMs from an information processing perspective? Essentially, yes, but I would go further in saying that embodiment is harder than intelligence in and of itself. I would argue that intelligence is a very simple and primitive mechanism compared to the evolved animal body, and the effectiveness of our own intelligence is circumstantial. We man…

> We manage to dominate the world mainly by using brute force to simplify our environment and then maintaining and building systems on top of that simplified environment. If we didn't have the proper tools to selectively ablate our environment's complexity… This is very interesting and I feel there is a lot to unpack here. Could you elaborate on this theory with a few more paragraphs (or books / blogs that elucidate…

> Why does our intelligence only operate on simplified forms?

Part of the issue with discussing this is that our understanding of complexity is subjective and adapted to our own capabilities. But the gist of it is that the difficulty of modelling and predicting the behavior of a system scales very sharply with its complexity. At the end of the scale, chaotic systems are basically unintelligible. Since modelling is the bread and butter of intelligence, any action that makes the environment more predictable has outsized utility. Someone else gave pretty good examples, but I think it's generally obvious when you observe how "symbolic-smart" people think (engineers, rationalists, autistic people, etc.) They try to remove as many uncontrolled sources of complexity as possible. And they will rage against those that cannot be removed, if they don't flat out pretend they don't exist. Because in order to realize their goals, they need to prove things about these systems, and it doesn't take much before that becomes intractable.

One example of a system that I suspect to be intractable is human society itself. It is made out of intelligent entities, but as a whole I don't think it is intelligent, or that it has any overarching intent. It is insanely complex, however, and our attempts to model its behavior do not exactly have a good record. We can certainly model what would happen if everybody did this or that (aka a simpler humanity), but everybody doesn't do this and that, so that's moot. I think it's an illuminating example of the limitations of symbolic intelligence: we can create technology (simple), but we have absolutely no idea what the long term consequences are (complex). Even when we do, we can't do anything about it. The system is too strong, it's like trying to flatten the tides.

> To me the natural test is how hard it is for each one to create the other.

I don't think so. We already observe that humans, the quintessential symbolic intelligences, have created symbolic intelligence before embodied intelligence. In and of itself, that's a compelling data point that embodied is harder. And it appears likely that if LLMs were tasked to create symbolic intelligences, even assuming no access to previous research, they would recreate themselves faster than they would create embodied intelligences. Possibly they would do so faster than evolution, but I don't see why that matters, if they also happen to recreate symbolic intelligence even faster than that. In other words, if symbolic is harder... how the hell did we get there so quick? You see what I mean? It doesn't add up.

On a related note, I'd like to point out an additional subtlety regarding intelligence. Intelligence (unlike, say, evolution) has goals and it creates things to further these goals. So you create a new synthetic life. That's cool. But do you control it? Does it realize your intent? That's the hard part. That's the chief limitation of intelligence. Creating stuff that is provably aligned with your goals. If you don't care what happens, sure, you can copy evolution, you can copy other methods, you can create literally anything, perhaps very quickly, but that's... not smart. If we create synthetic life that eats the universe, that's not an achievement, that's a failure mode. (And if it faithfully realizes our intent then yeah I'm impressed.)

Re: The Bitter Lesson Is Misunderstood

#208

Earlier quoted context omitted.

Thank you for this update. I vividly remember a few years ago the excitement of John Carmack announcing he was retreating into his cave to do some deep work on AGI, pushing the boundaries of the current AI research. I truly appreciate Carmack's intellectual honesty now at announcing "yeah, no, LLMs are not the way to go to recreate anything remotely close to human intelligence.". In fact, and I quote him, "we do not…

I don't think that quote from Carmack represents some deeply considered conclusion. He started off his efforts with embodiment. He either never considered LLMs a path towards AGI, or thought he didn't personally have anything to contribute to LLMs (he talked about it early on in his journey but I don't remember the specifics). He didn't spend a year investigating LLMs and then decide that they weren't the path to AGI…

Yes, I meant to say that, for Carmack, no type of modern AI research has figured out the path to actual general intelligence. I just didn't want to use the meaningless "AI" buzzword, and these days all the focus and money is on large language models, especially when talking about the end goal of AGI.

Re: The Bitter Lesson Is Misunderstood

#209

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Maybe human brains are constantly generating (and training on) massive amounts of synthetic data and that is how they get so smart?

What experiment could be run to test this hypothesis?

Re: The Bitter Lesson Is Misunderstood

#210
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

Since you seem to know your stuff, why do LLMs need so much data anyway? Humans don't. Why can't we make models aware of their own uncertainty, e.g. feeding the variance of the next token distribution back into the model, as a foundation to guide their own learning. Maybe with that kind of signal, LLMs could develop 'curiosity' and 'rigorousness' and seek out the data that best refines them themselves. Let the AI make and test its own hypotheses, using formal mathematical systems, during training.
Post reply on HN