Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

131–140 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#131
Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs.

There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just a problem for the current paradigm.

Re: The Bitter Lesson Is Misunderstood

#132
post #68

I really enjoyed reading this article as I found its content extremely insightful, but I fear I must whine for far too long about something entirely minor. As someone that didn't go to expensive maths club, the way people who did, talk about maths is disgraceful imho. Consider the equasion in this article: (C ~ 6 N⋅D) I can look up the symbol for "roughly equals", that was super cool and is a great part of curiousity…

I assume they use N⋅D rather than ND to make it explicit these are 2 different variables. That's not necessary for 6N because variable names don't start with a number by convention.

Its good we all learned this convention. Thanks for teaching to it to me though.

To clarify, if it read:

C ~ X N⋅D

you'd be as confused as me? Its because its a number it has special implied mechanics where we can skip operators because its "obvious".

Re: The Bitter Lesson Is Misunderstood

#133
post #100
post #24

I don't understand why we need more data for training. Assuming we've already digitized every book, magazine, research paper, newspaper, and other forms of media, why do we need this "second internet?" Legal issues aside, don't we already have the totality of human knowledge available to us for training?

The totality of human knowledge is a rounding error to what’s needed for AGI

What makes you think that? Especially given that fact that GI (without the 'A') is evidently very much possible with only a tiny fraction of the "totality of human knowledge".

Re: The Bitter Lesson Is Misunderstood

#134

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Now consider that the genius cannot physically interact with the world or the people therein, and uses her eyes only for reading text.

Re: The Bitter Lesson Is Misunderstood

#136
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

[deleted]

Re: The Bitter Lesson Is Misunderstood

#138
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

> there are some cool labs building foundational robotics models, but they're maybe ~5 years behind LMs today

Wouldn't the Bitter Lesson be to invest in those models over trying to be clever about ekeing out a little more oomph from today's language models (and langue-based data)?

Re: The Bitter Lesson Is Misunderstood

#139
post #134

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Now consider that the genius cannot physically interact with the world or the people therein, and uses her eyes only for reading text.

Yes - we train only on a subset of human communication, the one using written symbols (even voice has much much more depth to it), but human brains train on the actual physical world.

Human students who only learned some new words but have not (yet) even began to really comprehend a subject will just throw around random words and sentences that sound great but have no basis in reality too.

For the same sentence, for example, "We need to open a new factory in country XY", the internal model lighting up inside the brain of someone who has actually participated when this was done previously will be much deeper and larger than that of someone who only heard about it in their course work. That same depth is zero for an LLM, which only knows the relations between words and has no representation of the world. Words alone cannot even begin to represent what the model created from the real-world sensors' data, which on top of the direct input is also based on many times compounded and already-internalized prior models (nobody establishes that new factory as a newly born baby with a fresh neural net, actually, even the newly born has inherited instincts that are all based on accumulated real world experiences, including the complex very structure of the brain).

Somewhat similarly, situations reported in comments like this one (client or manager vastly underestimating the effort required to do something): https://news.ycombinator.com/item?id=45123810 The internal model for a task of those far removed from actually doing it is very small compared to the internal models of those doing the work, so trying to gauge required effort falls short spectacularly if they also don't have the awareness.

Re: The Bitter Lesson Is Misunderstood

#140
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

We think this because ten years ago we were all having our minds blown by DeepMind's game playing achievements and videos of dancing robots and thought this meant blue collar work would be solved imminently.

But most of these solutions were more crude than they let on, and you wouldn't really know unless you were working in AI already.

Watch John Carmack's recent talk at Upper Bound if you want him to see him destroy like a trillion dollars worth of AI hype.

https://m.youtube.com/watch?v=rQ-An5bhkrs&t=11303s&pp=2AGnWJ...

Spoiler: we're nowhere close to AGI

Post reply on HN