Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

61–70 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#62
post #60

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

Not a robotics guy, but to extent that the same fundamentals hold— I think it's a degrees of freedom question. Given the (relatively) low conditional entropy of natural language, there aren't actually that many degrees of (true) freedom. On the other hand, in the real world, there are massively more degrees of freedom both in general (3 dimensions, 6 degrees of movement per joint, M joints, continuous vs. discrete sp…

Also not a robotics guy, but that all sounds right to me...

What I do have deep experience in is market abstractions and jobs to be done theory. There are so many ways to describe intent, and it's extremely hard to describe intent precisely. So in addition to all the dimensions you brought up that relate to physical space, there is also the hard problem of mapping user intent to action with minimal "error", especially since the errors can have big consequences in the physical world. In other words, the "intent space" also has many dimensions to it, far beyond what LLMs can currently handle.

On one end of the spectrum of consequences is the robot loads my dishwasher such that there is too much overlap and a bunch of the dishes don't get cleaned (what I really want is for the dishes to be clean, not for the dishes to be in the dishwasher), and on the other end we get the robot that overpowers humanity and turns the universe into paperclips.

So maybe we have to master LLMs and probably a whole other paradigm before robots can really be general purpose and useful.

Re: The Bitter Lesson Is Misunderstood

#63
In any field where there is a creative element, progress comes in fits and starts that are difficult to predict in advance. No one can accurately predict when we'll get the cure for cancer, for example, in spite of people working on it.

But that isn't how investors operate. They want to know what they will get in exchange for giving a company a billion dollars. If you're running an AI business, you need to set expectations. How do you do that? Go do the thing you know you can do on a schedule, like standing up a new GPU data center.

I don't think the bitter lesson is misunderstood in quite the way the author describes. I think most are well aware we're approaching the data wall within a couple years. However, if you're not in academia you're not trying to solve that problem; you're trying to get your bag before it happens.

That may sound a little flip, but this is yet another incarnation of the hungry beast: https://stvp.stanford.edu/clips/the-hungry-beast-and-the-ugl...

Re: The Bitter Lesson Is Misunderstood

#64

I really enjoyed reading this article as I found its content extremely insightful, but I fear I must whine for far too long about something entirely minor. As someone that didn't go to expensive maths club, the way people who did, talk about maths is disgraceful imho. Consider the equasion in this article: (C ~ 6 N⋅D) I can look up the symbol for "roughly equals", that was super cool and is a great part of curiousity…

What diamond symbol?

Re: The Bitter Lesson Is Misunderstood

#65

Earlier quoted context omitted.

Very true. Living cells are ~4-5 orders of magnitude more functional-information-dense than the most advanced chips, and there is a lot more living mass than advanced chips. But the networking potential of digital compute is a fundamentally different paradigm than living systems. The human brain is constrained in size by the width of the female pelvis. So while it's expensive, we can trade scope-constrained robustnes…

> Living cells are ~4-5 orders of magnitude more functional-information-dense than the most advanced chips, and there is a lot more living mass than advanced chips. I believe you but I would love to know where this number came from just so I can read more about it

It's napkin math so take it with a pinch of salt, but I am calculating the information stored in genome, assuming 2 bits per base pair, reducing to estimated 88% coding fraction to get the functional bits, and then dividing by cell volume. Did this for a few different types of cells and then averaged the result to around 1–10 Mbit/μm³

# If there are any bioinformaticians around please come eviscerate or confirm this calc #

Then compared it to TSMC 2nm research macro of (38.1 Mbit/mm^2) normalized to cell scale: 0.00019 Mbit/μm³

Living Cells: 1–10 Mbit/μm³

Current best chips: 0.00019 Mbit/μm³

https://research.tsmc.com/page/memory/4.html

Re: The Bitter Lesson Is Misunderstood

#66
post #36
post #24

I don't understand why we need more data for training. Assuming we've already digitized every book, magazine, research paper, newspaper, and other forms of media, why do we need this "second internet?" Legal issues aside, don't we already have the totality of human knowledge available to us for training?

The point is that current methods are unable to get more than the current state-of-the-art models' degree of intelligence out of training on the totality of human knowledge. Previously, the amount of compute needed to process that much data was a limit, but not anymore. So now, in order to progress further, we either have to improve the methods, or synthetically generate more training data, or both.

What does synthetic training data actually mean? Just saying the same things in different ways? It seems like we're training in a way that's just not sustainable.

Re: The Bitter Lesson Is Misunderstood

#67
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

What do you mean about CLIP?

Re: The Bitter Lesson Is Misunderstood

#68

I really enjoyed reading this article as I found its content extremely insightful, but I fear I must whine for far too long about something entirely minor. As someone that didn't go to expensive maths club, the way people who did, talk about maths is disgraceful imho. Consider the equasion in this article: (C ~ 6 N⋅D) I can look up the symbol for "roughly equals", that was super cool and is a great part of curiousity…

I assume they use N⋅D rather than ND to make it explicit these are 2 different variables. That's not necessary for 6N because variable names don't start with a number by convention.

Re: The Bitter Lesson Is Misunderstood

#69
Has HRM really dramatically changed the landscape? My read of the paper thus far is that it is an impressive result, but there have been a few of those in the past that have fizzled out, so I'm still in wait-and-see mode

Re: The Bitter Lesson Is Misunderstood

#70

Earlier quoted context omitted.

Very true. Living cells are ~4-5 orders of magnitude more functional-information-dense than the most advanced chips, and there is a lot more living mass than advanced chips. But the networking potential of digital compute is a fundamentally different paradigm than living systems. The human brain is constrained in size by the width of the female pelvis. So while it's expensive, we can trade scope-constrained robustnes…

> The human brain is constrained in size by the width of the female pelvis. I think this is an old belief that isn't supported by modern research.

Thanks for pointing this out. I wasn't aware, and so I just dug into it.

From what I can tell, science used to point to this as the only/primary limit to human-brain size, but more recently the picture seems a lot less clear, with some indications that pelvis size doesn't place as hard of a constraint as we thought and there are other constraints such as metabolic (how many calories the mother can sustain during pregnancy and lactation).

So overall I'd say you are technically correct, even though this doesn't really materially change the point I was making; which is that the size of the human brain is constrained in ways that the size of data centers are not.

Post reply on HN