Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

151–160 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#151

Earlier quoted context omitted.

> So it makes me wonder, is embodiment (advanced robotics) 1000x harder than LLMs from an information processing perspective? Essentially, yes, but I would go further in saying that embodiment is harder than intelligence in and of itself. I would argue that intelligence is a very simple and primitive mechanism compared to the evolved animal body, and the effectiveness of our own intelligence is circumstantial. We man…

It took about the same amount of time to evolve human-level intelligence as human-level mobility. Pretty much no other animal walks on two legs...

Birds? Bears whose front paws got injured? https://youtu.be/kcIkQaLJ9r8

Re: The Bitter Lesson Is Misunderstood

#152
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

The problem is not the robot loading the diswasher, it is the dishwasher. The dishwasher (and general kitchen electronics) industry has not innovated in a long time.

My prediction is a new player will come in who vertically integrates these currently disjoint industries and product. The tableware used should be compatible with the dishwasher, the packaging of my groceries should be compatible with the cooking system. Like a mini-factory.

But current vendors have no financial incentive to do so, because if you take a step back the whole notion of putting one room of your apartment full with random electronics just to cook a meal once in a blue moon is deeply inefficient. End-to-end food automation is coming to the restaurant business, and I hope it pushes prices of meals so far down that having a dedicated room for a kitchen in the apartment is simply not worth it.

That's the "utopia" version of things.

In reality, we see prices for fast food (the most automated food business) going up while quality is going down. Does it make the established players more vulnerable to disruption? I think so.

Re: The Bitter Lesson Is Misunderstood

#153

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

> they've trained on Somewhat apples and oranges given billions of years of evolution behind that human. GPT-5 started off as a blank slate.

Re: The Bitter Lesson Is Misunderstood

#154
post #134

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Now consider that the genius cannot physically interact with the world or the people therein, and uses her eyes only for reading text.

Also the geniuses get beaten with a stick if they don't memorize and perfectly reproduce the text they've read.

Re: The Bitter Lesson Is Misunderstood

#155
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

What do you mean by "verifiable rewards"?

Do you mean challenges for which the answer is known?

Re: The Bitter Lesson Is Misunderstood

#156
post #60

Earlier quoted context omitted.

Not a robotics guy, but to extent that the same fundamentals hold— I think it's a degrees of freedom question. Given the (relatively) low conditional entropy of natural language, there aren't actually that many degrees of (true) freedom. On the other hand, in the real world, there are massively more degrees of freedom both in general (3 dimensions, 6 degrees of movement per joint, M joints, continuous vs. discrete sp…

This understates the complexity of the problem. I have built a career modeling/learning entity behavior in the physical world at scale. Language is almost a trivial case by comparison. Even the existence of most relationships in the physical world can only be inferred, never mind dimensionality. The correlations are often weak unless you are able to work with data sets that far exceed the entire corpus of all human t…

In other words, self driving cars and robot vacuum cleaners cannot exist. Hmm.

Re: The Bitter Lesson Is Misunderstood

#157
post #100
post #24

I don't understand why we need more data for training. Assuming we've already digitized every book, magazine, research paper, newspaper, and other forms of media, why do we need this "second internet?" Legal issues aside, don't we already have the totality of human knowledge available to us for training?

The totality of human knowledge is a rounding error to what’s needed for AGI

That is only true if your path to AGI is to take models similar to current models, and feed them with tons of data.

Advances in architecture and training protocols can and will easily dwarf "more data". I think that is quite obvious from the fact that humans learn to be quite intelligent using only a fraction of the data available to current LLMs. Our advantage is a very good pre-baked model, and feedback-based training.

Re: The Bitter Lesson Is Misunderstood

#158

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

We think this because ten years ago we were all having our minds blown by DeepMind's game playing achievements and videos of dancing robots and thought this meant blue collar work would be solved imminently. But most of these solutions were more crude than they let on, and you wouldn't really know unless you were working in AI already. Watch John Carmack's recent talk at Upper Bound if you want him to see him destroy…

> But most of these solutions were more crude than they let on, and you wouldn't really know unless you were working in AI already.

Same with LLMs. Despite having seen this play out before, and being aware of this, people are falling for it again.

Re: The Bitter Lesson Is Misunderstood

#159

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

The big robot AI issue is: no data! There is a lot of high quality text from diverse domains, there's a lot of audio or images or videos around. The largest robotics datasets are absolutely pathetic in size compared to that. We didn't collect or stockpile the right data in advance. Embodiment may be hard by itself, but doing embodiment in this data-barren wasteland is living hell. So you throw everything but the kitc…

> And all of that combined only gets you to "meh" real world performance - slow, flaky, fairly brittle, and on relatively narrow tasks. Often good enough for an impressive demo, but not good enough to replace human workers yet.

Sounds like LLMs to me.

Re: The Bitter Lesson Is Misunderstood

#160

Earlier quoted context omitted.

The big robot AI issue is: no data! There is a lot of high quality text from diverse domains, there's a lot of audio or images or videos around. The largest robotics datasets are absolutely pathetic in size compared to that. We didn't collect or stockpile the right data in advance. Embodiment may be hard by itself, but doing embodiment in this data-barren wasteland is living hell. So you throw everything but the kitc…

> And all of that combined only gets you to "meh" real world performance - slow, flaky, fairly brittle, and on relatively narrow tasks. Often good enough for an impressive demo, but not good enough to replace human workers yet. Sounds like LLMs to me.

It's like GPT-3.5 - a proof-of-concept tech demo more than a product.

I don't think further improvements are impossible, not at all. They're just hard to get at.

Post reply on HN