Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

171–180 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#171
The AI companies won't run out of data to train on. Almost every user interaction is a significant source of data. Chains of interactions are even more significant, especially the longer and more sophisticated they are. Yesterday I was given A/B tests from both GPT5-Thinking and Gemini 2.5 Pro, something neither of then had done before. OpenAI also just acquired Statsig for $1.1 billion. Statsig does A/B testing and other analytics.

The data scrapped from the Internet and scanned books served its purpose: it bootstrapped something that we all love talking to and discussing ANYTHING with. That's the new source of data and intelligence.

Re: The Bitter Lesson Is Misunderstood

#172

Earlier quoted context omitted.

> they've trained on Somewhat apples and oranges given billions of years of evolution behind that human. GPT-5 started off as a blank slate.

This comparison is absolute nonsense. "How could a telescope see saturn, human eyes have billions of years of evolution behind them, and we only made telescopes a few hundred years ago, so they should be much weaker than eyes" "How can StockFish play chess better than a human, the human brain has had billions of years of evolution" Evolution is random, slow, and does not mean we arrive at even a local optima.

What comparison? I was arguing against a comparison.

Re: The Bitter Lesson Is Misunderstood

#173

Earlier quoted context omitted.

This understates the complexity of the problem. I have built a career modeling/learning entity behavior in the physical world at scale. Language is almost a trivial case by comparison. Even the existence of most relationships in the physical world can only be inferred, never mind dimensionality. The correlations are often weak unless you are able to work with data sets that far exceed the entire corpus of all human t…

In other words, self driving cars and robot vacuum cleaners cannot exist. Hmm.

LOL. Both of those are very limited and work in 2D spaces in highly constrained environments especially designed for them.

Re: The Bitter Lesson Is Misunderstood

#174

In the bitter lesson essay [0], the word "data" is not mentioned a single time. The author fundamentally misunderstands the bitter lesson. [0] https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson...

"We have to learn the bitter lesson that building in how we think we think does not work in the long run. The bitter lesson is based on the historical observations that 1) AI researchers have often tried to build knowledge into their agents, 2) this always helps in the short term, and is personally satisfying to the researcher, but 3) in the long run it plateaus and even inhibits further progress, and 4) breakthrough…

What if the meta bitter lesson is that data scaling is just a more extreme form of the human-centric approach of building knowledge into agents? After all, we're telling the model what to say, think and how to behave.

A true general method wouldn't rely on humans at all! Human data would be worthless beyond bootstrapping!

Re: The Bitter Lesson Is Misunderstood

#175

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

> So it makes me wonder, is embodiment (advanced robotics) 1000x harder than LLMs from an information processing perspective? Essentially, yes, but I would go further in saying that embodiment is harder than intelligence in and of itself. I would argue that intelligence is a very simple and primitive mechanism compared to the evolved animal body, and the effectiveness of our own intelligence is circumstantial. We man…

> We manage to dominate the world mainly by using brute force to simplify our environment and then maintaining and building systems on top of that simplified environment. If we didn't have the proper tools to selectively ablate our environment's complexity…

This is very interesting and I feel there is a lot to unpack here. Could you elaborate on this theory with a few more paragraphs (or books / blogs that elucidate this)? In what ways do we use brute force to simplify the environment, and are there not ways in which we use highly sophisticated leveraged methods to simplify our environment tools? What proper tools allow us to selectively ablate complexity? Why does our intelligence only operate on simplified forms?

Also, what would convince you that symbolic intelligence is actually “harder” than embodied intelligence? To me the natural test is how hard it is for each one to create the other. We know it took a few billion years to go from embodied intelligence (ie organisms that can undergo evolution, with enough diversity to survive nearly any conditions on Earth) to sophisticated symbolic intelligence. What if it turns out that within 100 years, symbolic intelligence (contained in LLM like systems) could produce the insights to eg create new synthetic life from scratch that was capable of undergoing self-sustained evolution in diverse and chaotic environments? Would this convince you that actually symbolic intelligence is the harder problem?

Re: The Bitter Lesson Is Misunderstood

#176

Earlier quoted context omitted.

> So it makes me wonder, is embodiment (advanced robotics) 1000x harder than LLMs from an information processing perspective? Essentially, yes, but I would go further in saying that embodiment is harder than intelligence in and of itself. I would argue that intelligence is a very simple and primitive mechanism compared to the evolved animal body, and the effectiveness of our own intelligence is circumstantial. We man…

> We manage to dominate the world mainly by using brute force to simplify our environment and then maintaining and building systems on top of that simplified environment. If we didn't have the proper tools to selectively ablate our environment's complexity… This is very interesting and I feel there is a lot to unpack here. Could you elaborate on this theory with a few more paragraphs (or books / blogs that elucidate…

Not OP, but several examples:

A. instead of building a house on random terrain with random materials, first we prefer to flatten the place, then we use standard materials (e.g. bricks), which were produced from simple source (e.g. large and relatively homogenous deposit of clay).

B. For mental tasks it’s usual to said, that a person can handle only 7 items at a time (if you disagree multiply by 2-3). But when you ride a bike you process more inputs at the same time (you hear a car behind you, you see person on the right, you feel your balance, you anticipate your direction, if you feel strong wind or sun on your face you probably squint your eyes, you take a breath of air. On top of that all the processes of your body adjust and support your riding: heart, liver, stomach…)

C. “Spherical cows” in physics. (Google this if needed)

Re: The Bitter Lesson Is Misunderstood

#177
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

Embodiment is 1000x harder from a physical perspective.

Look at how hard it is for us to make reliable laptop hinges or the articulated car door handle trend (started by Tesla) where they constantly break.

These are simple mechanisms compared to any animal or human body. Our bodies last up to 80-100 years through not just constant regeneration but organic super-materials that rival anything synthetic in terms of durability within its spec range. Nature is full of this, like spider silk much stronger than steel or joints that can take repeated impacts for decades. This is what hundreds of millions to billions of years of evolution gets you.

We can build robots this good but they are expensive, so expensive that just hiring someone to do it manually is cheaper. So the problem is that good quality robots are still much more expensive than human labor.

The only areas where robots have replaced human labor is where the economics work, like huge volume manufacturing, or where humans can’t easily go or can’t perform. The latter includes tasks like lifting and moving things thousands of times larger than humans can or environments like high temperatures, deep space, the bottom of the ocean, radioactive environments, etc.

Re: The Bitter Lesson Is Misunderstood

#178

Is the data input into ChatGPT not a large enough source of new data to matter? People are constantly inputting novel data, telling ChatGPT about mistakes it made and suggesting approaches to try, and so on. For local tools, like claude code, it feels like there's an even bigger goldmine of data in that you can have a user ask claude code to do something, and when it fails they do it themselves... and then if only an…

NO! CLEARLY THE ENTIRE CORPUS OF HUMAN LITERATURE AND THE INTERNET DOESN'T CONTAIN ENOUGH INFORMATION TO EDUCATE AN EXPERT!!!! I JUST NEED ANOTHER BILLION DOLLARS PLS PLS PLS I PROMISE THE SCALING LAWS ARE ACTUALLY LAWS THIS TIME

Re: The Bitter Lesson Is Misunderstood

#179
post #106

Earlier quoted context omitted.

This seems so simple but I’m totally not understanding it.. If C = D^2, and you double compute, then 2C ==> 2D^2. How do you and the original author get 1.41D from 2D^2?

If C ~ D^2, then D ~ sqrt(C). In other words, the required amount of data scales with the square root of the compute. The square root of 2 ~= 1.414. If you double the compute, you need roughly 1.414 times more data.

Thanks for clarification!

Re: The Bitter Lesson Is Misunderstood

#180
post #134

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Now consider that the genius cannot physically interact with the world or the people therein, and uses her eyes only for reading text.

I'm not sure what point you are trying to make. Are you saying in order to make LLMs better at learning the missing piece is to make the capable to interact with the outside world? Give them actuators and sensors?
Post reply on HN