Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

221–230 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#221
post #60

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

Not a robotics guy, but to extent that the same fundamentals hold— I think it's a degrees of freedom question. Given the (relatively) low conditional entropy of natural language, there aren't actually that many degrees of (true) freedom. On the other hand, in the real world, there are massively more degrees of freedom both in general (3 dimensions, 6 degrees of movement per joint, M joints, continuous vs. discrete sp…

> in practice we currently have exponentially less real-world data for an exponentially harder problem

Is that where learning comes in? Any actual AGI machine will be able to learn. We should be able to buy a robot that comes ready to learn and we teach it all the things we want it to do. That might mean a lot of broken dishes at first, but it's about what you would expect if you were to ask a toddler to load your dishes into the dishwasher.

My personal bar for when we reach actual AGI is when it can be put in a robot body that can navigate our world, understand spatial relationships, and can learn from ordinary people.

Re: The Bitter Lesson Is Misunderstood

#222

Earlier quoted context omitted.

Plato's "Allegory of the cave" was uninteresting and uninformative when I first read it more than 50 years ago. It remains so today. https://en.wikipedia.org/wiki/Allegory_of_the_cave Also, other than in sculpture/dentistry/medicine I also find "ablation" to not be a particularly insightful metaphor either. Although I see ablation's application to LLMs I simply had to laugh when I first read about it: I envisioned st…

> Plato's "Allegory of the cave" was uninteresting and uninformative when I first read it more than 50 years ago. It remains so today. Anything can be uninteresting and uninformative when one doesn't see it's interestingness or can't grok its information. It however stood for millenia as a great device to describe multiple layers of abstractions, deeper reality vs appearance, and so on, with utility as such in countl…

No. the Allegory is a fragment of a poor unfinished story and little more. You don't need it to explain "multiple layers of abstractions, deeper reality vs appearance" as you say. In fact, you don't need it for anything at all except to explain Plato's "Allegory of the cave". Sheesh.

coldtea says "...with utility as such in countless domains." So when's the last time you referred to the "Allegory of the cave" in your day, other than on HN?

Re: The Bitter Lesson Is Misunderstood

#223

Earlier quoted context omitted.

I think we're talking past each other, I'll try once more. Suppose you train an LLM on a very small corpus of data, such as all the content of the library of congress. Then you have that LLM author new works. Then you train a new LLM on the original corpus plus this new material. Do you really think you've addressed the core issue in the SP? Can more parameters be meaningfully trained even if you add more GPU? To me,…

When it comes to logical reasoning, the difficulty isn't about having enough new information, but about ensuring the LLMs capture the right information. The problem LLMs have with learning logical reasoning from standard training is that they learn spurious relationships between the context and the next token, undermining its ability to learn fully general logical reasoning. Synthetic data helps because spurious asso…

I agree! DeepSeek has shown this is incredibly powerful. I think their Qwen 8B model may be as good as GPT4’s flagship. And I can run it on my laptop if it’s not on my lap. But the amount of synthetic data you can generate is bounded by the raw information, so I don’t think it’s an answer to the SP.

Re: The Bitter Lesson Is Misunderstood

#224

The AI companies won't run out of data to train on. Almost every user interaction is a significant source of data. Chains of interactions are even more significant, especially the longer and more sophisticated they are. Yesterday I was given A/B tests from both GPT5-Thinking and Gemini 2.5 Pro, something neither of then had done before. OpenAI also just acquired Statsig for $1.1 billion. Statsig does A/B testing and…

> it bootstrapped something that we all love talking to and discussing ANYTHING with.

We all? Speak for yourself, dude

Re: The Bitter Lesson Is Misunderstood

#225
post #128

Earlier quoted context omitted.

Plato's "Allegory of the cave" was uninteresting and uninformative when I first read it more than 50 years ago. It remains so today. https://en.wikipedia.org/wiki/Allegory_of_the_cave Also, other than in sculpture/dentistry/medicine I also find "ablation" to not be a particularly insightful metaphor either. Although I see ablation's application to LLMs I simply had to laugh when I first read about it: I envisioned st…

I don’t think that’s what ablation is about. It’s more like blowing parts off a bus until it ceases to be a bus. Then you find the minimal set of bus parts required to still be a bus, and that’s an indication that those parts are important to the central task of being a bus.

taneq SAYS "i don’t think that’s what ablation is about. It’s more like blowing parts off a bus until it ceases to be a bus."

Different people have different goals. You want some form of minimal bus and I want a Lotus 7. There's no guarantee either of us reach our goal.

Ablation is about disassembling something randomly, whether little by little or on an arbitrary scale until [SOMETHING INTERESTING OR DESIRABLE HAPPENS].

https://en.wikipedia.org/wiki/Ablation_(artificial_intellige...

Ablation is laughable but sometimes useful. It is also easy, mostly brainless, NOT guaranteed to provide any useful information (so you've an excuse for the wasted resources), and occasionally provides insight. It's a good tool for software engineers who have no (or seek no) understanding of their system, so I think of ablation as a "last resort" solutions (e.g., another being to randomly modify code until it "works") that I disdain.

But I'm old so I'm probably wrong! Burn those CPU towers down, boys and girls!

Re: The Bitter Lesson Is Misunderstood

#226

Earlier quoted context omitted.

It's napkin math so take it with a pinch of salt, but I am calculating the information stored in genome, assuming 2 bits per base pair, reducing to estimated 88% coding fraction to get the functional bits, and then dividing by cell volume. Did this for a few different types of cells and then averaged the result to around 1–10 Mbit/μm³ # If there are any bioinformaticians around please come eviscerate or confirm this…

You are comparing the fastest writable memory available (SRAM) vs biological non-volatile memory that is essentially read only. Samsung's 280 layer NAND reaches 28,5 Gbit per mm^2. I don't know how you would convert that to a volume, but if we simply multiply by 1000x for simplicity, it would be much closer to 0.19 Mbit/μm³, but even then you have to remember that NAND flash is still writable at pretty high speeds.

This is true, I was only comparing functional information density (unique functional genome bps), not read/write speeds.

I was also taking the information from a genome and then dividing it by the volume of a cell, but there are many instances of the genome in a cell. I didn't count all instances because they aren't unique.

There's a lot to unpack with this comparison and my approximation was crude, but the more I've dug into this comparison the more apparent how incredibly efficient life is at managing and processing and storing information. Especially if you also consider the amino acids, proteins, etc as information. No matter how you slice it, life seems orders of magnitude more efficient by every metric.

I'd like to think there's a paper somewhere where someone has carefully unpacked this and formally quantified it all.

Re: The Bitter Lesson Is Misunderstood

#227

Earlier quoted context omitted.

AlphaZero trained itself through chess games that it played with itself. Chess positions have something very close to an objective truth about the evaluation, the rules are clear and bounded. Winning is measurable. How do you achieve this for a language model? Yes, distillation is a thing but that is more about compression and filtering. Distillation does not produce new data in the same way that chess games produce…

Simple, you just need to turn language into a game. You make models talk to each other, create puzzles for each other's to solve, ask each other to make cases and evaluate how well they were made. Will some of it look like ramblings of pre-scientific philosophers? (or modern ones because philosophy never progressed after science left it in the dust) Sure! But human culture was once there too. And we pulled ourselves…

I feel like you’re glossing over some very thorny details that it’s not obvious we can solve. For example, if you just get two LLMs setting each other puzzles and scoring the others solutions how do you stop this just collapsing into nonsense? I.e. where does the source of actual truth come from for the puzzles?

Re: The Bitter Lesson Is Misunderstood

#228

Earlier quoted context omitted.

Note that wheels, steam engines, jet planes, spaceships wouldn't survive on their own in nature. Compared to natural structures, they are very simple, very straightforward. And while biological organisms are adapted to survive or thrive in complicated, ever-changing ecosystems, our machines thrive in sanitized environments. Wheels thrive on flat surfaces like roads, jet planes thrive in empty air devoid of trees, and…

Okay, but (1) we don't need to simulate physics faster than physics to make accurate-enough predictions to fly a plane, in our heads, or build a plane on paper, or to model flight in code. (2) If that's only because we've cleared out the trees and the Canada Geese and whatnot from our simplified model and "built the road" for the wheels, then necessity is also the mother of invention. "Hey, I want to fly but I keep c…

> we don't need to simulate physics faster than physics to make accurate-enough predictions to fly a plane

No, but that's only a small part of what you need to model. It won't help you negotiate a plane-saturated airspace, or avoid missiles being shot at you, for example, but even that is still a small part. Navigation models won't help you with supply chains and acquiring the necessary energy and materials for maintenance. Many things can -- and will -- go wrong there.

> In other words, why are we assuming that agents cannot shape the world

I'm not assuming anything, sorry if I'm giving the wrong impression. They could. But the "shapability" of the world is an environment constraint, it isn't fully under the agent's control. To take the paper clipper example, it's not operating with the same constraints we are. For one, unlike us (notwithstanding our best efforts to do just that), it needs to "simplify" humanity. But humanity is a fast, powerful, reactive, unpredictable monster. We are harder to cut than trees. Could it cull us with a supervirus, or by destroying all oxygen, something like that? Maybe. But it's a big maybe. Such brute force takes requires a lot of resources, the acquisition of which is something else it has to do, and it has to maintain supply chains without accidentally sabotaging them by destroying too much.

So: yes. It's possible that it could do that. But it's not easy, especially if it has to "simplify" humans. And when we simplify, we use our animal intelligence quite a bit to create just the right shapes. An entity that doesn't have that has a handicap.

Re: The Bitter Lesson Is Misunderstood

#229

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

Maybe human brains are constantly generating (and training on) massive amounts of synthetic data and that is how they get so smart?

You mean those like 8 hours of ~~nightmares~~ dreams I have every night?

Re: The Bitter Lesson Is Misunderstood

#230

Earlier quoted context omitted.

I believe he is referring to OpenAI proposal to move beyond training with pure text. Instead train with multi modal data. Instead of only the dictionary definition of an apple. Train it with a picture of an apple. Train it with a video of someone eating an apple etc.

Before this AI wave got going, I'd always assumed that AGI would be more about converting words, pictures, video, and lots of sensory data and who knows what else into a model of concepts that it would be putting together and hypothesizing about and testing as it grows. A database of what concepts have been learned and what data they were built from and what holes it needed to fill in. It would continually be working…

When I was young, my relatives would make fun of me. Saying I had a lot of book learning but yet to experience the absurdity of the real world. Wait, they said, when I try to apply my fancy book learning to a world controlled by good ole boys, gatekeepers, and double talk. Then I will learn reality is different from the idealized world of books.
Post reply on HN