Live data from Hacker News

The Bitter Lesson Is Misunderstood

obviouslywrong.substack.com

211–220 of 259 posts

Re: The Bitter Lesson Is Misunderstood

#211
post #64

Earlier quoted context omitted.

What diamond symbol?

Oh its a dot. Dots, diamonds,the absense of an operator, anything is multiplication it seems. While this comment might look like a paragraph, its actually a lot of maths.

But diamonds don't denote multiplication. That never happened. That was just you misreading.

Re: The Bitter Lesson Is Misunderstood

#212
post #60

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

Not a robotics guy, but to extent that the same fundamentals hold— I think it's a degrees of freedom question. Given the (relatively) low conditional entropy of natural language, there aren't actually that many degrees of (true) freedom. On the other hand, in the real world, there are massively more degrees of freedom both in general (3 dimensions, 6 degrees of movement per joint, M joints, continuous vs. discrete sp…

As I could see, classic methods (used in children teaching) could create at least magnitude more data than we have now, just paraphrasing text (classic NLP), but depends on language (I'll try explain).

Text really have lot of degrees of freedom, but depends on language, and even more on type of alphabet - modern English with phonetic alphabet is worst choice, because it is simplest, nearly nobody use second-third hidden meaning (I hear about 2-3 to 5-6 meanings depending on source); hieroglyphic languages are much more information rich (10-22 meanings); and what is interest, phonetic languages in totalitarian countries (like Russian) are also much more rich (8-12 meanings), because they used to hide few meanings from government to avoid punishment.

Language difference (more dimensions) could be explanation of current achievements of China, superior to Western, and it could also be hint, on how to boost Western achievements - I mean, use more scientists from Eastern Europe and give more attention to Eastern European languages.

For 3D robots, I see only one way - computational simulated environment.

Re: The Bitter Lesson Is Misunderstood

#213
post #49

Hey folks, OOP/original author and 20-year HN lurker here — a friend just told me about this and thought I'd chime in. Reading through the comments, I think there's one key point that might be getting lost: this isn't really about whether scaling is "dead" (it's not), but rather how we continue to scale for language models at the current LM frontier — 4-8h METR tasks. Someone commented below about verifiable rewards…

My focus lately is on the cost side of this. I believe strongly that it's possible to reduce the cost of compute for LLM type loads by 95% or more. Personally, it's been incredibly hard to get actual numbers for static and dynamic power in ASIC designs to be sure about this.

If I'm right (which I give a 50/50 odds to), and we can reduce the power of LLM computation by 95%, trillions can be saved in power bills, and we can break the need for Nvidia or other specialists, and get back to general purpose computation.

Re: The Bitter Lesson Is Misunderstood

#214

Just using common sense, if we had a genius, who had tremendous reasoning ability, total recall of memories, and an unlimited lifespan and patience, and he'd read what the current LLMs have read, we'd expect quite a bit more from him than what we're getting now from LLMs. There are teenagers that win gold medals on the math olympiad - they've trained on In other words, data scarcity is not a fundamental problem, just…

I think quantization is the simplest canary. If we can reduce the precision of the model parameters by 2~32x without much perceptible drop in performance, we are clearly dealing with something wildly inefficient. I'm open to the possibility that over parameterization is essential as part of the training process, much like how MSAA/SSAA over sample the frame buffer to reduce information aliasing in the final scaled re…

It’s not clear that the inefficiency of the current paradigm is in the neural net architectures. It seems just as likely that it’s in the training objective.

Re: The Bitter Lesson Is Misunderstood

#215

I don't think anyone has yet trained on all videos on the Internet. Plenty of petabytes left there to pretrain on, and likely just as useful once the text/audio/image pretraining is done.

It might have been trained on a select high quality of videos, say more than 10k views and only trained on its transcripts.

Seems like reading a transcript of the commentary from a football game, it’s obviously missing a lot of information.

Re: The Bitter Lesson Is Misunderstood

#216
isn’t the problem of not enough data just a problem of not having grounding in the world? world models and everything feel like they’re just dancing around the problem with thin veneers of human-in-the-loop and the verifiable domains we already have

Re: The Bitter Lesson Is Misunderstood

#217

Humans require a _lot_ less training data to become, for instance, fluent in English. If a given AI algorithm needs to be trained on the entire Internet to accomplish the same, then it seems safe to assume that the data has not really been "mined out". Generating more training data from the same original data should not be fundamentally problematic in that sense.

humans also have billions of years of evolution and trillions of organisms to develop a receptacle biased towards learning language

Re: The Bitter Lesson Is Misunderstood

#218
post #67

Earlier quoted context omitted.

What do you mean about CLIP?

I believe he is referring to OpenAI proposal to move beyond training with pure text. Instead train with multi modal data. Instead of only the dictionary definition of an apple. Train it with a picture of an apple. Train it with a video of someone eating an apple etc.

Before this AI wave got going, I'd always assumed that AGI would be more about converting words, pictures, video, and lots of sensory data and who knows what else into a model of concepts that it would be putting together and hypothesizing about and testing as it grows. A database of what concepts have been learned and what data they were built from and what holes it needed to fill in. It would continually be working on this and reaching out to test reality or discuss it's findings with people or other AIs instead of waiting for input like a chatbot. I haven't even seen anything like this yet, just ways of faking it by getting better at stringing words together or mashing pixels together based on text tokens.

No one seems to be working on building an AI model that understands, to any real degree, what it's saying or what it's creating. Without this, I don't see how they can even get to AGI.

Re: The Bitter Lesson Is Misunderstood

#219
post #217

Humans require a _lot_ less training data to become, for instance, fluent in English. If a given AI algorithm needs to be trained on the entire Internet to accomplish the same, then it seems safe to assume that the data has not really been "mined out". Generating more training data from the same original data should not be fundamentally problematic in that sense.

humans also have billions of years of evolution and trillions of organisms to develop a receptacle biased towards learning language

Billions of years of evolution, but still limited to the data that is replicated in human genome/DNA, which is about 3 gigabytes (+epigenome).

Re: The Bitter Lesson Is Misunderstood

#220

Earlier quoted context omitted.

10+ years ago I expected we would get AI that would impact blue collar work long before AI that impacted white collar work. Not sure exactly where I got the impression, but I remember some "rising tide of AI" analogy and graphic that had artists and scientists positioned on the high ground. Recently it doesn't seem to be playing out as such. The current best LLMs I find marvelously impressive (despite their flaws), a…

The problem is not the robot loading the diswasher, it is the dishwasher. The dishwasher (and general kitchen electronics) industry has not innovated in a long time. My prediction is a new player will come in who vertically integrates these currently disjoint industries and product. The tableware used should be compatible with the dishwasher, the packaging of my groceries should be compatible with the cooking system.…

> the whole notion of putting one room of your apartment full with random electronics just to cook a meal once in a blue moon is deeply inefficient

You don't use your kitchen? After the rooms we sleep in, the kitchen is probably the most used space in my home. We are planning an upcoming renovation of our home and the kitchen is where we plan on spending the most money.

> The tableware used should be compatible with the dishwasher

Aside from non-dishwasher safe items, what tableware is incompatible with a dishwasher?

Post reply on HN